Video data processing method and device, computer equipment and storage medium

By extracting and grouping images of videos, combining three-dimensional convolutional networks and self-attention models to extract features and performing fusion processing, the problem of insufficient robustness of video analysis in the prior art is solved, and a more stable and accurate video data processing effect is achieved.

CN120236221APending Publication Date: 2025-07-01SF TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311870523.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-30
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing video analysis algorithms are poorly robust when processing the relationship between video frames, resulting in unstable processing effects.

Method used

By performing image frame extraction and grouping processing on the target video, an image frame packet is constructed, and the features of the image frame packet are extracted using the three-dimensional convolutional network model and the self-attention model, and feature fusion and full connection processing are performed to obtain the video data processing results.

Benefits of technology

It improves the robustness of the video data analysis process and enhances the stability and accuracy of the algorithm when processing different video data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236221A_ABST
    Figure CN120236221A_ABST
Patent Text Reader

Abstract

The invention relates to a video data processing method and device, computer equipment, a storage medium and a computer program product. The method comprises the following steps: performing image frame extraction processing on a target video to obtain a video image frame; performing grouping processing on the video image frames according to the playing sequence, and determining an image frame packet formed by the video image frames in the same group; performing feature extraction processing on the image frames in each image frame packet through a three-dimensional convolutional network model to obtain three-dimensional convolutional features of the target video, and performing feature extraction processing on the image frames in each image frame packet through a self-attention model to obtain self-attention features of the target video; performing feature fusion processing on the three-dimensional convolution features and the self-attention features to obtain video fusion features; and performing full connection processing on the video fusion features to obtain a video data processing result. According to the method, the known information in the target video can be fully utilized to improve the algorithm precision, so that the robustness of the video data processing process is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular, to a video data processing method, apparatus, computer device, storage medium, and computer program product. Background Art

[0002] With the development of computer technology, artificial intelligence technology has emerged. It is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. Currently, video analysis can be performed through artificial intelligence technology. Video analysis is an important branch in the field of computer vision, aiming to obtain some key information through video data analysis. However, due to the variety of video types and the need to consider the relationship between video frames, video analysis has become one of the more difficult tasks in computer vision.

[0003] Most current video analysis algorithms use methods such as detection, segmentation, classification, and key points to process each single-frame image of the video separately to obtain corresponding processing results, and then add complex prior logical relationships to finally obtain the required video analysis results. However, this method has poor robustness. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a video data processing method, apparatus, computer device, computer-readable storage medium, and computer program product that can improve the robustness of the video data analysis process.

[0005] In a first aspect, the present application provides a video data processing method, including:

[0006] Performing image frame extraction processing on a target video to obtain video image frames;

[0007] Grouping the video image frames in the playing order to determine image frame packets composed of the video image frames with the same grouping;

[0008] Performing feature extraction processing on the image frames in each image frame packet through a three-dimensional convolutional network model to obtain three-dimensional convolutional features of the target video, and performing feature extraction processing on the image frames in each image frame packet through a self-attention model to obtain self-attention features of the target video. The three-dimensional convolutional network model is trained through image data with a time dimension in historical data, and the self-attention model is trained through image data fused with pixel means;

[0009] Performing feature fusion processing on the three-dimensional convolutional features and the self-attention features to obtain video fusion features;

[0010] Perform a fully connected process on the video fusion features to obtain the video data processing result.

[0011] In one embodiment, the extracting image frames from the target video to obtain video image frames includes:

[0012] Extract image frames from the target video based on a preset frame capture rate to obtain video image frames;

[0013] The grouping process of the video image frames in the playing order to determine the image frame packets composed of the video image frames with the same grouping includes:

[0014] Perform an average grouping process on the video image frames in chronological order to determine the image frame packets composed of the video image frames with the same grouping.

[0015] In one embodiment, the extracting features of the image frames in each image frame packet through a three-dimensional convolutional network model to obtain the three-dimensional convolutional features of the target video includes:

[0016] Extract image features of each image frame in the image frame packet through a three-dimensional convolutional network model to obtain an image frame packet convolutional feature including convolutional kernel number information, the number of images in the image frame packet information, and image frame feature information;

[0017] Perform a dimension mapping process on the image frame packet convolutional feature to compress the feature dimensions of the convolutional kernel number information and the image number information to obtain the three-dimensional convolutional features of the target video.

[0018] In one embodiment, the extracting features of the image frames in each image frame packet through a self-attention model to obtain the self-attention features of the target video includes:

[0019] Perform an average processing of the pixel values of the image frames in each image frame packet to obtain a mean fusion image of the image frame packet;

[0020] Segment the mean fusion image into fusion sub-images;

[0021] Expand the fusion sub-images based on RGB to obtain sub-image feature vectors;

[0022] Perform a linear embedding process on the sub-image feature vectors through a fully connected layer to obtain self-attention input features;

[0023] Perform self-attention processing and fully connected processing on the self-attention input features through a self-attention model to obtain self-attention features that match the dimensions of the three-dimensional convolutional features.

[0024] In one embodiment, the process of obtaining the mean fusion image of the image frame packet by performing pixel value averaging on each image frame within the image frame packet includes:

[0025] Obtain the channel pixel value data of each image frame in the image frame packet at each pixel point;

[0026] Perform average processing on each channel pixel value at the same pixel point position respectively to obtain the average channel pixel value at each pixel point position;

[0027] Combine the average channel pixel values at each pixel point position to obtain the mean fusion image of the image frame packet.

[0028] In one embodiment, the video data processing result includes a video data classification result;

[0029] The process of obtaining the video data processing result by performing fully connected processing on the video fusion feature includes:

[0030] Perform fully connected processing on the video fusion feature to obtain the fully connected feature;

[0031] Based on the fully connected feature, perform classification processing on the target video to obtain a video data classification result.

[0032] In a second aspect, the present application also provides a video data processing device, including:

[0033] A frame extraction processing module, configured to perform image frame extraction processing on a target video to obtain video image frames;

[0034] A frame grouping module, configured to perform grouping processing on the video image frames in the playing order to determine the image frame packets composed of the video image frames with the same grouping;

[0035] A feature extraction module, configured to perform feature extraction processing on each image frame in each image frame packet through a three-dimensional convolutional network model to obtain the three-dimensional convolutional features of the target video, and perform feature extraction processing on each image frame in each image frame packet through a self-attention model to obtain the self-attention features of the target video. The three-dimensional convolutional network model is trained through image data with a time dimension in historical data, and the self-attention model is trained through image data with pixel mean fusion;

[0036] A feature fusion module, configured to perform feature fusion processing on the three-dimensional convolutional features and the self-attention features to obtain video fusion features;

[0037] A fully connected processing module, configured to perform fully connected processing on the video fusion features to obtain video data processing results.

[0038] In a third aspect, the present application also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0039] Performing image frame extraction processing on a target video to obtain video image frames;

[0040] Grouping the video image frames in the playing order to determine image frame packets composed of the video image frames with the same grouping;

[0041] Performing feature extraction processing on the image frames in each image frame packet through a three-dimensional convolutional network model to obtain three-dimensional convolutional features of the target video, and performing feature extraction processing on the image frames in each image frame packet through a self-attention model to obtain self-attention features of the target video. The three-dimensional convolutional network model is trained through image data with a time dimension in historical data, and the self-attention model is trained through image data fused by pixel means;

[0042] Performing feature fusion processing on the three-dimensional convolutional features and the self-attention features to obtain video fusion features;

[0043] Performing fully connected processing on the video fusion features to obtain a video data processing result.

[0044] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0045] Performing image frame extraction processing on a target video to obtain video image frames;

[0046] Grouping the video image frames in the playing order to determine image frame packets composed of the video image frames with the same grouping;

[0047] Performing feature extraction processing on the image frames in each image frame packet through a three-dimensional convolutional network model to obtain three-dimensional convolutional features of the target video, and performing feature extraction processing on the image frames in each image frame packet through a self-attention model to obtain self-attention features of the target video. The three-dimensional convolutional network model is trained through image data with a time dimension in historical data, and the self-attention model is trained through image data fused by pixel means;

[0048] Performing feature fusion processing on the three-dimensional convolutional features and the self-attention features to obtain video fusion features;

[0049] Performing fully connected processing on the video fusion features to obtain a video data processing result.

[0050] In a fifth aspect, the present application also provides a computer program product, including a computer program, which when executed by a processor implements the following steps:

[0051] Performing image frame extraction processing on a target video to obtain video image frames;

[0052] Grouping the video image frames in the playing order to determine image frame packets composed of the video image frames with the same grouping;

[0053] Performing feature extraction processing on the image frames in each image frame packet through a three-dimensional convolutional network model to obtain three-dimensional convolutional features of the target video, and performing feature extraction processing on the image frames in each image frame packet through a self-attention model to obtain self-attention features of the target video. The three-dimensional convolutional network model is trained by image data with a time dimension in historical data, and the self-attention model is trained by image data fused with pixel means;

[0054] Performing feature fusion processing on the three-dimensional convolutional features and the self-attention features to obtain video fusion features;

[0055] Performing fully connected processing on the video fusion features to obtain a video data processing result.

[0056] The above video data processing method, device, computer device, storage medium, and computer program product perform image frame extraction processing on a target video to obtain video image frames; group the video image frames to determine image frame packets composed of the video image frames with the same grouping; perform feature extraction processing on the image frames in each image frame packet through a three-dimensional convolutional network model to obtain three-dimensional convolutional features of the image frame packet, and perform feature extraction processing on the image frames in each image frame packet through a self-attention model to obtain self-attention features of the image frame packet; perform feature fusion processing on the three-dimensional convolutional features and the self-attention features to obtain video fusion features; perform fully connected processing on the video fusion features to obtain a video data processing result. The present application constructs image frame packets through image frame extraction and image frame grouping processing, thereby obtaining image frame inputs in different time dimensions, and then extracts the features of the image frame packets through a three-dimensional convolutional network model and a self-attention model respectively, and performs feature fusion processing and fully connected processing on the extracted features to obtain a video data processing result, making full use of the known information in the target video to further improve the algorithm accuracy and enhance the robustness of the video data processing process. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0058] Figure 1 It is an application environment diagram of the video data processing method in an embodiment;

[0059] Figure 2 It is a schematic flowchart of the video data processing method in an embodiment;

[0060] Figure 3 It is a schematic flowchart of the image frame extraction step in an embodiment;

[0061] Figure 4 It is a schematic flowchart of the image frame grouping step in an embodiment;

[0062] Figure 5 It is a schematic flowchart of the 3D convolution step in an embodiment;

[0063] Figure 6 It is a schematic flowchart of the 3D convolution feature transformation step in an embodiment;

[0064] Figure 7 It is a schematic flowchart of the self-attention feature extraction step in an embodiment;

[0065] Figure 8 It is a schematic flowchart of the self-attention feature extraction step in another embodiment;

[0066] Figure 9 It is a schematic flowchart of the feature fusion and fully connected processing step in an embodiment;

[0067] Figure 10 It is a structural block diagram of the video data processing device in an embodiment;

[0068] Figure 11 It is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0069] In order to make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the following further details the present application in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0070] The video data processing method provided by the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed on the cloud or other network servers. When a user hopes to process video data, such as classifying a specified video, etc., the target video to be processed can be submitted to the server 104 through the terminal 102, and the server 104 will perform relevant processing on the video data. After obtaining the target video, the server 104 will first perform image frame extraction processing on the target video to obtain video image frames; group the video image frames in the playing order to determine the image frame packets composed of the video image frames with the same grouping; perform feature extraction processing on the image frames in each image frame packet through a three-dimensional convolutional network model to obtain the three-dimensional convolutional features of the target video, and perform feature extraction processing on the image frames in each image frame packet through a self-attention model to obtain the self-attention features of the target video; perform feature fusion processing on the three-dimensional convolutional features and the self-attention features to obtain video fusion features; perform fully connected processing on the video fusion features to obtain the video data processing result. Among them, the terminal 102 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0071] In an exemplary embodiment, as Figure 2 shown, a video data processing method is provided. Taking the server 104 in Figure 1 as an example for illustration, it includes the following steps 201 to step 209. Among them:

[0072] Step 201, perform image frame extraction processing on the target video to obtain video image frames.

[0073] Step 203, group the video image frames in the playing order to determine the image frame packets composed of the video image frames with the same grouping.

[0074] Among them, the target video is the video data processed by the video data processing method of this application. Image frame extraction processing refers to extracting individual image frames from a video sequence. That is, some frames in the video are cut out to form static images. These pictures can be used to make GIF animations, video clips, etc. Frame extraction can reduce the file size by reducing the number of frames in the video file without affecting the video quality. This is because frame extraction only selectively extracts some frames from the video without changing the time interval between frames. Grouping processing refers to processing the video image frames obtained by frame extraction in order, and each group contains multiple image frames arranged in order.

[0075] Exemplarily, when a user hopes to process a target video, such as summarizing the target video or classifying the target video, the video-related processing can be implemented through the solution of this application. The user can directly submit the target video to the server 104 through the terminal 102 to start the corresponding video data processing process. After receiving the video data, the server 104 can start the corresponding video data processing. First, the server 104 performs image frame extraction processing on the target video to obtain video image frames, and through image frame extraction, the video-type data is converted into image-type data for processing. After obtaining the video image frames, the video image frames can be grouped and processed in the playing order to obtain an image frame packet composed of each group of image frames. Through the image frame packet obtained by grouping, the time dimension information between adjacent and relatively distant video frames can be effectively utilized, thereby improving the robustness of the video data processing process.

[0076] Step 205, perform feature extraction processing on the image frames in each image frame packet through a three-dimensional convolutional network model to obtain the three-dimensional convolutional features of the target video, and perform feature extraction processing on the image frames in each image frame packet through a self-attention model to obtain the self-attention features of the target video. The three-dimensional convolutional network model is trained through the image data with a time dimension in historical data, and the self-attention model is trained through the image data fused by pixel means.

[0077] Among them, the three-dimensional convolutional network model, i.e., the 3D convolutional network model, compared with the 2D convolutional neural network, has one more time dimension in the input of the 3D convolutional network than the 2D one. That is to say, it can be regarded as performing 2D convolution on multiple pictures (or feature maps) simultaneously. During processing, by setting a 3D convolutional kernel containing the time dimension, in order to consider the correlation relationship of the entire video sequence at the same time, the input of the 3D convolution is the i-th picture in each image frame packet. Through the 3D convolutional network model, the features corresponding to each image frame packet are finally obtained from all the image frame packets of the target video, and then these features are processed through transformations such as dimensionality reduction to obtain the three-dimensional convolutional features of the target video. As for the self-attention model, it refers to a machine learning model constructed based on the self-attention mechanism. The basic idea of the self-attention mechanism is that when processing sequence data, each element can establish a connection with other elements in the sequence, rather than only relying on elements in adjacent positions. It adaptively captures the long-range dependence relationship between elements by calculating the relative importance between elements. The method of this application can effectively extract the long-range dependence relationship between the image frame packets extracted from the target video through the self-attention mechanism.

[0078] Exemplarily, the solution of this application specifically extracts the feature information in the target video through the three-dimensional convolutional network model and the self-attention model, so as to realize the processing of video data. Among them, for the feature extraction process of the three-dimensional convolutional network model, the feature information in each picture can be extracted through the three-dimensional convolutional network model, and then combined with the number dimension of the images in the image frame packet and the three-dimensional convolutional kernel with the time dimension to realize the preliminary feature extraction process, and then the extracted features are processed by dimensionality reduction to obtain the three-dimensional convolutional features containing time dimension information. As for the feature extraction process of the self-attention model, the images in the image frame packet can be processed by image mean fusion, so as to convert the image frame packet containing multiple images into a fused image, and then the self-attention features of the target video are extracted based on the fused image. In one embodiment, before the feature extraction process, the training processes of the three-dimensional convolutional network model and the self-attention model are also completed respectively. Among them, for the training process of the three-dimensional convolutional network model, the image data with time dimension in the historical data, which is also extracted from the video data and has corresponding data labels according to the task type of video data processing, can be used to complete the training and testing of the three-dimensional convolutional network model by dividing the data into training set, test set and validation set, etc., to obtain an available three-dimensional convolutional network model with adjusted parameters and ensure the effectiveness of the model. As for the self-attention model, it also includes a corresponding model training process. The image data for pixel mean fusion can be constructed by extracting frames and grouping from the historical video data, and then the training of the self-attention model is completed by these image data.

[0079] Step 207: Perform feature fusion processing on the 3D convolutional feature and the self-attention feature to obtain a video fusion feature.

[0080] Step 209: Perform a fully connected processing on the video fusion feature to obtain a video data processing result.

[0081] Among them, feature fusion processing refers to combining the 3D convolutional feature and the self-attention feature to obtain a new feature that can contain the complete information of both features. Feature fusion processing can be specifically implemented through bitwise addition operations. However, this processing requires ensuring that the feature dimensions of the 3D convolutional feature and the self-attention feature are the same. Therefore, during the feature extraction process of the 3D convolutional feature and the self-attention feature, feature dimension adjustment processes such as feature dimensionality reduction are also included to ensure the dimensional unity of these two features. Fully connected processing is to transform the video fusion feature into the final output result through a fully connected layer. The fully connected layer is a common layer type in neural networks, also known as a Dense Layer or a Fully Connected Layer. The fully connected layer can perform matrix multiplication and bias addition operations on the connection weights between the input feature and each neuron to obtain the output result. The role of the fully connected layer is to map the input feature to the output result, usually used in the last layer of a neural network for tasks such as classification and regression. The output result of the fully connected layer can be regarded as a non-linear transformation of the input feature. This transformation can map the input feature space to the output result space, thereby realizing the complexity and non-linear fitting ability of the model.

[0082] Exemplarily, after obtaining the 3D convolutional feature and the self-attention feature, the two features can be fused into the final video feature through feature fusion. This video feature can represent all the feature information of the target video, including not only the image information within the video frames extracted from the video frames, but also the time dimension information of the video frame packets and the long-range dependence features before and after the video frame packets. By performing fully connected processing on the video fusion feature, a video data processing result is obtained. In the solution of this application, the fully connected layer can be specifically designed according to the specific type of video processing task. For example, it can achieve video classification, video compliance, or video summarization, etc. Model training data can be constructed by searching historical video data in advance according to the video data processing task to be completed, so as to complete the training of each sub-model such as the 3D convolutional network model, the self-attention model, the feature fusion model, and the fully connected layer, and obtain a final usable complete video processing model. Then, according to the processing requirements of the video task, video processing functions such as video frame extraction and video frame grouping are encapsulated in the model to form a complete dedicated video processing model, and then the current video processing task is implemented through the video processing model.

[0083] The above video data processing method obtains video image frames by performing image frame extraction processing on a target video; performs grouping processing on the video image frames to determine image frame packets composed of video image frames with the same grouping; performs feature extraction processing on the image frames in each image frame packet through a three-dimensional convolutional network model to obtain the three-dimensional convolutional features of the image frame packet, and performs feature extraction processing on the image frames in each image frame packet through a self-attention model to obtain the self-attention features of the image frame packet; performs feature fusion processing on the three-dimensional convolutional features and the self-attention features to obtain video fusion features; and performs fully connected processing on the video fusion features to obtain the video data processing result. In this application, image frame extraction and image frame grouping processing are used to construct image frame packets, so as to obtain image frame inputs in different time dimensions. Then, the three-dimensional convolutional network model and the self-attention model are used to extract the features of the image frame packets respectively, and the extracted features are subjected to feature fusion processing and fully connected processing to obtain the video data processing result, making full use of the known information in the target video to further improve the algorithm accuracy and enhance the robustness of the video data processing process.

[0084] In an exemplary embodiment, step 201 includes: performing image frame extraction processing on the target video based on a preset frame capture rate to obtain video image frames. Step 203 includes: performing average grouping processing on the video image frames in the playback order to determine image frame packets composed of video image frames with the same grouping.

[0085] Exemplarily, when performing video frame extraction sampling and video grouping processing, the corresponding frame extraction processing can be specifically performed according to the preset frame capture rate. The frame capture rate can be used to determine the number of images collected per unit time and the interval of image collection, so that the target image can be subjected to image extraction according to the interval of image collection. After the image frame extraction processing is completed and a sequence of video image frames is obtained, the video image frames can be subjected to average grouping processing in the playback order to determine image frame packets composed of video image frames with the same grouping. Through the average grouping processing, it can be ensured that each image frame packet contains the same number of image frames arranged in the playback order, thus ensuring the consistency of the feature dimensions in the subsequent processing process. In one embodiment, the processing process of performing image extraction on the continuous image frames collected from the target video F_video according to a certain frame capture rate F_select can be referred to Figure 3 as shown. Frame extraction can obtain the video ancient town data arranged in order, and the process of performing average grouping processing on the video image frames can be referred to Figure 4As shown, the collected image frames are evenly divided into N groups, and each group of image frames is combined into an image frame packet I_group. In this embodiment, the target video is subjected to image frame extraction processing through the frame capture rate, and then the extracted video image frames are evenly grouped, which can effectively ensure the consistency of the video frame packet structure and retain the order and time information of the video frame packet, thereby effectively improving the effectiveness and accuracy of the subsequent feature extraction and data processing processes and improving the robustness of the data processing process.

[0086] In an exemplary embodiment, the image frames in each image frame packet are subjected to feature extraction processing through a three-dimensional convolutional network model to obtain the three-dimensional convolutional features of the target video, including: the image frames in the image frame packet are subjected to image feature extraction processing through the three-dimensional convolutional network model to obtain the convolutional features of the image frame packet, and the convolutional features of the image frame packet include the number of convolutional kernel information, the number of images in the image frame packet information, and the image frame feature information; by performing dimensional mapping processing on the convolutional features of the image frame packet, the feature dimensions of the number of convolutional kernel information and the number of image information are compressed to obtain the three-dimensional convolutional features of the target video.

[0087] Among them, the convolutional kernel is a commonly used tool in machine learning and computer vision for performing convolutional operations on data such as images, audio, and video. It can perform element-by-element multiplication operations with the original data and add the results to obtain a new value. The size and shape of the convolutional kernel can be adjusted according to needs to better capture the features in the data. In computer vision, convolutional kernels are usually used for image processing, such as edge detection, blurring, and sharpening. The role of the convolutional kernel is to extract the features in the image by performing convolutional operations on each pixel of the image. The size and shape of the convolutional kernel can be adjusted according to needs to better capture the features in the image. In the two-dimensional convolutional network for image processing, the output of the convolutional network is a two-dimensional input containing both length and width information, and the convolutional kernel is a two-dimensional matrix. When applied to a three-dimensional convolutional network, the input of the three-dimensional convolutional network has one more time dimension than the corresponding two-dimensional convolutional network, so a three-dimensional convolutional kernel containing the time dimension needs to be set correspondingly. The purpose of dimensional mapping processing is to map the convolutional features of the image frame packet obtained by three-dimensional convolutional processing into a fixed dimension, thereby getting rid of the problem of inconsistent feature dimensions caused by differences in video length.

[0088] Specifically, for the feature extraction process of the three-dimensional convolutional network model, it is similar to the process of two-dimensional convolution. The difference is that the input to the three-dimensional convolutional network model is multiple image frame packets, and each image frame packet contains multiple image frames. The input has one more time dimension than two-dimensional convolution, that is, it can be regarded as performing 2D convolution on multiple pictures (or feature maps) simultaneously. Therefore, for the corresponding N input image frame packets, a 3D convolutional kernel with a time dimension of N can be set. At the same time, in order to consider the correlation relationship of the entire video sequence, the input of the three-dimensional convolution is the i-th picture of each image frame packet. After the input is completed, the three-dimensional convolutional network model can be used to perform image feature extraction processing on each image frame in the image frame packet, and an image frame packet convolution feature containing three types of information: convolutional kernel quantity information, the number of images in the image frame packet, and image frame feature information can be obtained. In one embodiment, as Figure 5 shown, after being processed by the three-dimensional convolutional network model, a feature map of size (m*k)*h*w is finally generated as the image frame packet convolution feature, where m is the number of images contained in each image frame packet, k is the number of three-dimensional convolutional kernels, and h and w are the height and width of the generated feature map, that is, the image frame feature information. After generating the image frame packet convolution feature, since the total number of image frames contained in each video is different, that is, the number of images in the corresponding image frame packet is different, it also makes the extracted feature dimensions different (corresponding to m in the above content). Therefore, in order to make the extracted feature dimensions consistent for all videos, dimension mapping processing can be added after the three-dimensional convolution, so as to compress the feature dimensions of the convolutional kernel quantity information and the image quantity information. For example, a two-dimensional convolution can be added to compress the (m*k)*h*w image frame packet convolution feature into a three-dimensional convolution feature with a fixed dimension of c*h*w. The complete processing process can refer to Figure 6 shown. In this embodiment, through the image frame feature extraction processing and dimension mapping of the three-dimensional convolutional kernel, on the basis of extracting the features containing the time dimension, the dimensional unity of the output three-dimensional convolution features can be maintained, which can effectively ensure the efficiency and accuracy of feature extraction.

[0089] In an exemplary embodiment, as Figure 7 shown, through the self-attention model, feature extraction processing is performed on the image frames in each image frame packet, and the self-attention features of the target video obtained include:

[0090] Step 702, perform pixel value averaging processing on the image frames in each image frame packet to obtain the mean fusion image of the image frame packet.

[0091] Step 704, divide the mean fusion image into fusion sub-images.

[0092] Step 706, expand the fusion sub-images based on RGB to obtain sub-image feature vectors.

[0093] Step 708: Perform linear embedding processing on the sub-image feature vectors through a fully connected layer to obtain self-attention input features.

[0094] Step 710: Perform self-attention processing and fully connected processing on the self-attention input features through a self-attention model to obtain self-attention features that match the three-dimensional convolution feature dimension.

[0095] Among them, pixel value averaging processing means that for all image frames in the image frame packet, the pixel values of the pixel points at the same position in these image frames are averaged. After performing pixel value averaging processing on the image frame packet, the image frame packet can be converted into a mean fusion image. For channels, to describe a pixel point, if it is a grayscale image, only one value is needed to represent it, which is a single channel. If a pixel point is described by three colors RGB, it is a three-channel. The fused sub-image is unfolded based on RGB to obtain the three-channel feature values of each point in the fused sub-image, and the image feature vector corresponding to the fused sub-image is constructed. Linear embedding processing, that is, liner embeding, can be performed by inputting the sub-image feature vector into the linerembeding layer for linear transformation processing, so as to enhance feature expression and obtain self-attention input features.

[0096] Exemplarily, for the process of self-attention processing, when performing self-attention processing on the image frame packet, first, pixel value averaging processing needs to be performed on the image frames in each image frame packet to convert the image frame packet into a mean fusion image, and then self-attention processing is performed on the basis of the multiple mean fusion images obtained by converting the target video. The processing process first divides the mean fusion image into fused sub-images, N mean-fused images, then cuts the images into fused sub-images of size n*n, and then unfolds each fused sub-image into sub-image feature vectors according to the channel (rgb channel) dimension and sends them into the fully connected layer for linear embedding processing, and sends the obtained self-attention input features into the self-attention model. Finally, output features of dimension S are obtained, and finally, after another processing of the fully connected layer, the features obtained by self-attention processing are mapped into features of length c*h*w. After that, through resize processing, the obtained features can be converted into a feature map with the same shape as the three-dimensional convolution feature. The complete self-attention conversion processing process can be referred to Figure 8 as shown. In this embodiment, through the extraction and conversion processing of self-attention features, the consistency between the obtained self-attention features and the three-dimensional convolution features can be effectively guaranteed, thereby ensuring the accuracy of feature fusion processing.

[0097] In an exemplary embodiment, step 702 includes: obtaining the channel pixel value data of each image frame at each pixel point within the image frame packet; performing an averaging process on each channel pixel value at the same pixel point position to obtain the average channel pixel value at each pixel point position; combining the average channel pixel values at each pixel point position to obtain the mean fusion image of the image frame packet.

[0098] Exemplarily, when performing the pixel value averaging process, since the present application processes color images, the averaging process of only the pixel values can be performed channel by channel. When obtaining the channel pixel value data of each image frame at each pixel point within the image frame packet, at a pixel point position, by performing an averaging process on each channel pixel value respectively, the average channel pixel values of the three channels, namely the R channel, G channel, and B channel, at each pixel point position can be obtained; combining the average channel pixel values at each pixel point position, that is, combining the pixel points obtained by fusion, the mean fusion image of the image frame packet can be obtained. In one embodiment, for the group of image frame packets containing images, through the pixel value averaging process on the image frame packets, images can be fused into one image, and the pixel values at the corresponding positions of the newly generated sequential image in RGB format are the arithmetic averages of the channel pixel values of the original RGB images. Let the pixel value at the th row and th column of the sequential image be at the th row and th column of each image in the original image frame packet be

[0099]

[0100] Then the specific calculation method of these pixel values satisfies the following formula:

[0099]

[0100] After performing the pixel value fusion process on the images within the image frame packet through the above formula, the mean fusion image of the image frame packet can be obtained. In this embodiment, by performing an averaging process on each channel pixel value at the same pixel point position respectively, the features of the color image can be effectively averaged and combined, realizing the fusion process of the images within the image frame packet, obtaining the mean fusion image, thereby improving the accuracy of self-attention feature extraction.

[0101] In an exemplary embodiment, the video data processing result includes the video data classification result, and step 209 includes: performing a fully connected process on the video fusion feature to obtain a fully connected feature; classifying the target video based on the fully connected feature to obtain the video data classification result.

[0102] Exemplarily, the solution of the present application can be specifically applied to the field of video classification. For example, for the review of video content, after a user submits a video, the video data is processed and classified to determine whether the video is a violation video. Therefore, after feature fusion processing, a fully connected processing can be performed on the video fusion features to obtain fully connected features, and then the classification processing of the target video can be performed based on the features obtained by the fully connected operation to obtain the final processing result. In one embodiment, the complete process of feature fusion and fully connected processing can be referred to Figure 9 as shown. By performing feature fusion on the three-dimensional convolutional feature 3Dfeature and the self-attention feature self attention feature, specifically, feature fusion is performed here through a bitwise add operation, and then fully connected through a fully connected layer FC. Then, the result of the fully connected operation is mapped to the probability interval by a softmax function to obtain the final classification probability prediction result. In other embodiments, the solution of the present application can also be applied to video summarization and other processing. At this time, after the fully connected processing is performed through the fully connected layer FC, the fully connected processing can output multiple results, and these results correspond to the final video summary content. In addition, the present application has general applicability and can select the scaling size, video sampling rate, and number of image frame packets that match the computing power according to the computing power of each processor. This enables the algorithm engineering to flexibly deploy different solutions according to different task requirements. In this embodiment, fully connected features are obtained through fully connected processing, and then image classification processing is performed based on the fully connected features, which can effectively apply the video data processing method of the present application in the image classification task, thereby ensuring the robustness during the video classification processing.

[0103] The present application also provides an application scenario to apply the video data processing method of the present application. The application process of the video data processing method in this scenario includes:

[0104] When a user needs to perform summarization processing on specified video content to extract information from a large number of video contents with the highest efficiency, the video data processing method of the present application can be used to complete the video summarization processing process. First, the user needs to collect the video data and its corresponding summary data in the historical data, and then, based on these data, construct various sample sets for model training according to the model requirements. The training processing of model networks such as the three-dimensional convolutional network model, self-attention model, feature fusion layer, and model fully connected layer is completed through the sample sets. After the model training is passed, the summarization processing of the video content can be achieved through the trained model.

[0105] After obtaining the target video for which an abstract is required, the target video can be processed by extracting image frames based on a preset frame capture rate to obtain video image frames, and the video image frames can be evenly grouped according to the playback order to determine image frame packets composed of video image frames with the same grouping. Then, each image frame in the image frame packet is processed by a three-dimensional convolutional network model to obtain an image frame packet convolutional feature containing convolutional kernel number information, the number of images in the image frame packet, and image frame feature information; the image frame packet convolutional feature is processed by dimensional mapping to compress the feature dimensions of the convolutional kernel number information and the number of images to obtain the three-dimensional convolutional feature of the target video. At the same time, it is also necessary to obtain the channel pixel value data of each image frame in the image frame packet at each pixel point; the average processing is respectively performed on each channel pixel value at the same pixel point position to obtain the average channel pixel value at each pixel point position; the average channel pixel values at each pixel point position are combined to obtain the mean fusion image of the image frame packet. And the mean fusion image is segmented into fusion sub-images; the fusion sub-images are expanded based on RGB to obtain sub-image feature vectors; the sub-image feature vectors are linearly embedded by a fully connected layer to obtain self-attention input features; the self-attention input features are processed by a self-attention model for self-attention processing and fully connected processing to obtain self-attention features that match the dimensions of the three-dimensional convolutional features. Finally, the three-dimensional convolutional features and the self-attention features are processed by feature fusion to obtain video fusion features; the video fusion features are processed by a fully connected layer to obtain the abstract result of the target video.

[0106] It should be understood that although the steps in the flowcharts involved in the above embodiments are sequentially shown according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least some of the steps or stages in other steps or other steps.

[0107] Based on the same inventive concept, the embodiments of the present application also provide a video data processing device for implementing the video data processing method involved above. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the video data processing device provided below can refer to the limitations on the video data processing method in the above text, and will not be repeated here.

[0108] In an exemplary embodiment, asFigure 10 As shown, a video data processing device is provided, including:

[0109] A frame extraction processing module 1001, configured to perform image frame extraction processing on a target video to obtain video image frames.

[0110] A frame grouping module 1003, configured to perform grouping processing on the video image frames in the playback order to determine image frame packets composed of video image frames with the same grouping.

[0111] A feature extraction module 1005, configured to perform feature extraction processing on the image frames in each image frame packet through a three-dimensional convolutional network model to obtain three-dimensional convolutional features of the target video, and perform feature extraction processing on the image frames in each image frame packet through a self-attention model to obtain self-attention features of the target video. The three-dimensional convolutional network model is trained through image data with a time dimension in historical data, and the self-attention model is trained through image data fused by pixel means.

[0112] A feature fusion module 1007, configured to perform feature fusion processing on the three-dimensional convolutional features and the self-attention features to obtain video fusion features.

[0113] A fully connected processing module 1009, configured to perform fully connected processing on the video fusion features to obtain a video data processing result.

[0114] In one embodiment, the frame extraction processing module 1001 is specifically configured to: perform image frame extraction processing on the target video based on a preset frame capture rate to obtain video image frames. The frame grouping module 1003 is specifically configured to: perform average grouping processing on the video image frames in the playback order to determine image frame packets composed of video image frames with the same grouping.

[0115] In one embodiment, the feature extraction module 1005 is specifically configured to: perform image feature extraction processing on each image frame in the image frame packet through a three-dimensional convolutional network model to obtain image frame packet convolutional features, where the image frame packet convolutional features include convolutional kernel number information, the number of images in the image frame packet, and image frame feature information; perform dimension mapping processing on the image frame packet convolutional features to compress the feature dimensions of the convolutional kernel number information and the number of images information to obtain three-dimensional convolutional features of the target video.

[0116] In one embodiment, the feature extraction module 1005 is further configured to: perform pixel value averaging on the image frames within each image frame packet to obtain a mean fusion image of the image frame packet; segment the mean fusion image into fusion sub-images; expand the fusion sub-images based on RGB to obtain sub-image feature vectors; perform linear embedding processing on the sub-image feature vectors through a fully connected layer to obtain self-attention input features; and perform self-attention processing and fully connected processing on the self-attention input features through a self-attention model to obtain self-attention features that match the three-dimensional convolution feature dimension.

[0117] In one embodiment, the feature extraction module 1005 is further configured to: obtain the channel pixel value data of each image frame in the image frame packet at each pixel point; perform average processing on each channel pixel value at the same pixel point position respectively to obtain the average channel pixel value at each pixel point position; and combine the average channel pixel values at each pixel point position to obtain a mean fusion image of the image frame packet.

[0118] In one embodiment, the video data processing result includes a video data classification result; the fully connected processing module is specifically configured to: perform fully connected processing on the video fusion features to obtain fully connected features; and perform classification processing on the target video based on the fully connected features to obtain a video data classification result.

[0119] Each module in the above video data processing device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0120] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 11As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data related to video data processing. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a video data processing method.

[0121] Those skilled in the art can understand that Figure 11 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0122] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:

[0123] Perform image frame extraction processing on the target video to obtain video image frames;

[0124] Perform grouping processing on the video image frames in the playing order to determine image frame packets composed of video image frames with the same grouping;

[0125] Perform feature extraction processing on the image frames in each image frame packet through a three-dimensional convolutional network model to obtain three-dimensional convolutional features of the target video, and perform feature extraction processing on the image frames in each image frame packet through a self-attention model to obtain self-attention features of the target video. The three-dimensional convolutional network model is trained through image data with a time dimension in historical data, and the self-attention model is trained through image data fused by pixel means;

[0126] Perform feature fusion processing on the three-dimensional convolutional features and the self-attention features to obtain video fusion features;

[0127] Perform full connection processing on the video fusion features to obtain video data processing results.

[0128] In one embodiment, when the processor executes the computer program, the following steps are further implemented: performing image frame extraction processing on the target video based on a preset frame capture rate to obtain video image frames; performing average grouping processing on the video image frames in the playback order to determine image frame packets composed of video image frames with the same grouping.

[0129] In one embodiment, when the processor executes the computer program, the following steps are further implemented: performing image feature extraction processing on each image frame in the image frame packet through a three-dimensional convolutional network model to obtain image frame packet convolutional features, where the image frame packet convolutional features include convolutional kernel quantity information, the quantity information of images in the image frame packet, and image frame feature information; performing dimension mapping processing on the image frame packet convolutional features to compress the feature dimensions of the convolutional kernel quantity information and the quantity information of images, so as to obtain three-dimensional convolutional features of the target video.

[0130] In one embodiment, when the processor executes the computer program, the following steps are further implemented: performing pixel value averaging processing on the image frames in each image frame packet to obtain a mean fusion image of the image frame packet; segmenting the mean fusion image into fusion sub-images; expanding the fusion sub-images based on RGB to obtain sub-image feature vectors; performing linear embedding processing on the sub-image feature vectors through a fully connected layer to obtain self-attention input features; performing self-attention processing and fully connected processing on the self-attention input features through a self-attention model to obtain self-attention features that match the dimensions of the three-dimensional convolutional features.

[0131] In one embodiment, when the processor executes the computer program, the following steps are further implemented: obtaining channel pixel value data of each image frame in the image frame packet at each pixel point; respectively performing average processing on each channel pixel value at the same pixel point position to obtain the average channel pixel value at each pixel point position; combining the average channel pixel values at each pixel point position to obtain a mean fusion image of the image frame packet.

[0132] In one embodiment, when the processor executes the computer program, the following steps are further implemented: performing fully connected processing on the video fusion features to obtain fully connected features; performing classification processing on the target video based on the fully connected features to obtain a video data classification result.

[0133] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0134] Performing image frame extraction processing on the target video to obtain video image frames;

[0135] Performing grouping processing on the video image frames in the playback order to determine image frame packets composed of video image frames with the same grouping;

[0136] Feature extraction processing is performed on the image frames within each image frame packet through a three-dimensional convolutional network model to obtain the three-dimensional convolutional features of the target video, and feature extraction processing is performed on the image frames within each image frame packet through a self-attention model to obtain the self-attention features of the target video. The three-dimensional convolutional network model is trained through image data with a time dimension in historical data, and the self-attention model is trained through image data fused by pixel means;

[0137] Feature fusion processing is performed on the three-dimensional convolutional features and the self-attention features to obtain video fusion features;

[0138] Fully connected processing is performed on the video fusion features to obtain the video data processing result.

[0139] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: image frame extraction processing is performed on the target video based on a preset frame capture rate to obtain video image frames; the video image frames are evenly grouped according to the playback order, and it is determined that the image frame packets composed of video image frames with the same grouping.

[0140] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: image feature extraction processing is performed on each image frame within the image frame packet through a three-dimensional convolutional network model to obtain the convolutional features of the image frame packet. The convolutional features of the image frame packet include convolutional kernel number information, the number of images within the image frame packet, and image frame feature information; through dimension mapping processing on the convolutional features of the image frame packet, the feature dimensions of the convolutional kernel number information and the number of images are compressed to obtain the three-dimensional convolutional features of the target video.

[0141] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: pixel value averaging processing is performed on the image frames within each image frame packet to obtain the mean fusion image of the image frame packet; the mean fusion image is segmented into fusion sub-images; the fusion sub-images are expanded based on RGB to obtain sub-image feature vectors; linear embedding processing is performed on the sub-image feature vectors through a fully connected layer to obtain self-attention input features; self-attention processing and fully connected processing are performed on the self-attention input features through a self-attention model to obtain self-attention features that match the dimensions of the three-dimensional convolutional features.

[0142] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0143] To the channel pixel value data of each pixel point of each image frame within the image frame packet; average processing is respectively performed on each channel pixel value at the same pixel point position to obtain the average channel pixel value of each pixel point position; the average channel pixel values of each pixel point position are combined to obtain the mean fusion image of the image frame packet.

[0144] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: performing a fully-connected processing on the video fusion features to obtain fully-connected features; and classifying the target video based on the fully-connected features to obtain a video data classification result.

[0145] In one embodiment, a computer program product is provided, including a computer program which, when executed by a processor, implements the following steps:

[0146] Performing an image frame extraction process on the target video to obtain video image frames;

[0147] Grouping the video image frames in the playing order to determine image frame packets formed by video image frames with the same grouping;

[0148] Performing a feature extraction process on the image frames in each image frame packet through a three-dimensional convolutional network model to obtain three-dimensional convolutional features of the target video, and performing a feature extraction process on the image frames in each image frame packet through a self-attention model to obtain self-attention features of the target video. The three-dimensional convolutional network model is trained by using image data with a time dimension in historical data, and the self-attention model is trained by using image data with pixel mean fusion; performing a feature fusion process on the three-dimensional convolutional features and the self-attention features to obtain video fusion features;

[0149] Performing a fully-connected processing on the video fusion features to obtain a video data processing result.

[0150] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: performing an image frame extraction process on the target video based on a preset frame capture rate to obtain video image frames; and performing an average grouping process on the video image frames in the playing order to determine image frame packets formed by video image frames with the same grouping.

[0151] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: performing an image feature extraction process on each image frame in the image frame packet through a three-dimensional convolutional network model to obtain image frame packet convolutional features, where the image frame packet convolutional features include convolutional kernel number information, the number of images in the image frame packet, and image frame feature information; and performing a dimension mapping process on the image frame packet convolutional features to compress the feature dimensions of the convolutional kernel number information and the number of images information to obtain three-dimensional convolutional features of the target video.

[0152] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: performing pixel value averaging on the image frames within each image frame packet to obtain a mean fusion image of the image frame packet; segmenting the mean fusion image into fusion sub-images; expanding the fusion sub-images based on RGB to obtain sub-image feature vectors; performing linear embedding processing on the sub-image feature vectors through a fully connected layer to obtain self-attention input features; and performing self-attention processing and fully connected processing on the self-attention input features through a self-attention model to obtain self-attention features that match the three-dimensional convolution feature dimension.

[0153] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0154] obtaining the channel pixel value data of each image frame in the image frame packet at each pixel point; respectively performing averaging processing on each channel pixel value at the same pixel point position to obtain the average channel pixel value at each pixel point position; and combining the average channel pixel values at each pixel point position to obtain the mean fusion image of the image frame packet.

[0155] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: performing fully connected processing on the video fusion features to obtain fully connected features; and performing classification processing on the target video based on the fully connected features to obtain a video data classification result.

[0156] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0157] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0158] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0159] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A video data processing method, characterized in that, The method includes: Performing image frame extraction processing on the target video to obtain video image frames; Grouping the video image frames in the playing order and determining image frame packets composed of the video image frames with the same grouping; Performing feature extraction processing on the image frames in each image frame packet through a three-dimensional convolutional network model to obtain three-dimensional convolutional features of the target video, and performing feature extraction processing on the image frames in each image frame packet through a self-attention model to obtain self-attention features of the target video. The three-dimensional convolutional network model is trained through image data with a time dimension in historical data, and the self-attention model is trained through image data fused by pixel means; Performing feature fusion processing on the three-dimensional convolutional features and the self-attention features to obtain video fusion features; Performing a fully connected process on the video fusion features to obtain a video data processing result.

2. The method according to claim 1, characterized in that, The performing image frame extraction processing on the target video to obtain video image frames includes: Performing image frame extraction processing on the target video based on a preset frame capture rate to obtain video image frames; The grouping the video image frames in the playing order and determining image frame packets composed of the video image frames with the same grouping includes: Performing average grouping processing on the video image frames in the playing order and determining image frame packets composed of the video image frames with the same grouping.

3. The method according to claim 1, characterized in that, The performing feature extraction processing on the image frames in each image frame packet through a three-dimensional convolutional network model to obtain three-dimensional convolutional features of the target video includes: Performing image feature extraction processing on each image frame in the image frame packet through a three-dimensional convolutional network model to obtain image frame packet convolutional features, where the image frame packet convolutional features include convolutional kernel number information, the number of images in the image frame packet information, and image frame feature information; Performing dimension mapping processing on the image frame packet convolutional features to compress the feature dimensions of the convolutional kernel number information and the image number information to obtain three-dimensional convolutional features of the target video.

4. The method according to claim 1, characterized in that, The performing feature extraction processing on the image frames in each image frame packet through a self-attention model to obtain self-attention features of the target video includes: Performing pixel value averaging processing on the image frames in each image frame packet to obtain a mean fusion image of the image frame packet; Segmenting the mean fusion image into fusion sub-images; Expanding the fusion sub-images based on RGB to obtain sub-image feature vectors; Performing linear embedding processing on the sub-image feature vectors through a fully connected layer to obtain self-attention input features; Performing self-attention processing and fully connected processing on the self-attention input features through a self-attention model to obtain self-attention features that match the dimension of the three-dimensional convolutional features.

5. The method according to claim 4, wherein The performing pixel value averaging processing on the image frames in each image frame packet to obtain a mean fusion image of the image frame packet includes: Obtaining channel pixel value data of each image frame in the image frame packet at each pixel point; Performing average processing on each channel pixel value at the same pixel point position respectively to obtain average channel pixel values at each pixel point position. Combine the average channel pixel values at each pixel position to obtain the mean fusion image of the image frame packet.

6. The method according to any one of claims 1 to 5, characterized in that The video data processing result includes the video data classification result; The full connection processing of the video fusion feature to obtain the video data processing result includes: Perform full connection processing on the video fusion feature to obtain the full connection feature; Based on the full connection feature, perform classification processing on the target video to obtain the video data classification result.

7. A video data processing device, characterized in that, The device includes: A frame extraction processing module for performing image frame extraction processing on the target video to obtain video image frames; A frame grouping module for grouping the video image frames in the playing order to determine the image frame packets composed of the video image frames with the same grouping; A feature extraction module for performing feature extraction processing on the image frames in each image frame packet through a three-dimensional convolutional network model to obtain the three-dimensional convolutional features of the target video, and performing feature extraction processing on the image frames in each image frame packet through a self-attention model to obtain the self-attention features of the target video. The three-dimensional convolutional network model is trained by the image data with a time dimension in the historical data, and the self-attention model is trained by the image data of pixel mean fusion; A feature fusion module for performing feature fusion processing on the three-dimensional convolutional features and the self-attention features to obtain video fusion features; A full connection processing module for performing full connection processing on the video fusion features to obtain video data processing results.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.