Data Processing Method, Device, Storage Medium and Equipment

By extracting and fusion of video frames, generating time and pixel mixed feature maps, combining the classification model to realize automatic and accurate classification of video data, solving the accuracy problem of label-free video data classification.

CN114332678BActive Publication Date: 2025-07-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111480074.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-06
Publication Date
2025-07-11
Estimated Expiration
2041-12-06

AI Technical Summary

Technical Problem

The lack of effective methods in video data classification in the prior art leads to the inability to classify without video tags and introductions, and the accuracy is low when relying on manual classification.

Method used

By extracting the video frames, generating a feature map sequence, and using time and pixel mixing to generate a fusion feature map, combining the classification model to automatically classify the video content categories, avoiding relying on text information and manual experience.

Benefits of technology

It realizes accurate classification of video data without video tags and introductions, improves the accuracy and efficiency of classification, reduces the amount of calculation and avoids the loss of critical information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114332678B_ABST
    Figure CN114332678B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a data processing method, apparatus, storage medium, and device. The present application can be applied to the technical fields of artificial intelligence and intelligent transportation. The method includes: performing a first feature extraction process on M video frames in target video data to obtain a first feature map sequence, and performing a second feature extraction process on the M video frames to obtain a second feature map sequence; sampling the first feature map sequence according to a first time sampling parameter to obtain a target first feature map, and sampling the second feature map sequence according to a second time sampling parameter to obtain a target second feature map; generating a temporal fusion feature map according to the target first feature map and the target second feature map; generating a target fusion feature map according to the temporal fusion feature map, the first feature map sequence, and the second feature map sequence, and determining the video content category of the target video data according to the target fusion feature map. Through the present application, the accuracy of classifying the target video data can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a data processing method, apparatus, storage medium, and device. Background Art

[0002] With the development of artificial intelligence technology, more and more application scenarios call on classification models to classify video data to determine the category corresponding to the video. Video classification refers to classifying the content contained in a given video clip. In the fields of security, social media, intelligent transportation, etc., video classification has broad application prospects.

[0003] Currently, when classifying video data, generally, text information such as video tags and video introductions of the video data is obtained, and the type of the video is determined by identifying the text information. When the video data does not have video tags and video introductions, the video data cannot be classified, and even manual classification is required. Affected by subjective and other factors and with limited manual classification experience, the accuracy of video classification is relatively low. Summary of the Invention

[0004] The technical problem to be solved by the embodiments of this application is to provide a data processing method, apparatus, storage medium, and device, which can improve the accuracy of classifying target video data.

[0005] On the one hand, the embodiments of this application provide a data processing method, including:

[0006] Obtain M video frames in the target video data, perform first feature extraction processing on the M video frames to obtain a first feature map sequence, and perform second feature extraction processing on the M video frames to obtain a second feature map sequence;

[0007] Sample the first feature map sequence according to the first time sampling parameter to obtain a target first feature map, and sample the second feature map sequence according to the second time sampling parameter to obtain a target second feature map; the sum of the number of video frames corresponding to the target first feature map and the target second feature map is equal to M, and the video frames corresponding to the target first feature map and the target second feature map are different from each other;

[0008] Generate a time fusion feature map according to the target first feature map and the target second feature map;

[0009] Generate a target fusion feature map according to the time fusion feature map, the first feature map sequence, and the second feature map sequence, and classify the target fusion feature map to obtain the video content category of the target video data.

[0010] Among them, the first feature map sequence includes first feature maps corresponding to M video frames respectively, and the second feature map sequence includes second feature maps corresponding to M video frames respectively;

[0011] Generate a target fusion feature map according to the temporal fusion feature map, the first feature map sequence, and the second feature map sequence, and classify the target fusion feature map to obtain the video content category of the target video data, including:

[0012] In the first feature map sequence and the second feature map sequence, pixel-blend and splice the first feature map and the second feature map associated with the same video frame to obtain pixel-blend feature maps corresponding to M video frames respectively;

[0013] Generate a pixel fusion feature map according to the pixel-blend feature maps corresponding to M video frames respectively;

[0014] Generate a target fusion feature map according to the temporal fusion feature map and the pixel fusion feature map, and classify the target fusion feature map to obtain the video content category of the target video data.

[0015] Among them, obtaining M video frames in the target video data includes:

[0016] Obtain the original video data, and obtain the content attributes of each original video frame in the original video data;

[0017] Divide the original video data according to the content attributes of each original video frame to obtain N video segments; N is a positive integer;

[0018] Select a target video segment from the N video segments as the target video data;

[0019] Perform video frame sampling on the original video frames included in the target video data according to the number of sampled video frames M indicated by the video sampling rule to obtain M video frames in the target video data.

[0020] Among them, the method further includes:

[0021] Obtain an initial time sampling parameter; the number of sampling elements in the initial time sampling parameter is M;

[0022] Randomly determine the element values of M sampling elements with a positional order in the initial time sampling parameter to obtain a first time sampling parameter; the element values include a first element threshold for indicating sampling of a feature map and a second element threshold for indicating masking of a feature map;

[0023] Determine the second time sampling parameter according to the element values of the M sampling elements in the first time sampling parameter; the element values of the sampling elements in the same position in the first time sampling parameter and the second time sampling parameter are different from each other.

[0024] Among them, sampling the first feature map sequence according to the first time sampling parameter to obtain the target first feature map, and sampling the second feature map sequence according to the second time sampling parameter to obtain the target second feature map, including:

[0025] Call the target classification model. In the feature fusion layer of the target classification model, sample the associated feature maps in the first feature map sequence based on the first element threshold in the first time sampling parameter, and mask the associated feature maps in the first feature map sequence based on the second element threshold in the first time sampling parameter to obtain the target first feature map;

[0026] Sample the associated feature maps in the second feature map sequence according to the first element threshold in the second time sampling parameter, and mask the associated feature maps in the second feature map sequence according to the second element threshold in the second time sampling parameter to obtain the target second feature map.

[0027] Among them, generating the time fusion feature map according to the target first feature map and the target second feature map, including:

[0028] Obtain the first timestamp of the video frame corresponding to the target first feature map, and obtain the second timestamp of the video frame corresponding to the target second feature map;

[0029] Combine the target first feature map and the target second feature map according to the time order between the first timestamp and the second timestamp to obtain the time fusion feature map.

[0030] Among them, the M video frames include video frame M i , where i is a positive integer less than or equal to M;

[0031] In the first feature map sequence and the second feature map sequence, pixel-mix and splice the first feature map and the second feature map associated with the same video frame to obtain the pixel-mix feature maps corresponding to the M video frames, including:

[0032] Call the target classification model, and obtain the first feature map corresponding to video frame M in the first feature map sequence through the feature fusion layer in the target classification model, and obtain the second feature map corresponding to video frame M in the second feature map sequence i ; i corresponding second feature map;

[0033] According to the first pixel sampling parameter, for video frame M iPerform pixel sampling on the corresponding first feature map to obtain a first pixel-sampled feature map, and perform pixel sampling on the video frame M according to the second pixel-sampling parameter i Perform pixel sampling on the corresponding second feature map to obtain a second pixel-sampled feature map;

[0034] Perform pixel mixing and stitching on the first pixel-sampled feature map and the second pixel-sampled feature map to obtain a pixel-mixed feature map corresponding to the video frame M i

[0035] Among them, generating a target fusion feature map according to the temporal fusion feature map and the pixel fusion feature map, and classifying the target fusion feature map to obtain the video content category of the target video data, including:

[0036] Call the target classification model, and add the temporal fusion feature map and the pixel fusion feature map through the feature fusion layer in the target classification model to obtain the target fusion feature map;

[0037] Perform convolution processing on the target fusion feature map through the convolution layer in the target classification model to obtain the target fusion feature map after convolution processing;

[0038] Perform classification processing on the target fusion feature map after convolution processing through the classification layer in the target classification model to obtain the video content category of the target video data.

[0039] Among them, performing classification processing on the target fusion feature map after convolution processing through the classification layer in the target classification model to obtain the video content category of the target video data, including:

[0040] Input the target fusion feature map after convolution processing into the first classification sub-layer in the classification layer, and perform classification processing on the target fusion feature map after convolution processing to obtain a first classification result;

[0041] Input the target fusion feature map after convolution processing into the second classification sub-layer in the classification layer, and perform classification processing on the target fusion feature map after convolution processing to obtain a second classification result;

[0042] Obtain the average value of the first classification result and the second classification result, and determine the video content category of the target video data according to this average value.

[0043] One aspect of the embodiments of the present application provides a data processing method, including:

[0044] Through an initial classification model, perform first feature extraction processing on M first sample video frames in the first sample video data to obtain a first sample feature map sequence, and perform second feature extraction processing on M second sample video frames in the second sample video data to obtain a second sample feature map sequence; M is a positive integer; ​

[0045] Sample the first sample feature map sequence according to the first sample time sampling parameter to obtain a target first sample feature map, and sample the second sample feature map sequence according to the second sample time sampling parameter to obtain a target second sample feature map; the sum of the number of sample video frames corresponding to the target first sample feature map and the target second sample feature map is equal to M, and the position of the target first sample feature map in the first sample feature map sequence is different from the position of the target second sample feature map in the second sample feature map sequence;

[0046] Generate a sample time fusion feature map according to the target first sample feature map and the target second sample feature map, generate a target sample fusion feature map for predicting the video content category according to the sample time fusion feature map, the first sample feature map sequence, and the second sample feature map sequence, and adjust the parameters of the initial classification model according to the target sample fusion feature map to obtain a target classification model; the target classification model is used to predict the video content category of the target video data.

[0047] Among them, adjusting the parameters of the initial classification model according to the target sample fusion feature map to obtain a target classification model includes:

[0048] Predict the first predicted video content category of the first sample video data according to the target sample fusion feature map, and predict the second predicted video content category of the second sample video data according to the target sample fusion feature map;

[0049] Generate a first loss function according to the first video content category label and the first predicted video content category of the first sample video data;

[0050] Generate a second loss function according to the second video content category label and the second predicted video content category of the second sample video data;

[0051] Generate a total loss function according to the first loss function and the second loss function, and adjust the parameters of the initial classification model according to the total loss function. When the initial classification model after parameter adjustment meets the training convergence condition, determine the initial classification model after parameter adjustment as the target classification model.

[0052] Among them, generating a total loss function according to the first loss function and the second loss function includes:

[0053] Perform pixel sampling on the first sample feature map sequence according to the first sample pixel sampling parameter to obtain a first sample pixel sampling feature map sequence, call an information loss prediction model, and perform loss prediction on the first sample pixel sampling feature map sequence and the target first sample feature map to obtain a first information loss probability corresponding to the first sample feature map sequence;

[0054] Perform pixel sampling on the second sample feature map sequence according to the second sample pixel sampling parameters to obtain a second sample pixel sampling feature map sequence, and call an information loss prediction model to perform loss prediction on the second sample pixel sampling feature map sequence and the target second sample feature map to obtain a second information loss probability corresponding to the second sample feature map sequence;

[0055] Perform weighted processing on the first loss function according to the first information loss probability to obtain a weighted first loss function, and perform weighted processing on the second loss function according to the second information loss probability to obtain a weighted second loss function;

[0056] Perform a summation process on the weighted first loss function and the weighted second loss function to obtain a total loss function.

[0057] One aspect of the embodiments of the present application provides a data processing device, including:

[0058] A first feature extraction module, configured to obtain M video frames in the target video data, perform first feature extraction processing on the M video frames to obtain a first feature map sequence, and perform second feature extraction processing on the M video frames to obtain a second feature map sequence;

[0059] A first sampling module, configured to sample the first feature map sequence according to the first time sampling parameters to obtain a target first feature map, and sample the second feature map sequence according to the second time sampling parameters to obtain a target second feature map; the sum of the number of video frames corresponding to the target first feature map and the target second feature map is equal to M, and the video frames corresponding to the target first feature map and the target second feature map are different from each other;

[0060] A generation module, configured to generate a time fusion feature map according to the target first feature map and the target second feature map;

[0061] A classification module, configured to generate a target fusion feature map according to the time fusion feature map, the first feature map sequence, and the second feature map sequence, and classify the target fusion feature map to obtain the video content category of the target video data.

[0062] Wherein, the first feature map sequence includes first feature maps corresponding to the M video frames respectively, and the second feature map sequence includes second feature maps corresponding to the M video frames respectively;

[0063] The classification module includes:

[0064] A pixel mixing and splicing unit, configured to perform pixel mixing and splicing on the first feature map and the second feature map associated with the same video frame in the first feature map sequence and the second feature map sequence to obtain pixel mixing feature maps corresponding to the M video frames respectively;

[0065] The first generation unit is configured to generate a pixel fusion feature map based on the pixel mixed feature maps respectively corresponding to M video frames;

[0066] The classification unit is configured to generate a target fusion feature map based on the temporal fusion feature map and the pixel fusion feature map, classify the target fusion feature map, and obtain the video content category of the target video data.

[0067] Among them, the first feature extraction module includes:

[0068] The first acquisition unit is configured to acquire the original video data and acquire the content attributes of each original video frame in the original video data;

[0069] The partitioning unit is configured to partition the original video data according to the content attributes of each original video frame to obtain N video segments; N is a positive integer;

[0070] The selection unit is configured to select a target video segment from the N video segments as the target video data;

[0071] The video frame sampling unit is configured to perform video frame sampling on the original video frames included in the target video data according to the number M of sampled video frames indicated by the video sampling rule to obtain M video frames in the target video data.

[0072] Among them, the data processing device further includes:

[0073] The acquisition module is configured to acquire an initial time sampling parameter; the number of sampling elements in the initial time sampling parameter is M;

[0074] The first determination module is configured to randomly determine the element values of M sampling elements with a position order in the initial time sampling parameter to obtain a first time sampling parameter; the element values include a first element threshold for indicating sampling of a feature map and a second element threshold for indicating masking of a feature map;

[0075] The second determination module is configured to determine a second time sampling parameter according to the element values of the M sampling elements in the first time sampling parameter; the element values of the sampling elements in the same position in the first time sampling parameter and the second time sampling parameter are different from each other.

[0076] Among them, the first sampling module includes:

[0077] The first sampling unit is configured to call a target classification model, in the feature fusion layer of the target classification model, sample the associated feature maps in the first feature map sequence based on the first element threshold in the first time sampling parameter, and mask the associated feature maps in the first feature map sequence based on the second element threshold in the first time sampling parameter to obtain a target first feature map;

[0078] A second sampling unit, configured to sample the associated feature maps in the second feature map sequence according to the first element threshold in the second time sampling parameter, and mask the associated feature maps in the second feature map sequence according to the second element threshold in the second time sampling parameter, to obtain a target second feature map.

[0079] Among them, the generation module includes:

[0080] A second obtaining unit, configured to obtain the first timestamp of the video frame corresponding to the target first feature map, and obtain the second timestamp of the video frame corresponding to the target second feature map;

[0081] A combining unit, configured to combine the target first feature map and the target second feature map according to the time sequence between the first timestamp and the second timestamp, to obtain a time fusion feature map.

[0082] Among them, the M video frames include video frame M i , where i is a positive integer less than or equal to M;

[0083] Specifically, the pixel mixing and stitching unit is configured to:

[0084] Invoke the target classification model, and through the feature fusion layer in the target classification model, obtain the first feature map corresponding to video frame M in the first feature map sequence, and obtain the second feature map corresponding to video frame M in the second feature map sequence; i i Corresponding second feature map;

[0085] According to the first pixel sampling parameter, perform pixel sampling on the first feature map corresponding to video frame M, to obtain a first pixel sampling feature map, and according to the second pixel sampling parameter, perform pixel sampling on the second feature map corresponding to video frame M, to obtain a second pixel sampling feature map; i i Corresponding second pixel sampling feature map;

[0086] Perform pixel mixing and stitching on the first pixel sampling feature map and the second pixel sampling feature map, to obtain the pixel mixing feature map corresponding to video frame M i

[0087] Among them, the classification unit is specifically configured to:

[0088] Invoke the target classification model, and through the feature fusion layer in the target classification model, add the time fusion feature map and the pixel fusion feature map, to obtain a target fusion feature map;

[0089] Through the convolutional layer in the target classification model, perform convolutional processing on the target fusion feature map, to obtain a target fusion feature map after convolutional processing;

[0090] ​​​Through the classification layer in the target classification model, classify the target fusion feature map after convolution processing to obtain the video content category of the target video data.

[0091] Among them, the classification unit is specifically further configured to include:

[0092] Input the target fusion feature map after convolution processing into the first classification sub-layer in the classification layer, classify the target fusion feature map after convolution processing to obtain a first classification result;

[0093] Input the target fusion feature map after convolution processing into the second classification sub-layer in the classification layer, classify the target fusion feature map after convolution processing to obtain a second classification result;

[0094] Obtain the average value of the first classification result and the second classification result, and determine the video content category of the target video data according to this average value.

[0095] One aspect of the embodiments of the present application provides a data processing device, including:

[0096] A second feature extraction module, configured to perform first feature extraction processing on M first sample video frames in the first sample video data through an initial classification model to obtain a first sample feature map sequence, and perform second feature extraction processing on M second sample video frames in the second sample video data to obtain a second sample feature map sequence; M is a positive integer;

[0097] A second sampling module, configured to sample the first sample feature map sequence according to the first sample time sampling parameter to obtain a target first sample feature map, and sample the second sample feature map sequence according to the second sample time sampling parameter to obtain a target second sample feature map; the sum of the number of sample video frames corresponding to the target first sample feature map and the target second sample feature map is equal to M, and the position of the target first sample feature map in the first sample feature map sequence is different from the position of the target second sample feature map in the second sample feature map sequence;

[0098] A parameter adjustment module, configured to generate a sample time fusion feature map according to the target first sample feature map and the target second sample feature map, generate a target sample fusion feature map for predicting the video content category according to the sample time fusion feature map, the first sample feature map sequence, and the second sample feature map sequence, and adjust the parameters of the initial classification model according to the target sample fusion feature map to obtain a target classification model; the target classification model is used to predict the video content category of the target video data.

[0099] The parameter adjustment module includes:

[0100] A prediction unit, configured to predict a first predicted video content category of first sample video data according to a target sample fusion feature map, and predict a second predicted video content category of second sample video data according to the target sample fusion feature map;

[0101] A second generation unit, configured to generate a first loss function according to a first video content category label and a first predicted video content category of the first sample video data;

[0102] A third generation unit, configured to generate a second loss function according to a second video content category label and a second predicted video content category of the second sample video data;

[0103] A determination unit, configured to generate an overall loss function according to the first loss function and the second loss function, adjust parameters of an initial classification model according to the overall loss function, and determine the initial classification model with adjusted parameters as a target classification model when the initial classification model with adjusted parameters meets a training convergence condition.

[0104] Wherein, the determination unit is specifically configured to:

[0105] Perform pixel sampling on a first sample feature map sequence according to first sample pixel sampling parameters to obtain a first sample pixel sampling feature map sequence, call an information loss prediction model, perform loss prediction on the first sample pixel sampling feature map sequence and a target first sample feature map, and obtain a first information loss probability corresponding to the first sample feature map sequence;

[0106] Perform pixel sampling on a second sample feature map sequence according to second sample pixel sampling parameters to obtain a second sample pixel sampling feature map sequence, call an information loss prediction model, perform loss prediction on the second sample pixel sampling feature map sequence and a target second sample feature map, and obtain a second information loss probability corresponding to the second sample feature map sequence;

[0107] Perform weighted processing on the first loss function according to the first information loss probability to obtain a weighted first loss function, and perform weighted processing on the second loss function according to the second information loss probability to obtain a weighted second loss function;

[0108] Perform summation processing on the weighted first loss function and the weighted second loss function to obtain an overall loss function.

[0109] An embodiment of the present application provides a computer device on the one hand, including: a processor and a memory;

[0110] The processor is connected to the memory, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, the computer device executes the method provided by the embodiment of the present application.

[0111] One aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, which is adapted to be loaded and executed by a processor so that a computer device having the processor executes the method provided by the embodiments of the present application.

[0112] One aspect of the embodiments of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions so that the computer device executes the method provided by the embodiments of the present application.

[0113] In the embodiments of the present application, by acquiring M video frames in target video data, performing first feature extraction processing on the M video frames to obtain a first feature map sequence, and performing second feature extraction processing on the M video frames to obtain a second feature map sequence. By performing first feature extraction and second feature extraction on the M video frames, a first feature map sequence and a second feature map sequence of the M video frames can be obtained, and different feature information of the M video frames can be extracted from different perspectives. Further, the first feature map sequence is sampled according to a first time sampling parameter to obtain a target first feature map, and the second feature map sequence is sampled according to a second time sampling parameter to obtain a target second feature map. The sum of the number of video frames corresponding to the target first feature map and the target second feature map is equal to M, and the video frames corresponding to the target first feature map and the target second feature map are different from each other, and a time fusion feature map is generated according to the target first feature map and the target second feature map. It can be seen that the first feature map sequence and the second feature map sequence are sampled respectively from the time dimension to obtain a time fusion feature map, so as to perform feature enhancement on the target video data according to the timing information between each video frame and improve the feature enhancement effect of the target video data. Further, a target fusion feature map is generated according to the time fusion feature map, the first feature map sequence, and the second feature map sequence, and the target fusion feature map is classified to obtain the video content category of the target video data. It can be seen that the present application does not need to rely on text information such as video labels and video introductions of target video data, nor does it need to rely on manual experience analysis. By classifying the target fusion feature map, accurate classification of the target video data can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0114] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0115] Figure 1 It is a schematic structural diagram of a data processing system provided by an embodiment of the present application;

[0116] Figure 2 It is a schematic diagram of an application scenario of data processing provided by an embodiment of the present application;

[0117] Figure 3 It is a schematic flowchart of a data processing method provided by an embodiment of the present application;

[0118] Figure 4 It is a schematic diagram of predicting the category of video content by using a target classification model provided by an embodiment of the present application;

[0119] Figure 5 It is a schematic flowchart of a data processing method provided by an embodiment of the present application;

[0120] Figure 6 It is a schematic diagram of a method for obtaining a target fusion feature map provided by an embodiment of the present application;

[0121] Figure 7 It is a schematic flowchart of a data processing method provided by an embodiment of the present application;

[0122] Figure 8 It is a schematic diagram of obtaining the probability of information loss provided by an embodiment of the present application;

[0123] Figure 9 It is a schematic diagram of a method for training an initial classification model provided by an embodiment of the present application;

[0124] Figure 10 It is a schematic structural diagram of a data processing device provided by an embodiment of the present application;

[0125] Figure 11 It is a schematic structural diagram of a data processing device provided by an embodiment of the present application;

[0126] Figure 12 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0127] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0128] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.

[0129] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. In this application, machine learning technology can be used to perform feature fusion on the feature maps corresponding to the first sample video data and the second sample video data through an initial classification model to obtain a target sample fusion feature map. According to this target sample fusion feature map, the parameters of the initial classification model are adjusted to obtain a target classification model, which is used to predict the video content category of the target video data. In this way, by training the initial classification model, input feature enhancement and model integration can be carried out while training the initial classification model, improving the classification accuracy and robustness of the trained target classification model, that is, the generalization ability of the target classification model can be improved, and the video content categories of different video data can be accurately predicted.

[0130] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, vehicle networking, autonomous driving, intelligent transportation, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role. The solution provided in the embodiments of this application relates to the intelligent transportation technology of artificial intelligence, which will be specifically described through the following embodiments: This solution can use the camera component in the vehicle terminal to capture the road conditions around the vehicle to obtain target video data. The vehicle terminal sends the target video data to the server, and the target classification model in the server classifies the vehicle trajectories of the vehicles in the target video data to obtain the driving trajectories of the vehicles in the target video data (such as lane change, left turn, etc.), and then outputs the driving trajectories of the vehicles in the target video data through the vehicle terminal or the user terminal. In this way, the driver can make driving predictions based on the driving trajectories of the vehicles in the target video data, which provides convenience for the driver's driving.

[0131] Please refer to Figure 1 , Figure 1 which is a schematic structural diagram of a data processing system provided by an embodiment of this application. As Figure 1 shown, the data processing system may include a server 10 and a user terminal cluster. The user terminal cluster may include one or more user terminals, and the number of user terminals will not be limited here. As Figure 1 shown, it may specifically include user terminal 100a, user terminal 100b, user terminal 100c,..., user terminal 100n. As Figure 1 shown, user terminal 100a, user terminal 100b, user terminal 100c,..., user terminal 100n may be respectively connected to the above-mentioned server 10 through a network, so that each user terminal can perform data interaction with the server 10 through this network connection.

[0132] Among them, each user terminal in the user terminal cluster may include: intelligent terminals with data processing capabilities such as smart phones, tablets, laptops, desktop computers, wearable devices, smart homes, head-mounted devices, vehicle terminals, etc. It should be understood that each user terminal in the user terminal cluster as Figure 1 shown may be installed with a target application (i.e., application client). When the application client runs on each user terminal, it can perform data interaction with the above-mentioned Figure 1 shown server 10 respectively.

[0133] Among them, as Figure 1As shown, the server 10 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0134] For ease of understanding, in the embodiments of the present application, Figure 1 One user terminal can be selected from the multiple user terminals shown as the target user terminal. The target user terminal can include: intelligent terminals with data processing functions such as smart phones, tablet computers, laptop computers, desktop computers, and smart TVs. For example, for ease of understanding, in the embodiments of the present application, Figure 1 The user terminal 100a shown can be used as the target user terminal. The user terminal 100a can obtain M video frames in the target video data, where M is a positive integer. For example, M can take values of 1, 2, 3,.... The user terminal 100a can send the M video frames in the target video data to the server 10. The server 10 includes a target classification model for classifying video data. The server 10 can automatically classify the M video frames in the target video data uploaded by the user terminal 100a based on the target classification model, obtain the video content category of the target video data, and return the video content category of the target video data to the user terminal 100a, so as to quickly and accurately classify the target video data.

[0135] For ease of understanding, further, please refer to Figure 2 , Figure 2 which is a schematic diagram of an application scenario of data processing provided by the embodiments of the present application. Among them, as Figure 2 shown, the server 20e can be the above-mentioned server 10, and as Figure 2 shown, the target user terminal 20a can be any one of the user terminal clusters shown above Figure 1 , for example, the target user terminal 20a can be the above-mentioned user terminal 100a. As Figure 2As shown, the target user 20b can, in the video sharing interface 20c of the target user terminal 20a, perform an operation of clicking the upload button to upload and share the video data to be shared. The target user terminal 20a can respond to the upload operation of the target user 20b and obtain the video data uploaded by the target user 20b as the target video data. When the target user terminal 20a receives the target video data uploaded by the target user 20b, it can perform an interface jump, jump the video sharing interface 20c to the video sharing interface 20d, and display "Data uploading" to prompt the target user 20b that the video data uploaded by the target user 20b is currently being uploaded. During the process of the target user terminal 20a jumping from the video sharing interface 20c to the video sharing interface 20d, it can send the video data uploaded by the target user 20b as the target video data to the server 20e. The server 20e can obtain M video frames in the target video data, call the target classification model 20f to classify and process the M video frames in the target video data, and obtain the video content category of the target video data.

[0136] Furthermore, the server 20e can return the video content category of the target video data to the target user terminal 20a. The target user terminal 20a can output the video sharing interface 20g according to the video content category of the target video data. The video sharing interface 20g includes a prompt message 20h, and the prompt message 20h is used to prompt whether the target video data uploaded by the target user 20b is uploaded successfully. Among them, the target user terminal 20a can detect whether the video content category of the target video data is legal. If the target user terminal 20a detects that the video content category of the target video data is not legal, it outputs the video sharing interface 20g containing the prompt message 20h "The video content does not meet the requirements, please upload again". The prompt message 20h "The video content does not meet the requirements, please upload again" is used to prompt that the target video data uploaded by the target user 20b does not meet the requirements (i.e., is illegal), that is, the video data upload fails and needs to be uploaded again. If the target user terminal 20a detects that the video content category of the target video data is legal, it outputs the video sharing interface 20g containing the prompt message 20h "Upload successful".

[0137] For example, in an authentication scenario, the user needs to perform a target action according to the instructions. After the target user 20b uploads the target video data recorded according to the instructions, the server 20e can classify the target video data to obtain the user action (i.e., the video content category) in the target video data. When the user action in the target video data is the target action, it can be determined that the user action of the target user 20b is legal, and then the "upload successful" interface is output, that is, the authentication of the target user 20b is passed. If the user action of the target user 20b does not belong to the target action, it can be determined that the user action of the target user 20b is illegal, and then the "video action does not meet the requirements, please upload again" is output to prompt the user to perform the target action again according to the instructions for authentication. Through this application, the target video data can be data-augmented by the target classification model, and the video content category of the target video data can be accurately predicted, improving the accuracy of subsequent business processing based on the video content category of the target video data.

[0138] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of a data processing method provided by an embodiment of this application. This data processing method can be executed by a computer device, which can be a server (such as the server 10 above Figure 1 ), or a user terminal (such as any user terminal in the user terminal cluster above Figure 1 ), and this application does not make any restrictions. As Figure 3 shown, this data processing method may include but is not limited to the following steps:

[0139] S101, obtain M video frames in the target video data, perform a first feature extraction process on the M video frames to obtain a first feature map sequence, and perform a second feature extraction process on the M video frames to obtain a second feature map sequence.

[0140] Specifically, the computer device can perform data augmentation on the target video data in the time series dimension and the pixel dimension, which can improve the generalization of classifying the target video data and also improve the accuracy of classifying the target video data. Among them, the computer device can obtain M video frames in the target video data. The target video data can refer to the data obtained by the computer device through the camera component or the video data uploaded by the user. Among them, the M video frames in the target video data can refer to all the video frames in the target video data or some of the video frames in the target video data. M is a positive integer, and M can take values such as 1, 2, 3,.... Further, the computer device can perform a first feature extraction process on the M video frames to obtain a first feature map sequence, and perform a second feature extraction process on the M video frames to obtain a second feature map sequence.

[0141] Optionally, the specific manner in which the computer device obtains M video frames from the target video data may include: obtaining the original video data, and obtaining the content attributes of each original video frame in the original video data. According to the content attributes of each original video frame, the original video data is divided to obtain N video segments; N is a positive integer. The target video segment is selected from the N video segments as the target video data, and video frame sampling is performed on the original video frames included in the target video data according to the number M of sampled video frames indicated by the video sampling rule, to obtain M video frames in the target video data.

[0142] Specifically, the computer device may obtain the original video data, which may refer to the video data obtained by the computer device through the camera component by shooting the target object, or may refer to the video data uploaded by the user. The computer device may obtain the content attributes of each original video frame in the original video data, and the content attributes may refer to the video action type, the video language type, etc. The computer device may divide the original video data according to the video action type of each original video frame to obtain N video segments, so as to ensure that the video action types in the video frames in each video segment are the same, that is, to ensure that each video segment contains a single action content. N is a positive integer, and N may take values such as 1, 2, 3,.... Specifically, the computer device may sequentially divide the original video frames with the same video action type into the same video segment according to the time order between the shooting timestamps of each original video frame in the original video data. For example, the original video data includes original video frames 1, 2, 3,... 10, 11,... 20 arranged in sequence according to the shooting timestamp. If the video action types of the original video frames from 1 to 10 are the same (the same action), the original video frames from 1 to 10 may be divided into one video segment. If the video action types of the original video frames from 11 to 20 are the same, the original video frames from 11 to 20 may be divided into one video segment.

[0143] Among them, since the amount of original video data is large and there are many categories of video content in the original video data (such as including various action types such as cycling, rafting, walking, etc.), directly classifying the original video data involves a large amount of data, and the classification is inaccurate due to the large number of video content categories. Therefore, the original video data can be divided according to a single content attribute to obtain N video segments, and then each video segment is classified to obtain the video content category of each video segment, and the video content category of each video segment is determined as the video content category of the original video data, thereby improving the accuracy of classifying the original video data. Among them, to ensure the robustness of the target classification model for classifying each video segment, the duration segments of the target duration can be extended forward and backward for each video segment.

[0144] Furthermore, the computer device can randomly select a target video segment from the N video segments as the target video data. Of course, the computer device can also sequentially determine each video segment in the N video segments as the target video segment according to the time order of the timestamps of each video segment. The computer device can perform video frame sampling on the original video frames included in the target video data according to the number of sampled video frames M indicated by the video sampling rule to obtain M video frames in the target video data. The video sampling rule can refer to striding frame capture, that is, capturing every other frame. For example, if the target video data includes original video frame 1, original video frame 2, and original video frame 3, and the sampled video frame data M is 2, then the original video frames in the target video data can be stride-frame captured, the original video frame 1 can be sampled, the original video frame 2 can be discarded, and the original video frame 3 can be sampled to obtain 2 video frames in the target video data.

[0145] Optionally, the computer device can also perform a single action division on each video segment to obtain multiple video sub-segments, each video sub-segment having the same action posture, and then according to the video sampling rule, ensure that one or more video frames are sampled from each video sub-segment, so as to ensure that each action posture can be sampled, thereby avoiding the loss of key information caused by random sampling and improving the accuracy of subsequent video data classification.

[0146] Optionally, the computer device can also directly perform video frame sampling on the original video frames included in the original video data according to the number of sampled video frames M indicated by the video sampling rule, to obtain M video frames in the target video data. By performing video frame sampling on the original video data to obtain M video frames in the target video data, in this way, video data of different lengths can all be sampled into M video frames, thereby converting video data of different lengths into fixed-length video data, that is, converting video data of different lengths into a fixed-length video frame sequence (i.e., M video frames), which is convenient for the target classification model to perform subsequent classification service processing.

[0147] Specifically, as Figure 4 shown, Figure 4 FIG. is a schematic diagram of using a target classification model to predict the video content category provided by an embodiment of the present application. As Figure 4 shown, the computer device can call the target classification model 40b, and input the M video frames 40a in the target video data into the target classification model 40b. The target classification model 40b is used to classify the video data to obtain the video content category of the video data. The video content category can indicate the behavior category, scene category, etc. in the video. The video content category can also refer to video action categories (such as cars, rafting, driving, etc.), video categories (such as educational videos, entertainment videos, etc.), video language categories (such as domestic dramas, British dramas, Korean dramas, etc.), etc. The computer device can respectively perform first feature extraction processing on the M video frames through the first feature extraction layer 40c in the target classification model 40b to obtain first feature maps respectively corresponding to the M video frames, and combine the first feature maps respectively corresponding to the M video frames to obtain a first feature map sequence 40f. Among them, the computer device can perform permutation and combination on the first feature maps respectively corresponding to the M video frames according to the shooting timestamps respectively corresponding to the M video frames to obtain a first feature map sequence. For example, the computer device can arrange the first feature map of the video frame with the earliest shooting timestamp among the M video frames at the forefront according to the time order of the shooting timestamps respectively corresponding to the M video frames, and then sequentially arrange the first feature maps of the video frames according to the time order of the video frames to obtain the first feature map sequence 40f.

[0148] Furthermore, as Figure 4As shown in the figure, the computer device can perform second feature extraction processing on M video frames through the second feature extraction layer 40d in the target classification model 40b to obtain second feature maps corresponding to the M video frames respectively, and combine the second feature maps corresponding to the M video frames respectively to obtain a second feature map sequence 40g. Similarly, the computer device can perform permutation and combination on the second feature maps corresponding to the M video frames according to the shooting timestamps of the M video frames to obtain a second feature map sequence. The specific permutation and combination method of the second feature map sequence can refer to the content of obtaining the first feature map sequence above, and will not be elaborated in this embodiment of the present application.

[0149] Among them, both the first feature extraction layer 40c and the second feature extraction layer 40d in the target classification model 40b can be convolutional network layers, and can perform convolution and pooling processing on each of the M video frames to obtain a feature map corresponding to each video frame. Among them, the network parameters (i.e., model parameters) in the first feature extraction layer 40c and the second feature extraction layer 40d can be the same or different. If the network parameters in the first feature extraction layer 40c and the second feature extraction layer 40d are the same, then the first feature map sequence 40f is the same as the second feature map sequence 40g. If the network parameters in the first feature extraction layer 40c and the second feature extraction layer 40d are different, then the first feature map sequence 40f and the second feature map sequence 40g are different.

[0150] S102, sample the first feature map sequence according to the first time sampling parameter to obtain a target first feature map, and sample the second feature map sequence according to the second time sampling parameter to obtain a target second feature map.

[0151] Specifically, the computer device can sample the first feature map sequence according to the first time sampling parameter to obtain a target first feature map, and this first time sampling parameter is used to extract a first feature map from the first feature map sequence as the target first feature map. Sample the second feature map sequence according to the second time sampling parameter to obtain a target second feature map, and this second time sampling parameter is used to extract a second feature map from the second feature map as the target second feature map. The sum of the number of video frames corresponding to the target first feature map and the target second feature map is equal to M, and the video frames corresponding to the target first feature map and the target second feature map are different from each other.

[0152] Optionally, the computer device may obtain an initial time sampling parameter; the number of sampling elements in the initial time sampling parameter is M. Randomly determine the element values of M sampling elements with a position order in the initial time sampling parameter to obtain a first time sampling parameter; the element values include a first element threshold and a second element threshold, the first element threshold is used to indicate sampling of the feature map, and the second element threshold is used to indicate masking of the feature map. Determine a second time sampling parameter according to the element values of the M sampling elements in the first time sampling parameter; the element values of the sampling elements in the same position in the first time sampling parameter and the second time sampling parameter are different from each other.

[0153] Specifically, the computer device may obtain an initial time sampling parameter, and the number of sampling elements in the initial time sampling parameter is the same as the number of M video frames in the target video data, that is, the number of sampling elements in the initial time sampling parameter is equal to M. The initial time sampling parameter includes M sampling elements with a position order, and the initial element values of the M sampling elements with a position order may all be empty. The computer device may randomly determine the element values of the M sampling elements with a position order in the initial time sampling parameter to obtain a first time sampling parameter. Among them, the element value of the sampling element may be the first element threshold or the second element threshold. When the element value of the sampling element is the first element threshold, it indicates sampling of the feature map. When the element value of the sampling element is the second element threshold, it indicates masking (i.e., not sampling) of the feature map. Among them, the first element threshold may be 1, which is used to indicate sampling of the feature map, and the second element threshold is 0, which is used to indicate discarding (i.e., not sampling) of the feature map. The computer device may randomly set the element values of the M sampling elements in the initial time sampling parameter to the first element threshold or the second element threshold to obtain a first time sampling parameter.

[0154] Further, the computer device may determine a second time sampling parameter according to the element values of the M sampling elements in the first time sampling parameter. The element values of the sampling elements in the same position in the first time sampling parameter and the second time sampling parameter are different from each other. That is, if the element value of the sampling element in the first position in the first time sampling parameter is the first element threshold, then the element value of the sampling element in the first position in the second time sampling parameter is the second element threshold. For example, if the first time sampling parameter is 0110, then the second time sampling parameter is 1001.

[0155] Further, after the computer device obtains the first time sampling parameter and the second time sampling parameter, the specific manner of sampling the first feature map sequence according to the first time sampling parameter and sampling the second feature map sequence according to the second time sampling parameter may include: calling a target classification model. In the feature fusion layer of the target classification model, based on the first element threshold in the first time sampling parameter, sample the associated feature maps in the first feature map sequence. Based on the second element threshold in the first time sampling parameter, mask the associated feature maps in the first feature map sequence to obtain a target first feature map. According to the first element threshold in the second time sampling parameter, sample the associated feature maps in the second feature map sequence, and according to the second element threshold in the second time sampling parameter, mask the associated feature maps in the second feature map sequence to obtain a target second feature map.

[0156] Specifically, as Figure 4 shown, the computer device may sample the associated feature maps in the first feature map sequence 40f based on the first element threshold in the first time sampling parameter 40h in the feature fusion layer 40e of the target classification model 40b. Among them, the M sampling elements in the first time sampling parameter 40h have a position order, and the M first feature maps in the first feature map sequence 40f also have a position order. Then, the first feature maps with the same position order in the first feature map sequence 40f can be sampled according to the first element threshold with a position order in the first time sampling parameter 40h. It can be explained that if the element value ranked second in the first time sampling parameter 40h is the first element threshold, the first feature map ranked second in the first feature map sequence 40f can be sampled. The computer device may mask (i.e., not sample) the first feature maps with the same position order in the first feature map sequence 40f according to the second element threshold with a position order in the first time sampling parameter 40h to obtain the target first feature map 40j.

[0157] For example, as Figure 4 shown, if the first time sampling parameter 40h is 0110 (i.e., the element values ranked second and third are the first element thresholds), the first feature map sequence 40f includes the first feature map a1, the first feature map b1, the first feature map c1, and the first feature map d1 with a position order. The computer device may sample the first feature map b1 and the first feature map c1 with the same position in the first feature map sequence 40f according to the position of the first element threshold in the first time sampling parameter 40h. According to the position of the second element threshold in the first time sampling parameter 40h, mask (i.e., not sample) the first feature map a1 and the first feature map d1 with the same position in the first feature map sequence 40f to obtain the target first feature map 40k (i.e., the first feature map b1 and the first feature map c1).

[0158] Similarly, as Figure 4 shown, the computer device may sample the second feature maps with the same position order in the second feature map sequence 40g according to the first element threshold with a position order in the second time sampling parameter 40i, and mask (i.e., not sample) the second feature maps with the same position order in the second feature map sequence 40g according to the second element threshold with a position order in the second time sampling parameter 40i, to obtain a target second feature map. For example, as Figure 4 shown, if the first time sampling parameter is 0110, the second time sampling parameter 40i is 1001, and the second feature map sequence 40g includes second feature maps a2, b2, c2, and d2 with a position order. The computer device may sample the second feature maps a2 and d2 with the same position order in the second feature map sequence 40g according to the position of the first element threshold in the second time sampling parameter 40i, and mask (i.e., not sample) the second feature maps b2 and the first feature map c2 with the same position order in the second feature map sequence 40g according to the position of the second element threshold in the second time sampling parameter 40i, to obtain a target second feature map 40j (i.e., the second feature maps a2 and d2).

[0159] S103. Generate a time fusion feature map according to the target first feature map and the target second feature map.

[0160] Specifically, the computer device may fuse the target first feature map and the target second feature map to obtain a target time fusion feature map. For example, the computer device may perform permutation and combination on the target first feature map and the target second feature map to obtain a time fusion feature map.

[0161] Optionally, the specific manner in which the computer device generates a time fusion feature map according to the target first feature map and the target second feature map may include: obtaining a first timestamp of the video frame corresponding to the target first feature map, and obtaining a second timestamp of the video frame corresponding to the target second feature map. Combining the target first feature map and the target second feature map according to the time order between the first timestamp and the second timestamp to obtain a time fusion feature map.

[0162] Specifically, as Figure 4As shown, the computer device can obtain the first timestamp of the video frame corresponding to the target first feature map 40k, obtain the second timestamp of the video frame corresponding to the target second feature map 40j, and perform permutation and combination on the target first feature map 40k and the target second feature map 40j according to the chronological order between the first timestamp and the second timestamp to obtain a temporal fusion feature map. It can be understood that the computer device can perform permutation and combination on the target first feature map 40k and the target second feature map 40j according to the chronological order of shooting time to obtain a temporal fusion feature map. For example, as Figure 4 shown, the target first feature map 40k includes a first feature map b1 and a first feature map c1. The timestamp of the video frame corresponding to the first feature map b1 is 12:02, and the timestamp of the video frame corresponding to the first feature map c1 is 12:03. The target second feature map 40j includes a second feature map a2 and a second feature map d2. The timestamp of the video frame corresponding to the second feature map a2 is 12:01, and the timestamp of the video frame corresponding to the second feature map d2 is 12:04. The computer device can sort and combine according to the chronological order of the timestamps of the video frames corresponding to each feature map to obtain a temporal fusion feature map 40l (i.e., the second feature map a2, the first feature map b1, the first feature map c1, and the second feature map d2).

[0163] S104. Generate a target fusion feature map based on the temporal fusion feature map, the first feature map sequence, and the second feature map sequence, and classify the target fusion feature map to obtain the video content category of the target video data.

[0164] Specifically, as Figure 4 shown, the computer device can generate a target fusion feature map 40m based on the temporal fusion feature map, the first feature map sequence, and the second feature map sequence. As Figure 4 shown, the computer device can perform convolution processing on the target fusion feature map 40m through the convolution layer 40n in the target classification model 40b to obtain the target fusion feature map after convolution processing. The computer device can perform classification processing on the target fusion feature map after convolution processing through the classification layer 40o in the target classification model to obtain the video content category 40p of the target video data 40a. In this way, this application does not need to rely on text information such as video labels and video introductions of the target video data, nor does it need to rely on manual experience analysis. It enhances the features of the target video data in the time dimension, and classifies the target video data by using the target fusion feature map after feature enhancement, so as to accurately classify the target video data.

[0165] Optionally, the computer device may perform feature map addition on the temporal fusion feature map, the first feature map sequence, and the second feature map sequence to obtain a target fusion feature map, perform convolution processing on the target fusion feature map through the convolutional layer in the target classification model, and perform classification on the target fusion feature map after convolution processing through the splitting layer in the target classification model, so as to obtain the video content category of the target video data.

[0166] Optionally, the computer device may perform pixel mixing and splicing on the first feature map sequence and the second feature map sequence to obtain a pixel fusion feature map, and generate a target fusion feature map 40m according to the temporal fusion feature map and the pixel fusion feature map. The computer device may perform convolution processing on the target fusion feature map 40m through the convolutional layer 40n in the target classification model 40b, and perform classification processing on the target fusion feature map after convolution processing through the classification layer 40o in the target classification model 40b to obtain the video content category 40p of the target video data. In this way, by fusing the feature maps corresponding to the M video frames from the temporal dimension and the pixel dimension, that is, performing feature enhancement on the feature maps corresponding to the M video frames, the accuracy of classifying the target video data can be improved.

[0167] In the embodiments of the present application, M video frames in the target video data are obtained through a video sampling rule, and the target video data is preprocessed. While ensuring a reduction in the amount of computation, it is also possible to avoid the loss of key information in the target video data caused by sampling, which can improve the efficiency of subsequent classification of the target video data. Further, a first feature extraction process is performed on the M video frames to obtain a first sequence of feature maps, and a second feature extraction process is performed on the M video frames to obtain a second sequence of feature maps. By performing the first feature extraction and the second feature extraction on the M video frames, a first sequence of feature maps and a second sequence of feature maps of the M video frames are obtained, and different feature information of the M video frames can be extracted from different perspectives. Further, the first sequence of feature maps is sampled according to a first time sampling parameter to obtain a target first feature map, and the second sequence of feature maps is sampled according to a second time sampling parameter to obtain a target second feature map. The video frames corresponding to the target first feature map and the target second feature map respectively constitute the M video frames, and a time fusion feature map is generated based on the target first feature map and the target second feature map. It can be seen that by sampling the first sequence of feature maps and the second sequence of feature maps respectively in the time dimension, a time fusion feature map is obtained, and based on this, the feature enhancement of the target video data is performed according to the temporal information between each video frame, improving the feature enhancement effect of the target video data. Further, the first sequence of feature maps and the second sequence of feature maps can be feature-enhanced through other sampling methods (such as pixel sampling) to obtain fusion feature maps in other dimensions, and a target fusion feature map is generated based on the time fusion feature map in the time dimension and the fusion feature maps in other dimensions (such as the pixel dimension), and the target fusion feature map is classified to obtain the video content category of the target video data. It can be seen that the present application does not need to rely on text information such as video labels and video introductions of the target video data, nor does it need to rely on manual experience analysis. By performing feature enhancement on the target video data in the time dimension and other dimensions, and using the feature-enhanced target fusion feature map to classify the target video data, accurate classification of the target video data can be achieved.

[0168] Please refer to Figure 5 , Figure 5 which is a schematic flowchart of a data processing method provided by an embodiment of the present application. This data processing method can be executed by a computer device, and the computer device can be a server (such as server 10 in the above Figure 1 ), or a user terminal (such as any user terminal in the user terminal cluster in the above Figure 1 ), and the present application does not make any limitations in this regard. As Figure 5 shown, this data processing method may include but is not limited to the following steps:

[0169] S201, obtaining M video frames in the target video data, performing a first feature extraction process on the M video frames to obtain a first feature map sequence, and performing a second feature extraction process on the M video frames to obtain a second feature map sequence.

[0170] S202, sampling the first feature map sequence according to the first time sampling parameter to obtain a target first feature map, and sampling the second feature map sequence according to the second time sampling parameter to obtain a target second feature map.

[0171] S203: Generate a temporal fusion feature map according to the target first feature map and the target second feature map.

[0172] Specifically, the specific contents of steps S201 to S203 in the embodiment of the present application can be found in the above Figure 3 The specific contents of step S101 to step S103 are not repeated here in this embodiment of the present application.

[0173] S204, in the first feature map sequence and the second feature map sequence, pixel-mixing and splicing the first feature map and the second feature map associated with the same video frame to obtain pixel-mixing feature maps corresponding to M video frames respectively.

[0174] Specifically, the computer device may perform pixel mixing and splicing on the first feature map and the second feature map associated with the same video frame in the first feature map sequence and the second feature map sequence to obtain pixel mixing feature maps corresponding to the M video frames. It is understandable that the computer device may perform pixel mixing and splicing on the first feature map and the second feature map corresponding to each video frame to obtain a pixel mixing feature map corresponding to each video frame, that is, one video frame corresponds to one pixel mixing feature map.

[0175] Optionally, the M video frames include video frame M i , i is a positive integer less than or equal to M. If M is 3, i can be 1, 2, or 3. The specific method for the computer device to obtain the pixel mixed feature maps corresponding to the M video frames may include: calling the target classification model, obtaining the video frame M in the first feature map sequence through the feature fusion layer in the target classification model, i The corresponding first feature map, obtain the video frame M in the second feature map sequence i According to the first pixel sampling parameter, the video frame M i The corresponding first feature map is pixel sampled to obtain a first pixel sampling feature map, and the video frame M is sampled according to the second pixel sampling parameter. i The corresponding second feature map is pixel sampled to obtain a second pixel sampling feature map. The first pixel sampling feature map and the second pixel sampling feature map are pixel mixed and spliced ​​to obtain a video frame M iThe corresponding pixel-mixed feature map.

[0176] Specifically, the computer device can call the target classification model and obtain the video frame M from the first feature map sequence through the feature fusion layer in the target classification model. i The corresponding first feature map, and obtain the video frame M from the second feature map sequence. i The corresponding second feature map. Among them, the computer device can randomly extract a threshold within the interval (0, 1) as the first pixel sampling parameter, and the second pixel sampling parameter can be the difference between the first pixel sampling parameter and the threshold 1. The computer device can perform pixel sampling on the first feature map corresponding to the video frame M according to the first pixel sampling parameter to obtain the first pixel sampling feature map. It can be understood that the computer device can perform weighted processing on the first feature map corresponding to the video frame M using the first pixel sampling parameter to obtain the first pixel sampling feature map. The computer device can perform weighted processing on the second feature map corresponding to the video frame M using the second pixel sampling feature map to obtain the second pixel sampling feature map. The computer device can perform hybrid splicing on the corresponding first pixel sampling feature map and second pixel sampling feature map of the video frame M to obtain the pixel-mixed feature map corresponding to the video frame Mi. i i i i The corresponding first pixel sampling feature map and second pixel sampling feature map of the video frame M to obtain the pixel-mixed feature map corresponding to the video frame Mi.

[0177] Optionally, the calculation formula for the first pixel sampling parameter can be the following formula (1):

[0178]

[0179] Among them, r in formula (1) is a hyperparameter (i.e., tuning parameter, which needs to be set manually), and its value range is [0, +∞], k is the pixel sampling weight, and its value range is (0, 1). The second pixel sampling parameter is the difference between the threshold 1 and the first pixel sampling parameter.

[0180] Optionally, the specific way for the computer device to perform hybrid splicing on the corresponding first pixel sampling feature map and second pixel sampling feature map of the video frame M to obtain the pixel-mixed feature map corresponding to the video frame M can include: The computer device can determine the area to be filled in the first pixel sampling feature map, and determine the target area in the second pixel sampling feature map with the same area size as the area to be filled. The computer device can cut the target area in the second pixel sampling feature map and fill the area to be filled in the first pixel sampling area with the target area in the second pixel sampling feature map to obtain the pixel-mixed feature map corresponding to the video frame M. i i i The corresponding pixel-mixed feature map. ​​​​​

[0181] Optionally, the computer device may obtain an initial feature map. The size of the feature map of the initial feature map is the same as the size of the feature map of the first pixel sampling feature map and the second region sampling feature map, and the initial feature map is a blank feature map. The computer device may randomly determine a region to be filled in the initial feature map, cut the first pixel sampling feature map to obtain a first feature map region with the same region size as the region to be filled, and fill the first feature map region into the region to be filled in the initial feature map. The computer device may cut the second pixel sampling feature map to obtain a second feature map region with the same region size as the other regions in the initial feature map except the region to be filled, and fill the second feature map region into the other regions in the initial feature map to obtain video frame M i The corresponding pixel mixed feature map.

[0182] S205, generate a pixel fusion feature map according to the pixel mixed feature maps respectively corresponding to the M video frames.

[0183] Specifically, the computer device may splice and combine the pixel mixed feature maps respectively corresponding to the M video frames according to the time sequence between the timestamps of each video frame to obtain a pixel fusion feature map. For example, the computer device may arrange the pixel mixed feature map corresponding to the video frame with an earlier shooting time in the front, arrange the pixel mixed feature map corresponding to the video frame with a later shooting time in the back, and sequentially arrange and combine the corresponding pixel mixed feature maps according to the shooting timestamp of each video frame to obtain a pixel fusion feature map.

[0184] S206, generate a target fusion feature map according to the time fusion feature map and the pixel fusion feature map, and classify the target fusion feature map to obtain the video content category of the target video data.

[0185] Specifically, the computer device may generate a target fusion feature map according to the time fusion feature map and the pixel fusion feature map. In this way, by enhancing the features of the feature maps of the M video frames of the target video data in the time dimension and the pixel dimension, the accuracy of classifying the target video data can be improved. Further, the computer device may call a target classification model, perform convolution processing on the target fusion feature map through the convolution layer in the target classification model, and extract the feature information of the target fusion feature map. The computer device may perform classification processing on the target fusion feature map after convolution processing through the classification layer in the target classification model to obtain the video content category of the target video data.

[0186] As Figure 6 shown, Figure 6 is a schematic diagram of a method for obtaining a target fusion feature map provided by an embodiment of the present application. As Figure 6As shown, the computer device can perform temporal sampling on the first feature map sequence 60a using the first temporal sampling parameter 60c. As Figure 6 shown, the computer device samples the first feature map ranked first and the first feature map ranked fourth in the first feature map sequence 60a, and uses the first feature map ranked first and the first feature map ranked fourth as the target first feature maps. The computer device can perform temporal sampling on the second feature map sequence 60b using the second temporal sampling parameter 60d. As Figure 6 shown, the computer device samples the second feature map ranked second and the second feature map ranked third in the second feature map sequence 60b, and uses the second feature map ranked second and the second feature map ranked third as the target second feature maps. The computer device can combine the target first feature maps and the target second feature maps in the chronological order of the timestamps corresponding to the video frames to obtain the temporal fusion feature map 60e. Further, the computer device can perform pixel sampling (i.e., pixel weighting) on each first feature map in the first feature map sequence 60a using the first pixel sampling parameter 60f to obtain the first pixel sampling feature map corresponding to each first feature map. The computer device can perform pixel sampling (i.e., pixel weighting) on each second feature map in the second feature map sequence 60b using the second pixel sampling parameter 60g to obtain the second pixel sampling feature map corresponding to each second feature map. The computer device can perform hybrid splicing on the first pixel sampling feature map and the second pixel sampling feature map with the same position order to obtain the pixel hybrid feature map corresponding to each video frame.

[0187] Further, the computer device can arrange and combine the pixel hybrid feature maps corresponding to M video frames respectively in the chronological order of the timestamps of each video frame to obtain the pixel fusion feature map 60h. The specific content can refer to the content of step S206 above and will not be elaborated here. The computer device can add the temporal fusion feature map 60e and the pixel fusion feature map 60h to obtain the target fusion feature map 60i.

[0188] Optionally, the specific manner for the computer device to obtain the video content category of the target video data may include: calling the target classification model, adding the temporal fusion feature map and the pixel fusion feature map through the feature fusion layer in the target classification model to obtain the target fusion feature map. Performing convolution processing on the target fusion feature map through the convolution layer in the target classification model to obtain the target fusion feature map after convolution processing, and performing classification processing on the target fusion feature map after convolution processing through the classification layer in the target classification model to obtain the video content category of the target video data.

[0189] Specifically, the computer device can call the target classification model, and through the feature fusion layer in the target classification model, add the time fusion feature map and the pixel fusion feature map to obtain the target fusion feature map. Among them, since the time fusion feature map is obtained by arranging and combining the target first feature map and the target second feature map in the time order of the timestamps of the corresponding video frames, and the pixel fusion feature map is obtained by arranging and combining the pixel mixing feature maps corresponding to M video frames in the time order of the timestamps of the video frames. Therefore, the time fusion feature map includes M feature maps with a position order. The pixel fusion feature map also includes M feature maps with a position order. Therefore, the feature maps with the same position order in the time fusion feature and the pixel fusion feature map can be added (i.e., fused) to obtain the target fusion feature map.

[0190] Further, the computer device can perform convolution processing on the target fusion feature map through the convolution layer in the target classification model to obtain the target fusion feature map after convolution processing. The convolution layer in the target classification model can include multiple convolution sub-layers and fully connected sub-layers. Among them, the convolution layer is used to eliminate noise and enhance features for the target fusion feature map. Each convolution sub-layer in the convolution layer corresponds to one or more convolution kernels (kernels, which can also be called filters, or receptive fields). The number of channels of the convolution kernels in each convolution sub-layer is determined by the number of channels of the input data. The number of channels of the output data (i.e., the image feature information) of each layer is determined by the number of convolution kernels in the convolution sub-layer, and the image height H out and the image width W out (i.e., the second and third dimensions in the output data) are jointly determined by the size of the input data, the size of the convolution kernel, the stride, and the padding, that is, H out =(H in -H kernel +2*padding) / stride+1, W out =(W in -W kernel +2*padding) / stride+1. H in ,H kernel respectively represent the height of the input video frame and the height of the convolution kernel; W in ,W kernel respectively represent the width of the input video frame and the width of the convolution kernel. Through the fully connected sub-layer in the convolution layer, the feature information after convolution processing of multiple convolution sub-layers can be subjected to feature classification processing to find the key feature information.

[0191] Further, the computer device can perform convolution processing on the target fusion feature map through the convolution layer in the target classification model. After obtaining the target fusion feature map after convolution processing, the computer device can perform classification processing on the target fusion feature map after convolution processing through the classification layer in the target classification model to obtain the video content category of the target video data. Among them, the classification layer can include multiple fully connected layers, and these multiple fully connected layers can act as a "classifier" to map the learned "distributed feature representation" to the sample label space. It can be understood that the fully connected layer can be implemented by convolution operations, that is, the fully connected layer can be transformed into a convolution with a convolution kernel of 1x1 to linearly transform one feature space to another feature space, thereby achieving classification.

[0192] Optionally, the specific method for the computer device to perform classification processing on the target fusion feature map after convolution processing may include: inputting the target fusion feature map after convolution processing into the first classification sub-layer in the classification layer to perform classification processing on the target fusion feature map after convolution processing to obtain a first classification result. Inputting the target fusion feature map after convolution processing into the second classification sub-layer in the classification layer to perform classification processing on the target fusion feature map after convolution processing to obtain a second classification result. Obtaining the average value of the first classification result and the second classification result, and determining the video content category of the target video data according to this average value.

[0193] Specifically, the computer device can input the target fusion feature map after convolution processing into the first classification sub-layer in the classification layer. The first classification sub-layer can be a fully connected network structure. Through this first classification sub-layer, classification processing is performed on the target fusion feature map after convolution processing to obtain a first classification result. The computer device can input the target fusion feature map after convolution processing into the second classification sub-layer in the classification layer. The second classification sub-layer can also be a fully connected network structure. Through this second classification sub-layer, classification processing is performed on the target fusion feature map after convolution processing to obtain a second classification result. Among them, the network parameters in the first classification sub-layer are different from the network parameters in the second classification sub-layer, so that the target fusion feature map after convolution processing can be feature-classified from different angles to obtain results with different possibilities. The computer device can obtain the average value of the first classification result and the second classification result, and perform normalization processing (i.e., activation processing softmax) on this average value to obtain the video content category of the target video data. In this way, the target video data can be classified from multiple angles, and the video content category of the target video data can be predicted based on multiple classification results, which can improve the classification accuracy of the target video data.

[0194] In an embodiment of the present application, M video frames in target video data are obtained through a video sampling rule, and the target video data is preprocessed. While ensuring a reduction in the amount of computation, it is also possible to avoid the loss of key information in the target video data caused by sampling, which can improve the efficiency of subsequent classification of the target video data. Further, the M video frames are subjected to a first feature extraction process to obtain a first feature map sequence, and the M video frames are subjected to a second feature extraction process to obtain a second feature map sequence. By performing the first feature extraction and the second feature extraction on the M video frames, the first feature map sequence and the second feature map sequence of the M video frames are obtained, and different feature information of the M video frames can be extracted from different perspectives. Further, the first feature map sequence is sampled according to a first time sampling parameter to obtain a target first feature map, the second feature map sequence is sampled according to a second time sampling parameter to obtain a target second feature map, and a temporal fusion feature map is generated based on the target first feature map and the target second feature map. It can be seen that the first feature map sequence and the second feature map sequence are respectively sampled in the time dimension to obtain a temporal fusion feature map, and based on this, the feature enhancement of the target video data is performed according to the temporal information between each video frame, improving the feature enhancement effect of the target video data. Further, the first feature map sequence and the second feature map sequence are pixel-sampled to obtain a pixel fusion feature map, a target fusion feature map is generated based on the temporal fusion feature map in the time dimension and the pixel fusion feature map in the pixel dimension, and the target fusion feature map is classified to obtain the video content category of the target video data. It can be seen that the present application performs feature enhancement on the target video data in the time dimension and the pixel dimension, and classifies the target video data by sampling the target fusion feature map after feature enhancement, which can improve the accuracy of classifying the target video data.

[0195] Please refer to Figure 7 , Figure 7 which is a schematic flowchart of a data processing method provided by an embodiment of the present application. This data processing method can be executed by a computer device, which can be a server (such as server 10 in the above Figure 1 ), or a user terminal (such as any user terminal in the user terminal cluster in the above Figure 1 ), and the present application does not make any limitations in this regard. As Figure 7 shown, this data processing method may include but is not limited to the following steps:

[0196] S301, through an initial classification model, perform a first feature extraction process on M first sample video frames in the first sample video data to obtain a first sample feature map sequence, and perform a second feature extraction process on M second sample video frames in the second sample video data to obtain a second sample feature map sequence.

[0197] Specifically, data augmentation (i.e., image feature augmentation) can improve the generalization and robustness of the model, thereby improving the prediction effect and applicability of the model. Specifically, the computer device can obtain the initial classification model, M first sample video frames in the first sample video data, and M second sample video frames in the second sample video data, and obtain the first video content category label corresponding to the first sample video data and the second video content category label corresponding to the second sample video data. Among them, M is a positive integer, such as M can take values of 1, 2, 3... Among them, the first video content category label corresponding to the first sample video data and the second video content category label corresponding to the second sample video data can be manually labeled or obtained by other means.

[0198] Furthermore, the computer device can perform first feature extraction processing on the M first sample video frames in the first sample video data through the first feature extraction layer in the initial classification model to obtain a first sample feature map sequence corresponding to the M first sample video frames. The computer device can perform second feature extraction processing on the M second sample video frames in the second sample video data through the second feature extraction layer in the initial classification model to obtain a second sample feature map sequence corresponding to the M second sample video frames. Among them, the first feature extraction layer and the second feature extraction layer can be a convolutional neural network or an attention network. The convolutional neural network can perform convolution processing on the video frame (i.e., image) and the convolution kernel (i.e., filter) to obtain the feature map corresponding to the video frame, and the feature map can also perform convolution processing with the convolution kernel to generate a new feature map. The attention network (i.e., Transformer) can learn the sequential relationship between sequences and reduce the distance between any two positions in the sequence to a constant, so as to extract the correlation relationship between each video frame.

[0199] S302, sample the first sample feature map sequence according to the first sample time sampling parameter to obtain the target first sample feature map, and sample the second sample feature map sequence according to the second sample time sampling parameter to obtain the target second sample feature map.

[0200] Specifically, the computer device can sample the first sample feature map sequence according to the first sample time sampling parameter to obtain the target first sample feature map, and sample the second sample feature map sequence according to the second sample time sampling parameter to obtain the target second sample feature map. Among them, the first sample feature maps corresponding to the first sample video frames with earlier shooting times in the first sample feature map sequence are arranged in the front, and the first sample feature maps corresponding to the first sample video frames with later shooting times are arranged in the back. That is, each first sample feature map in the first sample feature map sequence is sorted and combined according to the time stamp of the corresponding first sample video frame. Similarly, the first sample feature maps corresponding to the second sample video frames with earlier shooting times in the second sample feature map sequence are arranged in the front, and the first sample feature maps corresponding to the second sample video frames with later shooting times are arranged in the back. That is, each first sample feature map in the first sample feature map sequence is sorted and combined according to the time stamp of the corresponding second sample video frame. Among them, the sum of the numbers of the sample video frames corresponding to the target first sample feature map and the target second sample feature map is equal to M. It can be understood that the sum of the number of the first sample video frames corresponding to the target first sample feature map and the number of the second sample video frames corresponding to the target second sample feature map is equal to M. It can be understood that the computer device can extract i first sample feature maps from the first sample feature map sequence according to the first time sampling parameter as the target first sample feature map, and the computer device can extract j second sample feature maps from the second sample feature map according to the second time sampling parameter as the target second sample feature map, and the sum of i and j is equal to M. Among them, the first sample feature map sequence includes M first sample feature maps with a positional order, that is, each first sample feature map has different prior and posterior position information. Similarly, the second sample feature map sequence includes M second sample feature maps with a positional order. The position of the target first sample feature map in the first sample feature map sequence is different from the position of the target second sample feature map in the second sample feature map sequence. For the specific content, reference can be made to the content of step S102 in the above Figure 3 and details are not described herein again in the embodiments of the present application.

[0201] For example, the first sample feature map sequence includes the first sample feature map p1 arranged in the first position, the first sample feature map p2 arranged in the second position, the first sample feature map p3 arranged in the third position, and the first sample feature map p4 arranged in the fourth position. The second sample feature map sequence includes the second sample feature map q1 arranged in the first position, the second sample feature map q2 arranged in the second position, the second sample feature map q3 arranged in the third position, and the second sample feature map q4 arranged in the fourth position. If the computer device can determine the first sample feature map p2 and the first sample feature map p3 as the target first sample feature maps from the first sample feature map sequence according to the first time sampling parameter, then the target second sample feature maps are the second sample feature map q1 and the second sample feature map q4, that is, the positions of the target first sample feature maps in the first sample feature map sequence are different from the positions of the target second sample feature maps in the second sample feature map sequence.

[0202] S303. Generate a sample time fusion feature map according to the target first sample feature map and the target second sample feature map, generate a target sample fusion feature map for predicting the video content category according to the sample time fusion feature map, the first sample feature map sequence, and the second sample feature map sequence, and adjust the parameters of the initial classification model according to the target sample fusion feature map to obtain a target classification model.

[0203] Specifically, the computer device can obtain the first sample timestamp of the first sample video frame corresponding to the target first sample feature map and obtain the second sample timestamp of the second sample video frame corresponding to the target second sample feature map. According to the time sequence of the first sample timestamp and the second sample timestamp, the target first sample feature map and the target second sample feature map are arranged and combined to obtain a sample time fusion feature map. Further, the computer device can perform pixel mixing and splicing on the feature maps with the same position in the first sample feature map sequence and the second sample feature map sequence to obtain M sample pixel mixing feature maps. The specific content can refer to the content of step S204 above. Figure 5 This application embodiment will not be elaborated here. The computer device can arrange and combine the M sample pixel mixing feature maps according to the position sequence corresponding to the M sample pixel mixing feature maps to obtain a sample pixel mixing feature map.

[0204] Further, the computer device can fuse the sample time fusion feature map and the sample pixel fusion feature map (i.e., add the feature maps) to obtain a target sample fusion feature map for predicting the video content category. According to the target sample fusion feature map, the parameters of the initial classification model are adjusted to obtain a target classification model, and the target classification model is used to predict the video content category of the target video data.

[0205] Optionally, the specific manner in which the computer device adjusts the parameters of the initial classification model according to the target sample fusion feature map to obtain the target classification model may include: predicting the first predicted video content category of the first sample video data according to the target sample fusion feature map, and predicting the second predicted video content category of the second sample video data according to the target sample fusion feature map. Generating a first loss function according to the first video content category label and the first predicted video content category of the first sample video data. Generating a second loss function according to the second video content category label and the second predicted video content category of the second sample video data. Generating a total loss function according to the first loss function and the second loss function, and adjusting the parameters of the initial classification model according to the total loss function. When the initial classification model after parameter adjustment meets the training convergence condition, the initial classification model after parameter adjustment is determined as the target classification model.

[0206] Specifically, the computer device may perform first feature classification on the target sample fusion feature map through the first classification layer in the initial classification model to obtain the first predicted video content category of the first sample video data. The computer device may perform second feature classification on the target sample fusion feature map through the second classification layer in the initial classification model to obtain the second predicted video content category of the second sample video data. Further, the computer device may generate a first loss function according to the first video content category label and the first predicted video content category of the first sample video data, and calculate the error between the first video content category label and the first predicted video content category based on the first loss function. Among them, the computer device may generate a second loss function according to the second video content category label and the second predicted video content category of the second sample video data, and calculate the error between the second video content category label and the second predicted video content category based on the second loss function.

[0207] Further, the computer device may generate a total loss function according to the first loss function and the second loss function, calculate the model loss of the initial classification model according to the total loss function, and adjust the parameters of the initial classification model according to the model loss. Among them, the computer device may detect whether the initial classification model after parameter adjustment meets the convergence condition. If the initial classification model after parameter adjustment meets the convergence condition, the initial classification model after parameter adjustment may be determined as the target classification model. If the initial classification model after parameter adjustment does not meet the convergence condition, continue to perform iterative training on the initial classification model after parameter adjustment until the initial classification model meets the convergence condition, and determine the initial classification model that meets the convergence condition as the target classification model. Among them, the convergence condition may refer to that the number of training times of the initial classification model reaches the target number of times, that is, one parameter adjustment of the initial classification model is one training, or the model loss of the initial classification model is less than or equal to the target loss value.

[0208] Optionally, the specific manner in which the computer device generates the total loss function based on the first loss function and the second loss function may include: performing pixel sampling on the first sample feature map sequence according to the first sample pixel sampling parameter to obtain a first sample pixel sampling feature map sequence. Invoking an information loss prediction model to perform loss prediction on the first sample pixel sampling feature map sequence and the target first sample feature map to obtain a first information loss probability corresponding to the first sample feature map sequence. Performing pixel sampling on the second sample feature map sequence according to the second sample pixel sampling parameter to obtain a second sample pixel sampling feature map sequence, invoking the information loss prediction model to perform loss prediction on the second sample pixel sampling feature map sequence and the target second sample feature map to obtain a second information loss probability corresponding to the second sample feature map sequence. Performing weighted processing on the first loss function according to the first information loss probability to obtain a weighted first loss function, performing weighted processing on the second loss function according to the second information loss probability to obtain a weighted second loss function. Summing the weighted first loss function and the weighted second loss function to obtain the total loss function.

[0209] Specifically, after feature sampling is performed on the first sample feature map sequence and the second sample feature map sequence, there will be varying degrees of information loss. To ensure the rationality of the initial classification model during training and avoid training collapse, the information loss prediction model can be used to predict the information loss degrees of the first sample feature map sequence and the second sample feature map sequence respectively. Based on the information loss degree, the loss function can be rationalized to improve the training efficiency of the initial classification model. Among them, the information loss prediction model can be based on the input sampled feature map input information loss probability, which is used to indicate the probability of obtaining the correct video content category of the corresponding video data based on the sampled feature map. For example, if it is necessary to predict the video action category of action video data, the information loss prediction model can refer to an action recognition model. By inputting the sampled feature map of the action video data into the action recognition model, the probability of obtaining the correct action category of the action video data based on the sampled feature map can be output, that is, this probability is used to indicate the probability of predicting the correct action category of the action video data based on the sampled feature map of the action video data. Specifically, the computer device can perform pixel sampling on each first sample feature map in the first sample feature map sequence according to the first sample pixel sampling parameters to obtain the first sample pixel sampling feature map corresponding to each first sample feature map. The first sample pixel sampling feature maps corresponding to each first sample feature map are arranged and combined in the corresponding position order to obtain the first sample pixel sampling feature map sequence. Among them, the corresponding position order is the position order of each first sample feature map in the first sample feature map sequence. The computer device can add the features with the same position order in the first sample pixel sampling feature map sequence and the target first sample feature map to obtain the total sampled feature map of the first sample feature map sequence. Among them, the part with sampling of 0 in the target first sample feature map is filled with 0 during addition, that is, if there is no sample feature map ranked second in the target first sample feature map, it is filled with 0 and added to the first sample pixel sampling feature map ranked second. The computer device can use the information loss prediction model to predict the information loss of the total sampled feature map of the first sample feature map sequence to obtain the first information loss probability. The first information loss probability can be used to indicate the probability of predicting the correct video content category of the first sample video data based on the total sampled feature map of the first sample feature map sequence.

[0210] As Figure 8 shown, Figure 8 Figure 1 is a schematic diagram of obtaining the information loss probability provided by an embodiment of the present application. As Figure 8As shown, the computer device may perform pixel sampling on the first sample feature map sequence 80a using the first sample pixel sampling parameters to obtain the first sample pixel sampling feature map sequence 80b. The computer device may perform temporal sampling on the first sample feature map sequence using the first sample temporal sampling parameters to obtain the target first sample feature map 80c. The computer device may add the feature maps with the same position order in the first sample pixel sampling feature sequence 80b and the target first sample feature map 80c to obtain the total sampling feature map 80d of the first sample feature map sequence 80a. The computer device may input the total sampling feature map 80d of the first sample feature map sequence 80a into the information loss prediction model 80e to perform loss prediction on the total sampling feature map 80d, obtaining the first information loss probability 80f corresponding to the first sample feature map sequence 80a.

[0211] Similarly, the computer device may perform pixel sampling on each second sample feature map in the second sample feature map sequence according to the second sample pixel sampling parameters to obtain the second sample pixel sampling feature map corresponding to each second sample feature map, and arrange and combine the second sample pixel sampling feature maps corresponding to each second sample feature map in the corresponding position order to obtain the second sample pixel sampling feature map sequence, where the corresponding position order is the position order of each second sample feature map in the second sample feature map sequence. The computer device may add the features with the same position order in the second sample pixel sampling feature map sequence and the target second sample feature map to obtain the total sampling feature map of the second sample feature map sequence. Among them, the part with a sampling of 0 in the target second sample feature map is filled with 0 during addition, that is, if there is no second sample feature map arranged in the second position in the target second sample feature map, it is filled with 0 and added to the second sample pixel sampling feature map arranged in the second position. The computer device may perform information loss prediction on the total sampling feature map of the second sample feature map sequence through the information loss prediction model to obtain the second information loss probability. This second information loss probability can be used to indicate the probability of predicting the correct video content category of the second sample video data based on the total sampling feature map of the second sample feature map sequence.

[0212] Optionally, the specific manner in which the computer device generates the total loss function based on the first loss function and the second loss function may further include: calling a feature loss prediction model, inputting the first sample feature map sequence and the target sample fusion feature map into the feature loss prediction model for the first loss prediction to obtain the first information loss probability corresponding to the first sample feature map sequence. Inputting the second sample feature map sequence and the target sample fusion feature map into the feature loss prediction model for the second loss prediction to obtain the second information loss probability corresponding to the second sample feature map sequence. Weighting the first loss function according to the first information loss probability to obtain the weighted first loss function, and weighting the second loss function according to the second information loss probability to obtain the weighted second loss function. Summing the weighted first loss function and the weighted second loss function to obtain the total loss function.

[0213] Specifically, after the feature fusion of the first sample feature map sequence and the second sample feature map sequence, there will be varying degrees of information loss. To ensure the rationality of the training of the initial classification model and avoid the situation of training collapse, the degree of information loss can be predicted according to the feature loss prediction model, and the loss function can be rationalized according to the degree of information loss to improve the training efficiency of the initial classification model. The computer device can call a feature loss prediction model, which is used to predict the information loss of the fused target sample fusion feature map to obtain the difference degree between the target sample fusion feature map and the first sample feature map sequence or the second sample feature map sequence before fusion. Among them, the feature loss prediction model can compare the first sample feature map sequence or the second sample feature map sequence before fusion with the fused target sample fusion feature map to determine the probability of the loss of key feature information in the first sample feature map sequence or the second sample feature map sequence. The computer device can input the first sample feature map sequence and the target sample fusion feature map into the feature loss prediction model to perform the first loss prediction on the target sample fusion feature map to obtain the first information loss probability of the first sample feature map sequence. It can be understood that the feature loss prediction model can predict how much useful information remains after the feature fusion of the first sample feature map sequence, and the first information loss probability can be used to indicate how much useful information of the first sample feature map sequence is still included in the target sample fusion feature map, that is, the first information loss probability can be used to indicate the probability of correctly predicting the video content category of the first sample video data based on the target sample fusion feature map.

[0214] Similarly, the computer device can input the second sample feature map sequence and the target sample fusion feature map into the feature loss prediction model to perform a second loss prediction on the target sample fusion feature map, and obtain the second information loss probability of the second sample feature map sequence. It can be understood that the computer device can predict, through the feature loss prediction model, how much useful information remains after the second sample feature map sequence undergoes feature fusion. This second information loss probability can be used to indicate how much useful information of the second sample feature map sequence is still included in the target sample fusion feature map, that is, this second information loss probability can be used to indicate the probability of correctly predicting the video content category of the second sample video data based on the target sample fusion feature map.

[0215] Furthermore, the computer device can perform a weighted processing on the first loss function using the first information loss probability to obtain a weighted first loss function, and perform a weighted processing on the second loss function using the first information loss probability to obtain a weighted second loss function. In this way, by performing a weighted processing on the first loss function using the first information loss probability and performing a weighted processing on the second loss function using the second information loss probability, it is possible to avoid the training collapse caused by information loss during feature fusion. It can be understood that when there is no useful information of the first sample feature map sequence left in the target sample fusion feature map, when classifying the first sample video data based on the target sample fusion feature map, the error between the obtained first predicted video content category and the first video content category label is much larger than the target error, which further leads to the initial classification model not reaching the convergence condition and resulting in training collapse. Furthermore, the computer device can perform a summation processing on the weighted first loss function and the weighted second loss function to obtain a total loss function. It can be seen that, based on the first information loss probability and the second information loss probability, performing a weighted processing on the first loss function and the second loss function can make the control of the model loss of the initial classification model more reasonable, and can avoid the training collapse caused by the loss of key information during feature fusion. At the same time, it can also accelerate the convergence speed of the initial classification model, improve the training efficiency of the initial classification model, and improve the accuracy of the target classification model obtained through training.

[0216] Among them, the computer formula of the total loss function can be shown as the following formula (2):

[0217] l mix =γ1×l CE (β1, fc1)+γ2×l CE (β2, fc2) (2)

[0218] Among them, γ1 in formula (2) refers to the first information loss probability of the first sample feature map sequence, l CE(β1, fc1) refers to the first loss function, β1 refers to the first video content category label, fc1 refers to the first predicted video content category, γ2 refers to the second information loss probability of the second sample feature map sequence, l CE (β2, fc2) refers to the second loss function, β2 refers to the second video content category label, fc2 refers to the second predicted video content category.

[0219] Optionally, the information loss prediction model can be pre-trained by a computer device and directly invoked when predicting the information loss of the target sample fusion feature map. The information loss prediction model does not participate in the parameter update of the initial classification model, that is, the parameters in the information loss prediction model do not need to be updated. It can be seen that in the embodiments of the present application, by performing information mixing enhancement on the first sample video data and the second sample video data in the time dimension and the pixel dimension, the data enhancement effect can be improved. At the same time, an information loss prediction model is used to predict the information loss of the fused target sample fusion feature map, and the loss function is weighted according to the information loss probability predicted by the information loss prediction model (that is, the useful information in the target sample fusion feature map is measured), avoiding the training collapse of the initial classification model and ensuring the rationality of the initial classification model training.

[0220] As Figure 9 shown, Figure 9 is a schematic diagram of an initial classification model training method provided by an embodiment of the present application. As Figure 9As shown in the figure, the computer device can obtain the first sample video data 90a and the second sample video data 90b. Among them, the first sample video data includes the first sample video frames T1, T2, T3, and T4, and the second sample video data includes the second sample video frames S1, S2, S3, and S4. The computer device can perform feature extraction processing on each first sample video frame in the first sample video data through the first feature extraction layer in the initial classification model to obtain the first sample feature map sequence 90c corresponding to the first sample video data. It can be understood that the computer device can perform feature extraction processing on the first sample video frames T1, T2, T3, and T4 respectively in the first feature extraction layer to obtain the first sample feature maps T1, T2, T3, and T4. Similarly, the computer device can perform feature extraction processing on each second sample video frame in the second sample video data through the second feature extraction layer in the initial classification model to obtain the second sample feature map sequence 90c corresponding to the second sample video data. It can be understood that the computer device can perform feature extraction processing on the second sample video frames S1, S2, S3, and S4 respectively in the second feature extraction layer to obtain the second sample feature maps S1, S2, S3, and S4.

[0221] Furthermore, as Figure 9 shown, the computer device can input the first sample feature map sequence 90c and the second sample feature map sequence 90d into the feature fusion layer 90e in the initial classification model to perform feature fusion on the first sample feature map sequence 90c and the second sample feature map sequence 90d, and obtain the target sample fusion feature map 90f. For the specific content, please refer to steps S102 - S104 in the above Figure 3 and steps S204 - S206 in the above Figure 5 . The embodiments of the present application will not elaborate here. The computer device can input the target sample fusion feature map 90f into the convolutional layer 90g in the initial classification model to perform convolutional processing on the target sample fusion feature map 90f to obtain the target sample fusion feature map after convolutional processing. For the specific content, please refer to the above Figure 5For the content of step S206 in the present application embodiment, it will not be elaborated here. The computer device may input the fused feature map of the target sample after convolution into the first classification layer 90h in the initial classification model to perform first feature classification on the fused feature map of the target sample after convolution, and obtain the first predicted video content category 90j of the first sample video data 90a. The computer device may input the fused feature map of the target sample after convolution into the second classification layer 90i in the initial classification model to perform second feature classification on the fused feature map of the target sample after convolution, and obtain the second predicted video content category 90k of the second sample video data 90b.

[0222] Further, the computer device may determine the first loss function 90m according to the first video content category label 90l and the first predicted video content category 90j of the first sample video data 90a, and determine the second loss function 90o according to the second video content category label 90n and the second predicted video content category 90k of the second sample video data 90b. Among them, the computer device may input the first sample pixel sampling feature map sequence of the first sample feature map sequence 90c and the target first sample feature map into the information loss prediction model 90p to predict and obtain the first information loss probability 90q of the first sample feature map sequence. The computer device may input the second sample pixel sampling feature map sequence of the second sample feature map sequence 90d and the target second sample feature map into the information loss prediction model 90p to output the second information loss probability 90r corresponding to the second sample feature map sequence 90d. The first loss function 90m is weighted by the first information loss probability 90q to obtain the weighted first loss function, and the second loss function 90o is weighted by the second information loss probability 90r to obtain the weighted second loss function. The computer device may perform a summation process on the weighted first loss function and the weighted second loss function to obtain the total loss function 90g, and may iteratively train the initial classification model according to the total loss function 90g to obtain the target classification model. For the specific content, reference may be made to the content of step S303 in the above Figure 7 and it will not be elaborated here.

[0223] In the embodiments of the present application, through an initial classification model, first feature extraction processing is performed on M first sample video frames in the first sample video data to obtain a first sample feature map sequence, and second feature extraction processing is performed on M second sample video frames in the second sample video data to obtain a second sample feature map sequence. By performing feature extraction on different sample video data, a first sample feature map sequence and a second sample feature map sequence are obtained. Sampling is performed on the first sample feature map sequence according to the first sample time sampling parameter to obtain a target first sample feature map, and sampling is performed on the second sample feature map sequence according to the second sample time sampling parameter to obtain a target second sample feature map. The sum of the number of sample video frames corresponding to the target first sample feature map and the target second sample feature map is equal to M, and the position of the target first sample feature map in the first sample feature map sequence is different from the position of the target second sample feature map in the second sample feature map sequence. A sample time fusion feature map is generated according to the target first sample feature map and the target second sample feature map. By using the timing information of each sample video frame, sampling is respectively performed on the first sample feature map sequence and the second sample feature map sequence, and feature fusion (i.e., feature enhancement) is performed on the sampled target first sample feature map and the target first sample feature map to obtain a sample time fusion feature map, thereby performing mutual feature enhancement on different video data according to the timing information.

[0224] A target sample fusion feature map for predicting the video content category is generated according to the sample time fusion feature map and the sample pixel fusion feature map. In this way, by performing hybrid feature enhancement on the first sample video data and the second sample video data in terms of the time dimension and the pixel dimension, the feature enhancement effect can be improved. The parameters of the initial classification model are adjusted according to the target sample fusion feature map to obtain a target classification model, which can improve the accuracy and robustness of the trained target classification model while integrating the initial classification model. In addition, this solution also uses an information loss prediction model to predict the information loss probability and weights the loss function of the initial classification model, so that the initial classification model can more reasonably control the sampled information during the training process. At the same time, it can also prevent the initial classification model from failing to meet the convergence condition and causing the training to crash, accelerate the convergence speed of the initial classification model, and improve the accuracy of the trained target classification model. It can be seen that the present application does not need to rely on text information such as video labels and video introductions of the target video data, nor does it need to rely on manual experience analysis, and can accurately classify the target video data.

[0225] Please refer to Figure 10 , Figure 10The following is a schematic structural diagram of a data processing device 1 provided by an embodiment of the present application. The above data processing device 1 may be a computer program (including program code) running in a computer device. For example, the data processing device 1 is an application software. The data processing device 1 may be used to execute corresponding steps in the data processing method provided by the embodiment of the present application. As Figure 10 shown, the data processing device 1 may include: a first feature extraction module 11, a first sampling module 12, a generation module 13, a classification module 14, an acquisition module 15, a first determination module 16, and a second determination module 17.

[0226] The first feature extraction module 11 is configured to obtain M video frames in the target video data, perform first feature extraction processing on the M video frames to obtain a first feature map sequence, and perform second feature extraction processing on the M video frames to obtain a second feature map sequence;

[0227] The first sampling module 12 is configured to sample the first feature map sequence according to a first time sampling parameter to obtain a target first feature map, and sample the second feature map sequence according to a second time sampling parameter to obtain a target second feature map; the sum of the number of video frames corresponding to the target first feature map and the target second feature map is equal to M, and the video frames corresponding to the target first feature map and the target second feature map are different from each other;

[0228] The generation module 13 is configured to generate a time fusion feature map according to the target first feature map and the target second feature map;

[0229] The classification module 14 is configured to generate a target fusion feature map according to the time fusion feature map, the first feature map sequence, and the second feature map sequence, and classify the target fusion feature map to obtain the video content category of the target video data.

[0230] Among them, the first feature map sequence includes first feature maps corresponding to M video frames respectively, and the second feature map sequence includes second feature maps corresponding to M video frames respectively;

[0231] The classification module 14 includes:

[0232] The pixel mixing and splicing unit 1401 is configured to perform pixel mixing and splicing on the first feature map and the second feature map associated with the same video frame in the first feature map sequence and the second feature map sequence to obtain pixel mixing feature maps corresponding to M video frames respectively;

[0233] The first generation unit 1402 is configured to generate a pixel fusion feature map according to the pixel mixing feature maps corresponding to M video frames respectively;

[0234] A classification unit 1403 is used to generate a target fusion feature map based on a temporal fusion feature map and a pixel fusion feature map, classify the target fusion feature map, and obtain the video content category of the target video data.

[0235] Among them, the first feature extraction module 11 includes:

[0236] A first acquisition unit 1101 is used to acquire the original video data and acquire the content attributes of each original video frame in the original video data.

[0237] A partitioning unit 1102 is used to partition the original video data according to the content attributes of each original video frame to obtain N video segments; N is a positive integer.

[0238] A selection unit 1103 is used to select a target video segment from the N video segments as the target video data.

[0239] A video frame sampling unit 1104 is used to perform video frame sampling on the original video frames included in the target video data according to the number M of sampled video frames indicated by the video sampling rule to obtain M video frames in the target video data.

[0240] Among them, the data processing device 1 further includes:

[0241] An acquisition module 15 is used to acquire an initial time sampling parameter; the number of sampling elements in the initial time sampling parameter is M.

[0242] A first determination module 16 is used to randomly determine the element values of M sampling elements with a position order in the initial time sampling parameter to obtain a first time sampling parameter; the element values include a first element threshold for indicating sampling of a feature map and a second element threshold for indicating masking of a feature map.

[0243] A second determination module 17 is used to determine a second time sampling parameter according to the element values of the M sampling elements in the first time sampling parameter; the element values of the sampling elements in the same position in the first time sampling parameter and the second time sampling parameter are different from each other.

[0244] Among them, the first sampling module 12 includes:

[0245] A first sampling unit 1201 is used to call a target classification model, sample the associated feature maps in the first feature map sequence based on the first element threshold in the first time sampling parameter, and mask the associated feature maps in the first feature map sequence based on the second element threshold in the first time sampling parameter to obtain a target first feature map.

[0246] The second sampling unit 1202 is configured to sample the associated feature maps in the second feature map sequence according to the first element threshold in the second time sampling parameter, and mask the associated feature maps in the second feature map sequence according to the second element threshold in the second time sampling parameter, so as to obtain a target second feature map.

[0247] Among them, the generation module 13 includes:

[0248] The second acquisition unit 1301 is configured to acquire the first timestamp of the video frame corresponding to the target first feature map, and acquire the second timestamp of the video frame corresponding to the target second feature map;

[0249] The combination unit 1302 is configured to combine the target first feature map and the target second feature map according to the time order between the first timestamp and the second timestamp, so as to obtain a time fusion feature map.

[0250] Among them, the M video frames include video frame M i , where i is a positive integer less than or equal to M;

[0251] The pixel mixing and stitching unit 1401 is specifically configured to:

[0252] Call the target classification model, and obtain the first feature map corresponding to video frame M in the first feature map sequence through the feature fusion layer in the target classification model, and obtain the second feature map corresponding to video frame M in the second feature map sequence i ; i Corresponding second feature map;

[0253] According to the first pixel sampling parameter, perform pixel sampling on the first feature map corresponding to video frame M i To obtain a first pixel sampling feature map, and perform pixel sampling on the second feature map corresponding to video frame M according to the second pixel sampling parameter i To obtain a second pixel sampling feature map;

[0254] Perform pixel mixing and stitching on the first pixel sampling feature map and the second pixel sampling feature map to obtain the pixel mixing feature map corresponding to video frame M i Corresponding pixel mixing feature map.

[0255] Among them, the classification unit 1403 is specifically configured to:

[0256] Call the target classification model, and add the time fusion feature map and the pixel fusion feature map through the feature fusion layer in the target classification model to obtain a target fusion feature map;

[0257] Perform convolution processing on the target fusion feature map through the convolution layer in the target classification model to obtain a convolution-processed target fusion feature map;

[0258] Through the classification layer in the target classification model, the target fusion feature map after convolution processing is classified to obtain the video content category of the target video data.

[0259] Among them, the classification unit 1403 is also specifically used to include:

[0260] Input the target fusion feature map after convolution processing into the first classification sub-layer in the classification layer, classify the target fusion feature map after convolution processing, and obtain the first classification result;

[0261] Input the target fusion feature map after convolution processing into the second classification sub-layer in the classification layer, classify the target fusion feature map after convolution processing, and obtain the second classification result;

[0262] Obtain the average value of the first classification result and the second classification result, and determine the video content category of the target video data according to this average value.

[0263] According to an embodiment of the present application, Figure 3 The steps involved in the data processing method shown can be executed by Figure 10 each module in the data processing device 1 shown. For example, Figure 3 the step S101 shown in Figure 10 can be executed by the first feature extraction module 11 in Figure 3 the step S102 shown in Figure 10 can be executed by the first sampling module 12 in Figure 3 the step S103 shown in Figure 10 can be executed by the generation module 13 in Figure 3 the step S104 shown in Figure 10 can be executed by the classification module 14 in, and so on.

[0264] According to an embodiment of the present application, Figure 10 Each module in the data processing device 1 shown can be separately or all combined into one or several units to form, or some of the units can be further split into multiple smaller sub-units in terms of function, and the same operations can be achieved without affecting the realization of the technical effects of the embodiments of the present application. The above modules are divided based on logical functions. In actual applications, the function of one module can also be realized by multiple units, or the functions of multiple modules can be realized by one unit. In other embodiments of the present application, the test device can also include other units. In actual applications, these functions can also be assisted by other units and can be realized by the cooperation of multiple units.

[0265] In the embodiments of the present application, M video frames in the target video data are obtained through a video sampling rule, and the target video data is preprocessed. While ensuring a reduction in the amount of computation, it is also possible to avoid the loss of key information in the target video data caused by sampling, and the efficiency of subsequent classification of the target video data can be improved. Further, the M video frames are subjected to a first feature extraction process to obtain a first sequence of feature maps, and the M video frames are subjected to a second feature extraction process to obtain a second sequence of feature maps. By performing the first feature extraction and the second feature extraction on the M video frames, the first sequence of feature maps and the second sequence of feature maps of the M video frames can be obtained, and different feature information of the M video frames can be extracted from different perspectives. Further, the first sequence of feature maps is sampled according to a first time sampling parameter to obtain a target first feature map, the second sequence of feature maps is sampled according to a second time sampling parameter to obtain a target second feature map, and a temporal fusion feature map is generated based on the target first feature map and the target second feature map. It can be seen that by sampling the first sequence of feature maps and the second sequence of feature maps respectively in the time dimension, a temporal fusion feature map is obtained, and based on this, the feature enhancement of the target video data is performed according to the temporal information between each video frame, improving the feature enhancement effect of the target video data. Further, the first sequence of feature maps and the second sequence of feature maps are pixel-sampled to obtain a pixel fusion feature map, and a target fusion feature map is generated based on the temporal fusion feature map in the time dimension and the pixel fusion feature map in the pixel dimension, and the target fusion feature map is classified to obtain the video content category of the target video data. It can be seen that the present application performs feature enhancement on the target video data in the time dimension and the pixel dimension, and samples the target fusion feature map after feature enhancement to classify the target video data, which can improve the accuracy of classifying the target video data. It can be seen that the present application does not need to rely on text information such as video tags and video introductions of the target video data, nor does it need to rely on manual experience analysis, and can accurately classify the target video data.

[0266] Please refer to Figure 11 , Figure 11 which is a schematic structural diagram of a data processing device 2 provided in the embodiments of the present application. The above data processing device 2 may be a computer program (including program code) running in a computer device. For example, the data processing device 2 is an application software; the data processing device 2 may be used to execute the corresponding steps in the data processing method provided in the embodiments of the present application. As Figure 11 shown, the data processing device 2 may include: a second feature extraction module 21, a second sampling module 22, and a parameter adjustment module 23.

[0267] The second feature extraction module 21 is configured to perform first feature extraction processing on M first sample video frames in the first sample video data through an initial classification model to obtain a first sample feature map sequence, and perform second feature extraction processing on M second sample video frames in the second sample video data to obtain a second sample feature map sequence; M is a positive integer;

[0268] The second sampling module 22 is configured to sample the first sample feature map sequence according to the first sample time sampling parameter to obtain a target first sample feature map, and sample the second sample feature map sequence according to the second sample time sampling parameter to obtain a target second sample feature map; the sum of the number of sample video frames corresponding to the target first sample feature map and the target second sample feature map is equal to M, and the position of the target first sample feature map in the first sample feature map sequence is different from the position of the target second sample feature map in the second sample feature map sequence;

[0269] The parameter adjustment module 23 is configured to generate a sample time fusion feature map according to the target first sample feature map and the target second sample feature map, generate a target sample fusion feature map for predicting the video content category according to the sample time fusion feature map, the first sample feature map sequence, and the second sample feature map sequence, and adjust the parameters of the initial classification model according to the target sample fusion feature map to obtain a target classification model; the target classification model is used to predict the video content category of the target video data.

[0270] The parameter adjustment module 23 includes:

[0271] The prediction unit 2301 is configured to predict a first predicted video content category of the first sample video data according to the target sample fusion feature map, and predict a second predicted video content category of the second sample video data according to the target sample fusion feature map;

[0272] The second generation unit 2302 is configured to generate a first loss function according to the first video content category label and the first predicted video content category of the first sample video data;

[0273] The third generation unit 2303 is configured to generate a second loss function according to the second video content category label and the second predicted video content category of the second sample video data;

[0274] The determination unit 2304 is configured to generate a total loss function according to the first loss function and the second loss function, adjust the parameters of the initial classification model according to the total loss function, and when the initial classification model after parameter adjustment meets the training convergence condition, determine the initial classification model after parameter adjustment as the target classification model.

[0275] Wherein, the determination unit 2304 is specifically configured to:

[0276] Perform pixel sampling on the first sample feature map sequence according to the first sample pixel sampling parameters to obtain the first sample pixel sampling feature map sequence. Invoke the information loss prediction model to perform loss prediction on the first sample pixel sampling feature map sequence and the target first sample feature map, and obtain the first information loss probability corresponding to the first sample feature map sequence;

[0277] Perform pixel sampling on the second sample feature map sequence according to the second sample pixel sampling parameters to obtain the second sample pixel sampling feature map sequence. Invoke the information loss prediction model to perform loss prediction on the second sample pixel sampling feature map sequence and the target second sample feature map, and obtain the second information loss probability corresponding to the second sample feature map sequence;

[0278] Perform weighted processing on the first loss function according to the first information loss probability to obtain the weighted first loss function. Perform weighted processing on the second loss function according to the second information loss probability to obtain the weighted second loss function;

[0279] Perform summation processing on the weighted first loss function and the weighted second loss function to obtain the total loss function.

[0280] According to an embodiment of the present application, Figure 7 the steps involved in the data processing method shown can be performed by Figure 11 each module in the data processing device 2 shown. For example, Figure 7 the step S301 shown in can be performed by Figure 11 the second feature extraction module 21 in, Figure 7 the step S302 shown in can be performed by Figure 11 the second sampling module 22 in, Figure 7 the step S303 shown in can be performed by Figure 11 the parameter adjustment module 23 in, and so on. The second feature extraction module 21, the second sampling module 22, and the parameter adjustment module 23.

[0281] According to an embodiment of the present application, Figure 11 each module in the data processing device 2 shown can be separately or all combined into one or several units to form, or a certain one (or some) of the units can be further split into multiple smaller sub-units in terms of function, and the same operations can be achieved without affecting the realization of the technical effects of the embodiments of the present application. The above modules are divided based on logical functions. In practical applications, the function of one module can also be realized by multiple units, or the functions of multiple modules can be realized by one unit. In other embodiments of the present application, the test device can also include other units. In practical applications, these functions can also be assisted by other units and can be realized by the cooperation of multiple units.

[0282] In the embodiment of the present application, through the initial classification model, the first feature extraction process is performed on M first sample video frames in the first sample video data to obtain the first sample feature map sequence, and the second feature extraction process is performed on M second sample video frames in the second sample video data to obtain the second sample feature map sequence. By performing feature extraction on different sample video data, the first sample feature map sequence and the second sample feature map sequence are obtained. The first sample feature map sequence is sampled according to the first sample time sampling parameter to obtain the target first sample feature map, and the second sample feature map sequence is sampled according to the second sample time sampling parameter to obtain the target second sample feature map. The sum of the number of sample video frames corresponding to the target first sample feature map and the target second sample feature map is equal to M, and the position of the target first sample feature map in the first sample feature map sequence is different from the position of the target second sample feature map in the second sample feature map sequence. The sample time fusion feature map is generated according to the target first sample feature map and the target second sample feature map. By using the timing information of each sample video frame, the first sample feature map sequence and the second sample feature map sequence are respectively sampled, and the sampled target first sample feature map and the target first sample feature map are subjected to feature fusion (i.e., feature enhancement) to obtain the sample time fusion feature map, so as to perform mutual feature enhancement on different video data according to the timing information.

[0283] According to the sample time fusion feature map and the sample pixel fusion feature map, the target sample fusion feature map for predicting the video content category is generated. In this way, by performing hybrid feature enhancement on the first sample video data and the second sample video data in terms of the time dimension and the pixel dimension, the feature enhancement effect can be improved. According to the target sample fusion feature map, the parameters of the initial classification model are adjusted to obtain the target classification model, which can improve the accuracy and robustness of the trained target classification model while integrating the initial classification model. In addition, this solution also uses an information loss prediction model to predict the information loss probability and weights the loss function of the initial classification model, so that the initial classification model can more reasonably control the sampled information during the training process. At the same time, it can also avoid the training collapse caused by the initial classification model not meeting the convergence condition, accelerate the convergence speed of the initial classification model, and improve the accuracy of the trained target classification model. It can be seen that the present application does not need to rely on text information such as video labels and video introductions of the target video data, nor does it need to rely on manual experience analysis, and can accurately classify the target video data.

[0284] Please refer to Figure 12 , Figure 12 is a schematic structural diagram of a computer device provided by an embodiment of the present application. As Figure 12As shown in the figure, the above computer device 1000 may include: a processor 1001, a network interface 1004, and a memory 1005. In addition, the above computer device 1000 may further include: a user interface 1003 and at least one communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. Among them, the user interface 1003 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. Optionally, the memory 1005 may also be at least one storage device located far from the aforementioned processor 1001. As Figure 12 shown, the memory 1005, as a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.

[0285] In Figure 12 the computer device 1000 shown in the figure, the network interface 1004 can provide network communication functions; while the user interface 1003 is mainly used to provide an input interface for the target user; and the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement the description of the data processing method in the corresponding embodiment described above Figure 3 and can also execute the description of the data processing device 1 in the corresponding embodiment described above Figure 10 and will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either.

[0286] In addition, the computer device 1000 described in the embodiments of the present application can also execute the description of the data processing method in the corresponding embodiment described above Figure 7 and can also execute the description of the data processing device 2 in the corresponding embodiment described above Figure 11 and will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either.

[0287] In addition, it should be pointed out here that: the embodiments of the present application also provide a computer-readable storage medium, and the computer-readable storage medium stores the computer programs executed by the aforementioned data processing device 1 and data processing device 2, and the computer programs include program instructions. When the processor executes the program instructions, it can execute the above Figure 3 , Figure 5 and Figure 7The description of the data processing method in the corresponding embodiments will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in this application, please refer to the description of the method embodiments of this application. As an example, the program instructions can be deployed to be executed on a computing device, or on multiple computing devices located at one location, or on multiple computing devices distributed at multiple locations and interconnected by a communication network. The multiple computing devices distributed at multiple locations and interconnected by a communication network can form a blockchain system.

[0288] In addition, it should be noted that: The embodiments of the present application also provide a computer program product or a computer program. The computer program product or the computer program may include computer instructions, and the computer instructions may be stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor may execute the computer instructions, so that the computer device executes the description of the data processing method in the foregoing Figure 3 , Figure 5 and Figure 7 corresponding embodiments. Therefore, it will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated. For the technical details not disclosed in the embodiments of the computer program product or the computer program involved in this application, please refer to the description of the method embodiments of this application.

[0289] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0290] The steps in the method embodiments of this application can be adjusted, combined, and deleted according to actual needs.

[0291] The modules in the device embodiments of this application can be combined, divided, and deleted according to actual needs.

[0292] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0293] The above-disclosed are only the preferred embodiments of the present application. Of course, the scope of the rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.

Claims

1. A data processing method, characterized in that, Including: Obtain M video frames in the target video data, perform first feature extraction processing on the M video frames to obtain a first feature map sequence, and perform second feature extraction processing on the M video frames to obtain a second feature map sequence; Obtain an initial time sampling parameter; the number of sampling elements in the initial time sampling parameter is M; Randomly determine the element values of M sampling elements with a position order in the initial time sampling parameter to obtain a first time sampling parameter; the element values include a first element threshold for indicating sampling of a feature map and a second element threshold for indicating masking of a feature map; Determine a second time sampling parameter according to the element values of the M sampling elements in the first time sampling parameter; the element values of the sampling elements in the same position in the first time sampling parameter and the second time sampling parameter are different from each other; Sample the first feature map sequence according to the first time sampling parameter to obtain a target first feature map, and sample the second feature map sequence according to the second time sampling parameter to obtain a target second feature map; the sum of the numbers of video frames corresponding to the target first feature map and the target second feature map is equal to M, and the video frames corresponding to the target first feature map and the target second feature map are different from each other; Generate a time fusion feature map according to the target first feature map and the target second feature map; Generate a target fusion feature map according to the time fusion feature map, the first feature map sequence, and the second feature map sequence, and classify the target fusion feature map to obtain the video content category of the target video data.

2. The method according to claim 1, wherein The first feature map sequence includes first feature maps corresponding to the M video frames respectively, and the second feature map sequence includes second feature maps corresponding to the M video frames respectively; The generating a target fusion feature map according to the time fusion feature map, the first feature map sequence, and the second feature map sequence, and classifying the target fusion feature map to obtain the video content category of the target video data includes: In the first feature map sequence and the second feature map sequence, perform pixel mixing and splicing on the first feature map and the second feature map associated with the same video frame to obtain pixel mixing feature maps corresponding to the M video frames respectively; Generate a pixel fusion feature map according to the pixel mixing feature maps corresponding to the M video frames respectively; Generate a target fusion feature map according to the time fusion feature map and the pixel fusion feature map, and classify the target fusion feature map to obtain the video content category of the target video data.

3. The method according to claim 1, wherein The obtaining M video frames in the target video data includes: Obtain the original video data, and obtain the content attributes of each original video frame in the original video data; Divide the original video data according to the content attributes of each original video frame to obtain N video segments; N is a positive integer; Select a target video segment from the N video segments as the target video data; Perform video frame sampling on the original video frames included in the target video data according to the number of sampled video frames M indicated by the video sampling rule, to obtain M video frames in the target video data.

4. The method according to claim 1, characterized in that, The sampling the first feature map sequence according to the first time sampling parameter to obtain a target first feature map, and sampling the second feature map sequence according to the second time sampling parameter to obtain a target second feature map includes: Invoke a target classification model, and in the feature fusion layer of the target classification model, sample the associated feature maps in the first feature map sequence based on the first element threshold in the first time sampling parameter, and mask the associated feature maps in the first feature map sequence based on the second element threshold in the first time sampling parameter, to obtain a target first feature map; Sample the associated feature maps in the second feature map sequence according to the first element threshold in the second time sampling parameter, and mask the associated feature maps in the second feature map sequence according to the second element threshold in the second time sampling parameter, to obtain a target second feature map.

5. The method according to claim 1, characterized in that The generating a temporal fusion feature map according to the target first feature map and the target second feature map includes: Obtain the first timestamp of the video frame corresponding to the target first feature map, and obtain the second timestamp of the video frame corresponding to the target second feature map; Combine the target first feature map and the target second feature map according to the time order between the first timestamp and the second timestamp, to obtain a temporal fusion feature map.

6. The method according to claim 2, wherein The M video frames include video frame M i , where i is a positive integer less than or equal to M; In the first feature map sequence and the second feature map sequence, pixel-blend and splice the first feature map and the second feature map associated with the same video frame, to obtain pixel-blended feature maps corresponding to the M video frames, includes: Call the target classification model, and obtain video frame M in the first feature map sequence through the feature fusion layer in the target classification model i The corresponding first feature map, and obtain the video frame M in the second feature map sequence i The corresponding second feature map; Perform pixel sampling on the first feature map corresponding to the video frame M according to the first pixel sampling parameter to obtain a first pixel sampling feature map, and perform pixel sampling on the second feature map corresponding to the video frame M according to the second pixel sampling parameter to obtain a second pixel sampling feature map; i i ​​ Perform pixel mixing and stitching on the first pixel sampling feature map and the second pixel sampling feature map to obtain the video frame M i The corresponding pixel mixing feature map.

7. The method according to claim 2, wherein The generating a target fusion feature map according to the temporal fusion feature map and the pixel fusion feature map, and classifying the target fusion feature map to obtain the video content category of the target video data includes: Invoke a target classification model, and through the feature fusion layer in the target classification model, add the temporal fusion feature map and the pixel fusion feature map to obtain a target fusion feature map; Perform convolution processing on the target fusion feature map through the convolution layer in the target classification model, to obtain a target fusion feature map after convolution processing; Perform classification processing on the target fusion feature map after convolution processing through the classification layer in the target classification model, to obtain the video content category of the target video data.

8. The method according to claim 7, characterized in that, The performing classification processing on the target fusion feature map after convolution processing through the classification layer in the target classification model to obtain the video content category of the target video data includes: Input the target fusion feature map after convolution processing into the first classification sub-layer in the classification layer, and perform classification processing on the target fusion feature map after convolution processing to obtain a first classification result; Input the target fusion feature map after convolution processing into the second classification sublayer in the classification layer, perform classification processing on the target fusion feature map after convolution processing, and obtain a second classification result; Obtain the average value of the first classification result and the second classification result, and determine the video content category of the target video data according to this average value.

9. A data processing method, characterized in that Including: Through an initial classification model, perform first feature extraction processing on M first sample video frames in the first sample video data to obtain a first sample feature map sequence, and perform second feature extraction processing on M second sample video frames in the second sample video data to obtain a second sample feature map sequence; M is a positive integer; Obtain an initial time sampling parameter; the number of sampling elements in the initial time sampling parameter is M; Randomly determine the element values of M sampling elements with a position order in the initial time sampling parameter to obtain a first sample time sampling parameter; the element values include a first element threshold for indicating sampling of the feature map and a second element threshold for indicating masking of the feature map; Determine a second sample time sampling parameter according to the element values of M sampling elements in the first sample time sampling parameter; the element values of the sampling elements in the same position in the first sample time sampling parameter and the second sample time sampling parameter are different from each other; Sample the first sample feature map sequence according to the first sample time sampling parameter to obtain a target first sample feature map, and sample the second sample feature map sequence according to the second sample time sampling parameter to obtain a target second sample feature map; the sum of the number of sample video frames corresponding to the target first sample feature map and the target second sample feature map is equal to M, and the position of the target first sample feature map in the first sample feature map sequence is different from the position of the target second sample feature map in the second sample feature map sequence; Generate a sample time fusion feature map according to the target first sample feature map and the target second sample feature map, generate a target sample fusion feature map for predicting the video content category according to the sample time fusion feature map, the first sample feature map sequence, and the second sample feature map sequence, and adjust the parameters of the initial classification model according to the target sample fusion feature map to obtain a target classification model; the target classification model is used to predict the video content category of the target video data.

10. The method according to claim 9, wherein The adjusting the parameters of the initial classification model according to the target sample fusion feature map to obtain a target classification model includes: Predict a first predicted video content category of the first sample video data according to the target sample fusion feature map, and predict a second predicted video content category of the second sample video data according to the target sample fusion feature map; Generate a first loss function according to the first video content category label of the first sample video data and the first predicted video content category; Generate a second loss function according to the second video content category label of the second sample video data and the second predicted video content category; Generate a total loss function according to the first loss function and the second loss function, and adjust the parameters of the initial classification model according to the total loss function. When the initial classification model after parameter adjustment meets the training convergence condition, determine the initial classification model after parameter adjustment as the target classification model.

11. The method according to claim 10, wherein The generating the total loss function according to the first loss function and the second loss function includes: Perform pixel sampling on the first sample feature map sequence according to the first sample pixel sampling parameter to obtain a first sample pixel sampling feature map sequence, call the information loss prediction model, perform loss prediction on the first sample pixel sampling feature map sequence and the target first sample feature map, and obtain the first information loss probability corresponding to the first sample feature map sequence; Perform pixel sampling on the second sample feature map sequence according to the second sample pixel sampling parameter to obtain a second sample pixel sampling feature map sequence, call the information loss prediction model, perform loss prediction on the second sample pixel sampling feature map sequence and the target second sample feature map, and obtain the second information loss probability corresponding to the second sample feature map sequence; Perform weighted processing on the first loss function according to the first information loss probability to obtain a weighted first loss function, and perform weighted processing on the second loss function according to the second information loss probability to obtain a weighted second loss function; Perform summation processing on the weighted first loss function and the weighted second loss function to obtain a total loss function.

12. A data processing device, characterized in that, It includes: A feature extraction processing module, configured to obtain M video frames in the target video data, perform first feature extraction processing on the M video frames to obtain a first feature map sequence, and perform second feature extraction processing on the M video frames to obtain a second feature map sequence; An acquisition module, configured to acquire an initial time sampling parameter; the number of sampling elements in the initial time sampling parameter is M; A first determination module, configured to randomly determine the element values of M sampling elements with a position order in the initial time sampling parameter to obtain a first time sampling parameter; The element values include a first element threshold and a second element threshold. The first element threshold is used to indicate sampling of the feature map, and the second element threshold is used to indicate masking of the feature map; A second determination module, configured to determine a second time sampling parameter according to the element values of M sampling elements in the first time sampling parameter; the element values of the sampling elements in the same position in the first time sampling parameter and the second time sampling parameter are different from each other; A sampling module, configured to sample the first feature map sequence according to the first time sampling parameter to obtain a target first feature map, and sample the second feature map sequence according to the second time sampling parameter to obtain a target second feature map; the sum of the number of video frames corresponding to the target first feature map and the target second feature map is equal to M, and the video frames corresponding to the target first feature map and the target second feature map are different from each other; A generation module, configured to generate a time fusion feature map according to the target first feature map and the target second feature map; A classification module, configured to generate a target fusion feature map according to the time fusion feature map, the first feature map sequence, and the second feature map sequence, and classify the target fusion feature map to obtain the video content category of the target video data.

13. A computer device, characterized in that, Comprising: A processor and a memory; The processor is connected to the memory. Wherein, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method according to any one of claims 1-11.

14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and the computer program is adapted to be loaded and executed by a processor so that a computer device having the processor executes the method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Video classification method and device, electronic equipment and storage medium

    CN113010735A

  • Video classification method, device, electronic equipment and storage medium

    CN113010736A

  • Automatic and intelligent video sorting

    US20180336931A1