Video processing network training method, device, equipment and readable storage medium
Through multiple rounds of iterative training and feature similarity adjustment methods, the feature extraction accuracy of video processing networks is improved, and the problem of low feature extraction accuracy in the prior art is solved.
Patent Information
- Application Number
- CN202110218759.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-26
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-02-26
AI Technical Summary
In the prior art, video processing networks based on supervised learning pay too much attention to pixel-level details during feature extraction, resulting in low accuracy of feature extraction.
The first and second networks are trained by multiple rounds of iterative training methods, and parameter adjustments are made based on the similarity and contribution values of candidate reference features and historical reference features to improve the feature extraction accuracy of the video processing network.
The accuracy of video processing network extraction of video feature is improved, more video timing information can be retained, and feature representation is independent of auxiliary tasks.
Smart Images

Figure CN113705291B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a training method, apparatus, device and readable storage medium for a video processing network. Background Art
[0002] With the rapid development of artificial intelligence and machine learning technologies, rich feature extraction and feature representation of videos can be performed based on trained neural networks. Among them, the above neural networks are often trained based on supervised learning or unsupervised learning. However, due to the low richness of information provided by video samples in supervised learning and the excessive focus on pixel-level details of the video in supervised learning, the accuracy of feature extraction of videos by the trained neural networks is low. Therefore, how to improve the accuracy of feature extraction of videos based on neural networks is an issue that needs to be considered. Summary of the invention
[0003] The embodiments of the present application provide a video processing network training method, apparatus, device and readable storage medium for improving the accuracy of feature extraction from videos.
[0004] In a first aspect of the present application, a method for training a video processing network is provided, comprising:
[0005] Based on the training video, the first network and the second network are subjected to multiple rounds of iterative training, and the first network outputted by the last round of iterative training is determined as the target video processing network; the first network and the second network are networks with the same network structure and different network parameters, wherein one round of iterative training includes:
[0006] Inputting a basic video obtained based on the training video into the current first network to obtain corresponding basic features, and inputting a candidate reference video obtained based on the training video into the current second network to obtain corresponding candidate reference features;
[0007] Based on the candidate reference features and at least one historical reference feature, a target reference feature set is determined, and according to the similarity between each target reference feature in the target reference feature set and the basic feature and the corresponding contribution value, the first influence value corresponding to each target reference feature is determined, wherein the contribution value represents the influence of the corresponding target reference feature on the one round of iterative training, and the contribution value is obtained based on the extraction time of the corresponding target reference feature; when the one round of iterative training is the first round of iterative training in the multiple rounds of iterative training, the historical reference feature is obtained based on the current second network, and when the one round of iterative training is not the first round of iterative training in the multiple rounds of iterative training, the historical reference feature is obtained based on the historical second network;
[0008] According to the first influence values corresponding to the respective target reference features, the parameters of the current first network are adjusted to obtain the first network output by the one-round iterative training; and based on the first network output by the one-round iterative training, the parameters of the current second network are adjusted.
[0009] In a possible implementation, the reference feature sequence is obtained by sorting the target reference features in order of extraction time from earliest to latest, and a contribution value corresponding to one target reference feature is positively correlated with the arrangement position; or
[0010] The reference feature sequence is obtained by sorting the target reference features in order of extraction time from late to early, and the contribution value corresponding to one target reference feature is negatively correlated with the arrangement position.
[0011] In a second aspect of the present application, a training device for a video processing network is provided, comprising:
[0012] A network training unit, the network training unit is used to: perform multiple rounds of iterative training on the first network and the second network based on the training video, and determine the first network output from the last round of iterative training as the target video processing network; the first network and the second network are networks with the same network structure and different network parameters, and the network training unit includes a feature extraction subunit, an influence value determination subunit and a network adjustment subunit, wherein:
[0013] The feature extraction subunit is used to input the basic video obtained based on the training video into the current first network in one round of iterative training to obtain the corresponding basic features, and input the candidate reference video obtained based on the training video into the current second network to obtain the corresponding candidate reference features;
[0014] The influence value determination subunit is used to determine a target reference feature set based on the candidate reference features and at least one historical reference feature in a round of iterative training, and determine the first influence value corresponding to each target reference feature in the target reference feature set according to the similarity between each target reference feature and the basic feature and the corresponding contribution value, wherein the contribution value represents the influence degree of the corresponding target reference feature on the round of iterative training, and the contribution value is obtained based on the extraction time of the corresponding target reference feature; when the round of iterative training is the first round of iterative training in the multiple rounds of iterative training, the historical reference feature is obtained based on the current second network, and when the round of iterative training is not the first round of iterative training in the multiple rounds of iterative training, the historical reference feature is obtained based on the historical second network;
[0015] The network adjustment subunit is used to adjust the parameters of the current first network in a round of iterative training according to the first influence values corresponding to each of the target reference features to obtain the first network output by the round of iterative training; and adjust the parameters of the current second network based on the first network output by the round of iterative training.
[0016] In a possible implementation, a contribution value corresponding to a first target reference feature in the target reference feature set is smaller than a contribution value corresponding to a second target reference feature, wherein an extraction time of the first target reference feature is earlier than an extraction time of the second target reference feature.
[0017] In a possible implementation, the base video is obtained by performing a first data augmentation process on a target training video, and the candidate reference video is obtained by performing a second data augmentation process on the training video; wherein the target training video is obtained by performing interference processing on the training video based on a third network; or
[0018] The basic video is obtained by performing the first data augmentation process on the training video, and the candidate reference video is obtained by performing the second data augmentation process on the target training video; or
[0019] The basic video is obtained by performing the first data augmentation process on the target training video, and the candidate reference video is obtained by performing the second data augmentation process on the target training video.
[0020] In one possible implementation, the feature extraction subunit is further used to: before performing interference processing on the training video based on the third network, determine whether the model loss value of the current first network satisfies a preset loss value stability condition, and the model loss value represents the degree of loss of the current first network in extracting features from the video.
[0021] In a possible implementation manner, the network adjustment subunit is further used to: perform parameter adjustment on the third network according to the first influence values corresponding to the respective target reference features.
[0022] In a possible implementation manner, the network adjustment subunit is specifically configured to:
[0023] Determine the sum of the first influence values corresponding to each of the target reference features as the negative sample influence value;
[0024] Determine a model loss value of the current first network based on the negative sample influence value and the first influence value corresponding to the target reference feature; the model loss value is negatively correlated with the negative sample influence value, and the model loss value is positively correlated with the first influence value corresponding to the target reference feature;
[0025] Based on the model loss value, adjusting parameters of the current first network;
[0026] Based on the model loss value, parameters of the third network are adjusted.
[0027] In a possible implementation, the influence value determination subunit is specifically used to: perform the following operations for each target reference feature in the target reference feature set: determine a contribution value corresponding to a target reference feature in the target reference feature set based on an extraction time of the target reference feature; determine the similarity between the target reference feature and the basic feature; and perform weighted processing on the determined similarity based on the contribution value corresponding to the target reference feature to obtain a first influence value corresponding to the target reference feature.
[0028] In a possible implementation, the influence value determination subunit is specifically used to: determine the contribution value corresponding to the target reference feature based on the arrangement position of the target reference feature in the reference feature sequence; wherein the reference feature sequence is obtained by sorting the target reference features based on the extraction time of the target reference features.
[0029] In a possible implementation manner, the first network is created based on an R(2+1)D encoder or an S3D-G encoder.
[0030] In a possible implementation, the interference processing includes one or any combination of the following: deleting the first type of video frames in the training video; adjusting the positions of the second type of video frames in the training video; and modifying the color information of the third type of video frames in the training video.
[0031] In a possible implementation, the feature extraction subunit is further used to: before deleting the first category of video frames in the training video, based on the feature extraction subnetwork in the third network, extract the image features of each of the video frames; input the image features of each of the video frames into the long short-term memory subnetwork in the third network to determine the importance corresponding to each of the video frames, wherein the importance represents the degree of semantic influence of each video frame on the training video; and determine the first category of video frames among the video frames according to the importance corresponding to each of the video frames.
[0032] In a possible implementation, the reference feature sequence is obtained by sorting the target reference features in order of extraction time from early to late, and the contribution value corresponding to one target reference feature is positively correlated with the arrangement position; or the reference feature sequence is obtained by sorting the target reference features in order of extraction time from late to early, and the contribution value corresponding to one target reference feature is negatively correlated with the arrangement position.
[0033] In a third aspect of the present application, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect when executing the program.
[0034] In a fourth aspect of the present application, a computer program product is provided, the computer program product comprising computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method provided in the first aspect above.
[0035] In a fifth aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions. When the computer instructions are executed on a computer, the computer executes the method described in the first aspect.
[0036] Since the embodiment of the present application adopts the above technical solution, it has at least the following technical effects:
[0037] In the embodiment of the present application, based on the idea of contrastive learning, the video processing network is trained for multiple rounds of iterations, and in the process, based on the extraction time of each target reference feature, the contribution value characterizing the degree of influence of each target reference feature on the current training is determined, so that when the trained target video processing network extracts features from the video, the feature representation of the video obtained is independent of the specific auxiliary task, and is an identical feature representation. In addition, more video timing information can be retained in the obtained video representation, thereby improving the accuracy of feature extraction of the video by the target video processing network. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 A schematic diagram of an application scenario of training a video processing network provided in an embodiment of the present application;
[0039] Figure 2 A flowchart of training a video processing network provided in an embodiment of the present application;
[0040] Figure 3 A schematic diagram of obtaining a basic video and a candidate reference video provided in an embodiment of the present application;
[0041] Figure 4 A diagram showing an example structure of a third network provided in an embodiment of the present application;
[0042] Figure 5 A schematic diagram of the principle of training a video processing network provided in an embodiment of the present application;
[0043] Figure 6 A flowchart of training a video processing network provided in an embodiment of the present application;
[0044] Figure 7 A schematic diagram of another principle of training a video processing network provided in an embodiment of the present application;
[0045] Figure 8 A flowchart of another video processing network training provided in an embodiment of the present application;
[0046] Fig. 9 A schematic diagram of the structure of a video processing network training provided in an embodiment of the present application;
[0047] Fig.10 A structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0048] In order to better understand the technical solution provided by the embodiments of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods; in order to facilitate technical personnel in the field to better understand the technical solution of the present application, some concepts involved in the present application are described below.
[0049] 1) First network, second network and target video processing network
[0050] In the embodiment of the present application, the first network and the second network are neural networks for extracting features from videos, and the first network and the second network are networks with the same network structure and different network parameters; the above-mentioned neural networks may include but are not limited to convolutional neural networks (CNN), recurrent neural networks (RNN), deep neural networks (DNN), etc., and technicians in this field can set them according to actual needs.
[0051] The target video processing network in the embodiment of the present application is the first network output after training.
[0052] 2) Self-supervised learning
[0053] Self-supervised learning does not require manually annotated category label information, but directly uses the data itself as supervision information to learn the feature expression of sample data and apply it to downstream tasks. Self-supervised learning can be divided into two main technical routes: contrastive learning and generative learning. Among them, the core idea of contrastive learning is to compare positive samples and negative samples in the feature space and learn the feature representation of samples. Contrastive learning first learns the general representation of videos on unlabeled datasets, and then can use a small amount of labeled videos to fine-tune the network to improve the performance of given tasks (such as video classification). That is, contrastive representation learning can be considered as learning through comparison. Relatively speaking, generative learning is to learn the discriminant model of the mapping of certain (pseudo) labels and then reconstruct the input samples.
[0054] The embodiments of the present application involve artificial intelligence (AI) and machine learning technology, and are designed based on computer vision technology and machine learning (ML) in artificial intelligence; artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence.
[0055] Artificial intelligence is the study of the design principles and implementation methods of various intelligent machines, which enable machines to have the functions of perception, reasoning and decision-making. It mainly includes computer vision technology, natural language processing technology, and machine learning or deep learning. With the research and progress of artificial intelligence technology, artificial intelligence has been studied and applied in many fields, such as common smart homes, smart customer service, virtual assistants, smart speakers, smart marketing, unmanned driving, automatic driving, robots, smart medical care, etc. It is believed that with the development of technology, artificial intelligence will be applied in more fields and play an increasingly important role.
[0056] Computer Vision (CV) is a science that studies how to make machines "see", that is, using cameras and computers to replace human eyes to identify, track and measure targets, and further perform image processing to make computer processing into images that are more suitable for human observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data; computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning and map construction, and other technologies, as well as common biometric recognition technologies such as face recognition and fingerprint recognition.
[0057] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.
[0058] The design concept of this application is described below.
[0059] When training a neural network for extracting features from a video, the training is often performed through supervised learning or unsupervised learning. In supervised learning, a large number of video samples for training need to be annotated, and the neural network is trained by back-propagating the cross-entropy loss of the predicted results of the video samples and the annotated true results. However, the information provided by the video samples in this method is low in richness, resulting in poor stability of the trained neural network and poor accuracy in extracting features from the video. In unsupervised learning, such as principal component analysis (PCA) and autoencoders, the input video samples are dimensional compressed to discard and merge redundant information, and the most critical information in the video samples is selected for training. However, in this process, excessive attention is paid to the pixel-level details of the video frames in the video, while the spatial dimensional features of the video frames are ignored, which results in low accuracy in extracting features from the video by the trained neural network.
[0060] In view of this, the inventors have designed a training method, apparatus, device and readable storage medium for a video processing network, which are used to improve the accuracy of feature extraction from videos; in the embodiments of the present application, based on the ideas of self-supervised learning and contrastive learning, a network for feature extraction from videos is trained, and the training is performed on unlabeled training videos; specifically, in the embodiments of the present application, multiple rounds of iterative training are performed on the first network and the second network, and the first network output by the last round of iterative training is determined as the target video processing network. In each round of iterative training: at least one historical reference feature and a candidate reference feature obtained based on the current second network are regarded as target reference features, and based on the extraction of each target reference feature time, determine the contribution value characterizing the influence of each target reference feature on the round of iterative training, and determine the first influence value corresponding to each target reference feature based on the contribution value of each target reference feature and the similarity between each target reference feature and the basic feature obtained based on the first network, and adjust the parameters of the current first network and the second network based on the first influence value corresponding to each target reference feature, wherein the basic video is obtained based on the basic video obtained from the training video, the candidate reference features are obtained based on the candidate reference videos obtained from the training video, and the historical reference features are obtained based on the historical reference videos obtained from the historical videos, and the historical videos are videos different from the above-mentioned training videos.
[0061] Furthermore, in order to increase the training difficulty of the target video processing network, in the embodiment of the present application, the training video can also be interfered based on the third network, and then at least one of the above-mentioned basic video and candidate reference video is obtained according to the training video after the interference processing.
[0062] Furthermore, in order to improve the accuracy of feature extraction of the target video processing network, in an embodiment of the present application, when adjusting the parameters of the first network and the second network, the parameters of the third network can be adjusted based on the first influence values corresponding to each target reference feature, and then the third network after the parameter adjustment can be used to perform interference processing on the training video in the subsequent training process.
[0063] In order to more clearly understand the design ideas of the present application, the application scenarios in the embodiments of the present application are introduced as follows.
[0064] See also Figure 1 , an application scenario of video processing network training is provided, the application scenario may include a terminal device 110 and a server 120; the terminal device 110 and the server 120 may communicate with each other through a network, wherein:
[0065] As an embodiment, the communication network is a wired network or a wireless network; the terminal device 110 and the server 120 may be directly or indirectly connected via wired or wireless communication, which is not limited in the present application.
[0066] In the embodiment of the present application, the terminal device 110 (such as but not limited to 110-1 or 110-2 shown in the figure) can be a mobile terminal, a fixed terminal or a portable terminal. For example, the terminal device 110 can be an electronic device used by a user, and the electronic device can be a personal computer, a mobile phone, a tablet computer, a notebook, a car device, an e-book reader, a smart home, etc., which has a certain computing power and runs instant messaging software and websites or image processing software and websites. Each terminal device 110 communicates with the server 120 through a communication network. The server 120 (such as but not limited to 120-1, 120-2 or 120-3 shown in the figure) can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and basic cloud computing services such as big data and artificial intelligence platforms.
[0067] The target video processing model can be but is not limited to being deployed on the server 120 for training. The server 120 can store and train a video set, and the training video set can include multiple training videos for training. After the target video processing model is trained based on the training method in the embodiment of the present application, the trained target video processing model can be directly deployed on the terminal device 110, or the target video processing model can be deployed on the server 120.
[0068] In the embodiment of the present application, when the target image video model is deployed on the terminal device 110, the input video to be processed is received and the video to be processed is processed. When the target video processing model is deployed on the server 120, the terminal device 110 can obtain the video to be processed and upload it to the server 120, and the server 120 processes the video to be processed and returns the processing result to the terminal device 110, and then the terminal device 110 can receive the above processing result.
[0069] As an embodiment, the present application embodiment does not limit the method for obtaining training samples, nor does it limit the specific content of the training video. Those skilled in the art can set it according to actual needs. For example, the above-mentioned training video set can include but is not limited to the following first category video A1 to third category video A3:
[0070] Video type A1: Video containing a single person's behavior; the single person's behavior can be performed by a single person, such as but not limited to painting, drinking, gambling, and punching.
[0071] Video type A2: Video containing human behaviors; wherein the human behaviors may be behaviors performed by at least two people together, such as but not limited to hugging, kissing, shaking hands, etc.
[0072] Video type A3: Video containing human behaviors; wherein the human behaviors may be interactive actions between humans and objects (such as animals, plants, static objects, etc.), such as but not limited to opening gifts, mowing the lawn, washing dishes, etc.
[0073] As an embodiment, the embodiments of the present application can but are not limited to using Kinetics as a training video set, or using UCF101 or HMDB51 as the above-mentioned training video set, wherein Kinetics is a video data set with a large amount of data provided by YouTube, UCF101 is a video data set of 101 categories from YouTube, and HMDB51 is a video data set of 51 categories from YouTube. In the embodiments of the present application, videos can also be obtained from video libraries other than the above-mentioned YouTube, and the obtained videos can be used as training videos in the above-mentioned training video set.
[0074] As an embodiment, cloud storage technology can be used in the embodiments of the present application to save the above-mentioned training video set; wherein cloud storage (Cloud Storage) is a new concept extended and developed from the concept of cloud computing, and a distributed cloud storage system (hereinafter referred to as the storage system) refers to a storage system that uses cluster applications, grid technology, and distributed storage file systems and other functions to bring together a large number of various types of storage devices (storage devices are also called storage nodes) in the network through application software or application interfaces to work together and jointly provide data storage and business access functions to the outside world.
[0075] The following is a detailed introduction to the training method of the video processing network in the embodiment of the present application. It should be noted that the above application scenarios are only shown to facilitate understanding of the spirit and principle of the present application, and the implementation of the present application is not limited in this respect.
[0076] based on Figure 1 The following is an example of a video processing network training method involved in an embodiment of the present application;
[0077] In the embodiment of the present application, the first network and the second network can be subjected to multiple rounds of iterative training based on the training video, and the first network output from the last round of iterative training is determined as the target video processing network; the first network and the second network are networks with the same network structure and different network parameters.
[0078] As an embodiment, the first network and the second network in the embodiment of the present application can be created based on the R(2+1)D encoder or the S3D-G encoder; when the first network and the second network are created based on the R(2+1)D encoder, in the architecture of the R(2+1)D encoder, Ni 3D convolution filters (Ni-1×t×d×d) are replaced by Mi spatial 2D convolution filters (Ni-1×1×d×d) and Ni temporal convolution filters (Mi×t×1×1); Mi determines the dimension of the intermediate subspace and makes the parameter amount of the network parameters of R(2+1)D equal to that of R3D; the encoder created based on R(2+1)D Compared with the full three-dimensional convolution, the first network has at least the following two advantages: first, without changing the number of parameters of the first network, the additional Relu between the two-dimensional and one-dimensional convolutions in each block doubles the number of nonlinearities in the first network. Increasing the number of nonlinearities increases the complexity of representable functions, such as the VGG network. The first network approximates the effect of a large filter by applying multiple smaller filters (with additional nonlinearities between them); second, the three-dimensional convolution is forced to separate the spatial and temporal components, making optimization easier. This shows that the first network created based on the R(2+1)D network structure has a lower training error than a three-dimensional convolutional network of the same capacity.
[0079] See also Figure 2 , wherein one round of iterative training in the above-mentioned multiple rounds of iterative training specifically includes the following steps S201 to S203.
[0080] Step S201, input the basic video obtained based on the above training video into the current first network to obtain the corresponding basic features, and input the candidate reference video obtained based on the above training video into the current second network to obtain the corresponding candidate reference features.
[0081] As an example, see Figure 3In this step, the training video may be subjected to a first data augmentation process to obtain the above-mentioned basic video, and the basic video is input into the current first network; the training video may be subjected to a second data augmentation process to obtain the above-mentioned candidate reference video, and the candidate reference video is input into the current second network; wherein the first data augmentation process and the second data augmentation process are different data augmentation processes, and the data augmentation process in the embodiment of the present application may include but is not limited to at least one of the following data IO, image mirroring, color space transformation, image cropping and scaling, image linear transformation, etc., or any combination thereof, wherein:
[0082] Data IO can include but is not limited to ToTensor, ToPILImage, PILToTensor, etc.; image mirroring classes can include but are not limited to RandomHorizontalFlip, RandomVerticalFlip, TenCrop, FiveCrop, etc.; color space transformation can include but is not limited to ColoeJitter, Grayscale, etc.; image cropping and scaling classes can include but are not limited to Resize, CenterCrop, RandomResizedCrop, TenCrop, FiveCrop, etc.; image linear transformation classes can include but are not limited to LinearTransForm, RandomRotation, RandomPerspective, etc.
[0083] Step S202, based on the above-mentioned candidate reference features and at least one historical reference feature, determine a target reference feature set, and determine the first influence value corresponding to each of the above-mentioned target reference features according to the similarity between each target reference feature in the above-mentioned target reference feature set and the above-mentioned basic features and their respective corresponding contribution values, wherein the contribution value represents the influence of the corresponding target reference feature on the above-mentioned round of iterative training, and the above-mentioned contribution value is obtained based on the extraction time of the corresponding target reference feature; when the round of iterative training is the first round of iterative training in the above-mentioned multiple rounds of iterative training, the above-mentioned historical reference features are obtained based on the above-mentioned current second network, and when the round of iterative training is not the first round of iterative training in the above-mentioned multiple rounds of iterative training, the above-mentioned historical reference features are obtained based on the historical second network.
[0084] As an embodiment, in the first round of iterative training of multiple rounds of iterative training, it is possible but not limited to obtaining part of the training videos from the above-mentioned training video set, inputting the obtained part of the training videos into the current second network, obtaining the candidate reference features corresponding to each training video in the part of the training videos, and then using the candidate reference features corresponding to each training video in the part of the training videos as the at least one historical reference feature mentioned above; wherein the above-mentioned part of the training videos may include the training videos in step S202, and the part of the training videos may also not include the training videos in step S202; in addition, there is no limitation on the number of training videos in the above-mentioned part of the training videos, and technicians in this field can set it according to actual needs.
[0085] As an embodiment, the at least one historical reference feature mentioned above may be a historical reference feature in the current batch. The historical reference features in the current batch may be saved in a queue, and a preset number of historical reference features may be saved in the queue. In each round of iterative training in multiple rounds of iterative training, the obtained candidate reference features may be added to the queue, and the historical reference features that first entered the queue may be removed from the queue. That is, in non-first iterative training in multiple rounds of iterative training, the latest obtained candidate reference features may be added to the queue, and the historical reference features that first entered the queue may be removed from the queue, so that the preset number of historical reference features may be saved in the queue as the at least one historical reference feature mentioned above.
[0086] As an embodiment, as the acquisition time of historical reference features decreases from late to early, their influence on this round of iterative training continues to decay, that is, the earlier the acquisition time of a historical reference feature, the earlier it is acquired by the second network, and therefore the weaker its influence on this round of iterative training. Therefore, in the embodiment of the present application, the contribution value corresponding to the first target reference feature in the above-mentioned target reference feature set is less than the contribution value corresponding to the second target reference feature, wherein the extraction time of the above-mentioned first target reference feature is earlier than the extraction time of the above-mentioned second target reference feature. Specifically, technical personnel in this field can flexibly obtain the contribution values corresponding to the above-mentioned each target reference feature.
[0087] As an embodiment, in this step, the following operations may be performed respectively for each target reference feature in the target reference feature set:
[0088] According to the extraction time of a target reference feature in the above-mentioned target reference feature set, the contribution value corresponding to the above-mentioned target reference feature is determined; and the similarity between the above-mentioned target reference feature and the above-mentioned basic feature is determined; and then based on the contribution value corresponding to the above-mentioned target reference feature, the determined similarity is weighted to obtain a first influence value corresponding to the above-mentioned target reference feature; wherein, when the above-mentioned target reference feature and the basic feature are in the form of feature vectors, in the embodiment of the present application, the similarity between the above-mentioned target reference feature and the basic feature can be determined based on but not limited to the distance between the above-mentioned target reference feature and the basic feature, such as the L1 distance between the above-mentioned target reference feature and the basic feature can be determined as the similarity between the above-mentioned target reference feature and the basic feature; in the embodiment of the present application, the specific method for determining the similarity between the two features (i.e., the target reference feature and the basic feature) is not limited, and technicians in this field can set it according to actual needs.
[0089] As an embodiment, the specific method of weighting the determined similarity based on the contribution value corresponding to the above-mentioned target reference feature is not limited. Technical personnel in this field can set it according to actual needs. For example, please refer to the following formula (1a). It can be but not limited to determining the product of the contribution value corresponding to the above-mentioned target reference feature and the above-mentioned determined similarity as the first influence value corresponding to the above-mentioned target reference feature. It can also be based on the principle of the following formula (1b) to obtain the first influence value corresponding to the above-mentioned target reference feature.
[0090] Inf_j=con_j×exp(k j ×q) Formula (1a)
[0091] Inf_j=M con_j ×exp(k j ×q) Formula (1b)
[0092] In formulas (1a) and (1b), j is the identifier of the target reference feature, inf_j is the first influence value corresponding to the target reference feature identified as j, M is a value greater than 1, con_j is the contribution value corresponding to the target reference feature identified as j, and k j is the target reference feature identified as j, and q is the above basic feature.
[0093] Step S203, adjusting the parameters of the current first network according to the first influence values corresponding to the above-mentioned each target reference feature, to obtain the first network output by the above-mentioned round of iterative training; and adjusting the parameters of the current second network based on the first network output by the above-mentioned round of iterative training.
[0094] As an embodiment, in order to improve the accuracy of feature extraction of the target video processing network obtained by training, the embodiment of the present application is trained in the direction of improving the similarity between the basic features of the training video and the candidate reference features, and reducing the similarity between the basic features obtained based on the training video and the historical candidate reference features obtained based on the historical candidate reference video; specifically, in the embodiment of the present application, the sum of the first influence values corresponding to each of the above-mentioned target reference features can be determined as the negative sample influence value; based on the above-mentioned negative sample influence value and the first influence value corresponding to the above-mentioned target reference feature, the model loss value of the above-mentioned current first network is determined; the above-mentioned model loss value is negatively correlated with the above-mentioned negative sample influence value, and the above-mentioned model loss value is positively correlated with the first influence value corresponding to the above-mentioned target reference feature; and based on the above-mentioned model loss value, the parameters of the above-mentioned current first network are adjusted.
[0095] As an embodiment, for ease of understanding, the embodiment of the present application provides a specific method for determining the above-mentioned model loss value. In the embodiment of the present application, the above-mentioned model loss value can be determined based on but not limited to the following formula (2).
[0096]
[0097] In formula (2), L q is the above model loss value, j is the identifier of the target reference feature, Inf_0 is the first influence value of the candidate reference feature, K is the number of at least one historical reference feature, and Inf_j is the first influence value of the target reference feature identified as j.
[0098] As an embodiment, in step S203, L is minimized q , the parameters of the first network are adjusted through post-feedback, and then based on the second network after the parameter adjustment, the parameters of the second network can be adjusted based on the principle of the following formula (3).
[0099] θ k ←mθ k +(1-m)θ q Formula (3)
[0100] In formula (3), θ k refers to the parameters of the second network, θ q refers to the parameters of the first network, and m is a momentum coefficient between 0.0 and 1.0.
[0101] In the embodiment of the present application, other methods may be used to adjust the parameters of the current second network based on the first network after the parameter adjustment.
[0102] As an embodiment, when creating the initial first network and the second network, the network parameters of the first network and the second network may be the same or different; when the network parameters of the initial first network and the second network are the same, in the first round of iterative training, the network parameters of the first network and the second network are the same, but after the first round of iterative training, the network parameters of the first network and the second network will be different.
[0103] As an embodiment, the method of determining the contribution value corresponding to each target reference feature in step S202 is further described below.
[0104] In order to improve the efficiency of determining contribution values, in an embodiment of the present application, the above-mentioned target reference features can be sorted based on the extraction time of the above-mentioned target reference features to obtain a reference feature sequence, and then when determining the contribution value corresponding to a target reference feature in the target reference feature set, the contribution value corresponding to the above-mentioned target reference feature can be determined based on the arrangement position of the above-mentioned target reference feature in the reference feature sequence.
[0105] As an embodiment, if the above-mentioned reference feature sequence is a first reference feature sequence, then the contribution value corresponding to a target reference feature is positively correlated with the arrangement position of the target reference feature in the first reference feature sequence; wherein the first reference feature sequence is obtained by sorting the above-mentioned target reference features in the order of extraction time from early to late; if the extraction time of target reference features 1 to 5 is from early to late: target reference feature 1, target reference feature 2, target reference feature 3, target reference feature 4 and target reference feature 5, then the first reference feature sequence is {target reference feature 1, target reference feature 2, target reference feature 3, target reference feature 4, target reference feature 5}, the arrangement positions of target reference features 1 to 5 in the first reference feature sequence are 1, 2, 3, 4 and 5 respectively, and the contribution values corresponding to the target reference features 1 to 5 are from large to small: target reference feature 5, target reference feature 4, target reference feature 3, target reference feature 2 and target reference feature 1; specifically, the contribution value corresponding to each target reference feature can be determined based on, but not limited to, the principle of the following formula (4).
[0106]
[0107] In formula (4), i is the arrangement position of the target reference feature in the first reference feature sequence, con_i is the contribution value corresponding to the target reference feature at the arrangement position i in the first reference feature sequence, and t 1 Is a value greater than 1.
[0108] As an embodiment, if the above-mentioned reference feature sequence is a second reference feature sequence, then the contribution value corresponding to a target reference feature is negatively correlated with the arrangement position of the target reference feature in the second reference feature sequence; wherein the second reference feature sequence is obtained by sorting the above-mentioned target reference features in order of extraction time from late to early; if the extraction time of target reference features 1 to 5 is from early to late: target reference feature 1, target reference feature 2, target reference feature 3, target reference feature 4 and target reference feature 5, then the second reference feature sequence is {target reference feature 5, target reference feature 4, target reference feature 3, target reference feature 2, target reference feature 1}, and the arrangement positions of target reference features 1 to 5 in the first reference feature sequence are 5, 4, 3, 2 and 1 respectively, and the contribution values corresponding to the target reference features 1 to 5 are from large to small: target reference feature 5, target reference feature 4, target reference feature 3, target reference feature 2 and target reference feature 1; specifically, the contribution value corresponding to each target reference feature can be determined based on but not limited to the principle of the following formula (5a) or formula (5b).
[0109]
[0110]
[0111] In formulas (5a) and (5b), i is the position of the target reference feature in the second reference feature sequence, con_i is the contribution value corresponding to the target reference feature at position i in the second reference feature sequence, and t 2 is a value greater than 0 and less than 1, t 3 Is a value greater than 1.
[0112] As an embodiment, when determining the contribution value corresponding to each target reference feature based on the above formula (5a), the above formula 2 can also be converted into the form of the following formula (2a).
[0113]
[0114] In formula (2a), L q is the loss value of the above model, i is the identifier of the target reference feature, Inf_0 is the first influence value of the candidate reference feature, K is the number of at least one historical reference feature, Inf_i is the first influence value of the target reference feature identified as i, and k 0 is a candidate reference feature, k i is the target reference feature identified as i, q is the basic feature, τ is a constant, and t is a value greater than 0 and less than 1.
[0115] As an embodiment, in order to improve the efficiency of determining the model loss value, the value range of the first impact value can also be narrowed in the embodiment of the present application. Specifically, the above formula (2a) can be transformed into formula (2b), and then based on formula (2b), the above model loss value can be determined.
[0116]
[0117] As an embodiment, in order to increase the training difficulty of the target video processing network, in the embodiment of the present application, before performing the above steps S201 to S203, the training video may be interfered, and then the training process of the above steps S201 to S203 may be performed according to the training video after the interference processing and the training video before the interference training; specifically, in the embodiment of the present application, the training video may be but is not limited to being processed as follows to obtain the above-mentioned basic video and candidate reference video:
[0118] Based on the third network, the training video is interfered with to obtain a target training video; and then any one of the video processing C1 to video processing C3 in the following operations is performed to obtain the basic video and the candidate reference video:
[0119] Video processing C1: performing a first data augmentation process on the target training video to obtain the basic video, and performing a second data augmentation process on the training video to obtain the candidate reference video; Video processing C2: performing the first data augmentation process on the training video to obtain the basic video, and performing the second data augmentation process on the target training video to obtain the candidate reference video; Video processing C3: performing a first data augmentation process on the target training video to obtain the basic video, and performing a second data augmentation process on the target training video to obtain the candidate reference video; wherein, the specific contents of the first data augmentation process and the second data augmentation process involved in the above-mentioned video processing C1 to video processing C3 can be found in the above description and will not be repeated here.
[0120] As an embodiment, in order to further improve the accuracy of feature extraction of the target video processing network obtained by training, in the embodiment of the present application, in each round of iterative training in the above-mentioned multiple rounds of iterative training, the training samples can be interfered with in the above-mentioned manner, and the above-mentioned basic video and candidate reference video can be obtained by any one of the above-mentioned video processing C1 to video processing C3; the embodiment of the present application can also first perform training for a period of time through the above-mentioned steps S201 to step S203, and after determining that the model loss value of the current first network meets the preset loss value stability condition, in each round of iterative training, the training samples are interfered with in the above-mentioned manner, and the above-mentioned basic video and candidate reference video can be obtained by any one of the above-mentioned video processing C1 to video processing C3; wherein the above-mentioned model loss value represents the degree of loss of the above-mentioned current first network in extracting features from the video, and the method for determining the model loss value can refer to the above description, which will not be repeated here.
[0121] As an embodiment, the above-mentioned preset loss value stability condition is not excessively limited, and technicians in this field can set it according to actual needs, such as but not limited to setting the preset loss value stability condition to one or any combination of the following losses: the training time of the first network and the second network reaches the first time threshold; the current round of iterative training reaches the first training round threshold; the current model loss value of the first network reaches the first loss value threshold; the difference between the model loss value of the first network after the parameters of the first network are adjusted in the current round and the model loss value of the first network after the parameters of the first network are adjusted in the previous round is less than the first loss difference threshold; the difference in the model loss value of the first network in n consecutive rounds of iterative training is less than the first loss difference threshold.
[0122] As an embodiment, in order to further improve the accuracy of feature extraction of the target video processing network, in the embodiment of the present application, when adjusting the parameters of the first network and the second network, the parameters of the third network can be adjusted based on the first influence values corresponding to each target reference feature, and then the third network after the parameter adjustment can be used to perform interference processing on the training video in the subsequent training process. Specifically, in each round of iterative training, after determining the first influence values corresponding to each of the above-mentioned target reference features, the parameters of the third network can be further adjusted according to the first influence values corresponding to each of the above-mentioned target reference features.
[0123] As an embodiment, the parameters of the current third network can be adjusted based on the above-mentioned model loss value; wherein, the method for determining the model loss value can refer to the above description, which will not be repeated here, or it can be based on the following formula (6), based on the first influence value corresponding to each target reference feature, to determine the network loss value of the third network, and then based on the above-mentioned network loss value, the parameters of the above-mentioned third network are adjusted.
[0124]
[0125] In formula (6), L is the network loss value of the third network, j is the identifier of the target reference feature, Inf_0 is the first influence value of the candidate reference feature, K is the number of at least one historical reference feature, and Inf_j is the first influence value of the target reference feature identified as j.
[0126] The following content of the embodiment of the present application describes the specific manner of performing interference processing on the training video, wherein the interference processing involved in the embodiment of the present application may include but is not limited to one or any combination of the following:
[0127] The first interference processing: delete the first type of video frames in the above training video.
[0128] Specifically, the above-mentioned first category of videos refers to video frames that can be deleted in the training video. The above-mentioned first category of video frames can be one or more video frames randomly determined from the video frames in the training video. The first category of videos can also be determined based on the degree of semantic influence of each video frame on the training video. There is no limitation on the number of first category videos. Technical personnel in this field can set it according to actual needs. For example, in each round of iterative training, the same number of first category video frames in the training video can be deleted, and as the number of iterative training rounds increases, the number of first category video frames deleted in the training video can also increase accordingly.
[0129] As an embodiment, in order to further increase the difficulty of training so as to enhance the ability of the first network to learn the features of the video, in the embodiment of the present application, before deleting the first category of video frames in the above-mentioned training video, the image features of each of the above-mentioned video frames can be further extracted based on the feature extraction subnetwork in the above-mentioned third network; the image features of each of the above-mentioned video frames are input into the long short-term memory subnetwork in the above-mentioned third network to determine the importance corresponding to each of the above-mentioned video frames, and the above-mentioned importance represents the degree of semantic influence of each video frame on the above-mentioned training video; according to the importance corresponding to each of the above-mentioned video frames, the first category of video frames among the above-mentioned video frames are determined.
[0130] The network structure of the third network is not limited in detail, and those skilled in the art can set it according to actual needs. Figure 4, a structural example diagram of a third network is given, in which the ConLSTM network is used as the above, and ConvLayers and reshape in the ConLSTM network are used as feature extraction subnetworks, and ConvLayers are used to extract features of each video frame in the training video, and reshape is used to output the feature representation of each video frame (such as but not limited to the feature representation of video frame 1, the feature representation of video frame 2, the feature representation of video frame 3, etc. shown in the figure), and the feature representation can be but not limited to a feature vector; and then through the long short-term memory subnetwork LSTM, based on the feature representation of each video frame, the importance of each video frame is determined. After determining the importance of each video frame, a time mask sequence mask can be generated based on the importance of each video frame, and then the time mask sequence is used to delete the first type of video in the training video.
[0131] The second interference processing: adjusting the position of the second type of video frame in the training video.
[0132] Specifically, the positions of some video frames (i.e., second-category video frames) in the training video can be randomly adjusted, or the positions of all video frames in the training video can be adjusted, such as reversing the order of all video frames in the entire training video; the above-mentioned first-category video frames can also be regarded as second-category video frames, and then the positions of the first-category video frames in the training video can be adjusted.
[0133] The third interference processing: modify the color information of the third type of video frames in the above training video.
[0134] Specifically, there are no excessive restrictions on the method of modifying the color information of the video frame, and technical personnel in this field can set it according to actual needs; among them, there are no excessive restrictions on the third category of video frames, and technical personnel in this field can set it according to actual needs, such as but not limited to randomly selecting some video frames from the training video as the above-mentioned third category video frames, or based on the above-mentioned third network, based on the importance of each video frame in the training video, selecting certain video frames from the training video as the third category video frames, etc.
[0135] As an embodiment, a complete example of a training method for a video processing network is provided below.
[0136] In this example, the encoding network (encoder) in the MoCo network (hereinafter referred to as the discriminator D) is used as the first network, the momentum encoding network (momentum encoder) in the MoCo network is used as the second network, and the ConvLSTM network is used as the third network as an example for explanation; and in this example, the network structure of the encoder and momentum encoder is constructed based on the network structure of R(2+1)D.
[0137] In this example, the process of training the video processing network mainly includes two training stages. In the first training stage, the training video is not interfered, and the contribution value corresponding to each target reference feature is determined based on the extraction time of each target reference feature. Based on the above contribution value, the discriminator D and the momentum encoder are iteratively trained for multiple rounds until the current model loss value of the discriminator D meets the above preset loss value stability condition. In the second training stage, after determining that the current model loss value of the discriminator D meets the above preset loss value stability condition, the discriminator D is trained for adversarial learning based on the ConvLSTM network (hereinafter referred to as the generator G). In each round of training, the training video is interfered based on the generator G, and the discriminator D and the generator G are trained for adversarial learning using the interfered training video.
[0138] The following content describes the first training stage in detail.
[0139] Specifically, see Figure 5 In the first training stage, the training video is directly subjected to the first data augmentation process to obtain the basic video x query , the basic video x query Input the current encoder to obtain the basic feature q; directly perform the second data augmentation process on the training video to obtain the basic video to obtain the candidate reference video x key , the candidate reference video x key Input the current momentum encoder to obtain the candidate reference feature k; in this example, after the historical reference video obtained based on the historical video is input into the historical momentum encoder, the corresponding historical reference feature is obtained and added to the memory queue, wherein the historical reference feature and the candidate reference feature are used as the target reference feature, and based on the extraction time of each target reference feature, the corresponding contribution value of each target reference feature is temporally decayed;
[0140] Specifically, each target reference feature can be sorted, and the contribution value corresponding to each target reference feature can be determined according to the arrangement position of each target reference feature. As shown in the figure, the target reference features are sorted in order from the latest to the earliest extraction time, and the target reference feature k with the arrangement position i is i The corresponding contribution value is t i , and then based on each target reference feature k i The corresponding contribution value t i And each target reference feature k i The similarity between the target feature q and the basic feature q is used to determine the first influence value corresponding to each target reference feature, and the contrast loss of the encoder (that is, the above-mentioned model loss value) is determined based on the first influence value corresponding to each target reference feature, and then the contrast loss and gradient are used to adjust the parameters of the encoder, and based on the encoder after the parameter adjustment, the parameters of the momentum encoder are adjusted.
[0141] See also Figure 6 One round of iterative training in the multiple rounds of iterative training in the first training stage may include, but is not limited to, the following steps S601 to S609:
[0142] Step S601, obtaining a training video from a training video set.
[0143] Specifically, currently unused training videos can be randomly obtained from the training video set, or currently unused training videos can be obtained from the training video set in sequence according to the arrangement order of the training videos in the training video set, etc. Technical personnel in this field can set it according to actual needs.
[0144] Step S602: Perform a first data augmentation process on the acquired training video to obtain a basic video, and perform a second data augmentation process on the acquired training video to obtain a candidate reference video.
[0145] Step S603: input the basic video into the current first network to obtain the corresponding basic features, and input the base candidate reference video into the current second network to obtain the corresponding candidate reference features.
[0146] Step S604: determine a target reference feature set based on the candidate reference features and at least one historical reference feature obtained from the historical second network, and determine a contribution value corresponding to each target reference feature based on the extraction time of each target reference feature in the target reference feature set.
[0147] Step S605 , based on the contribution values corresponding to the respective target reference features, weighted processing is performed on the similarities between the respective target reference features and the basic features to obtain the first influence values corresponding to the respective target reference features.
[0148] Step S606: Determine the model loss value of the first network according to the first influence values corresponding to the above-mentioned respective target reference features, and adjust the parameters of the current first network based on the above-mentioned model loss values.
[0149] Step S607: Based on the first network after parameter adjustment, adjust the parameters of the current second network.
[0150] Step S608, determining whether the network training end condition is met, if so, the training ends, otherwise, proceeding to step S609.
[0151] There are no excessive restrictions on the above-mentioned network training termination conditions. Technical personnel in this field can set them according to actual needs, such as but not limited to setting the above-mentioned network training termination adjustment element to one of the following or any combination: the training duration reaches the second duration threshold; the current round of iterative training reaches the second training round threshold; the current model loss value of the first network reaches the second loss value threshold, etc.
[0152] Step S609, determining whether the current model loss value of the first network meets the preset loss value stability condition, if not, proceeding to step S601, otherwise entering the second training stage.
[0153] The following content describes the second training phase in detail.
[0154] Specifically, see Figure 7 In the second training stage, in each round of iterative training, the generator G (Generator) (i.e., the ConLSTM network) performs interference processing on the training video, extracts features of each video frame in the training video through ConvLayers, and outputs the image features of each video frame through reshape, and then determines the importance of each video frame based on the image features of each video frame through LSTM, generates a time mask sequence mask based on the importance of each video frame, and then uses the mask to delete the first type of video in the training video to obtain the target training video; after performing the first data augmentation processing on the target training video, the basic video x is obtained. query , the basic video x query Input the current discriminator D (i.e. the encoder mentioned above) to obtain the basic feature q; directly perform the second data augmentation process on the training video to obtain the basic video to obtain the candidate reference video x key , the candidate reference video x keyInput the current momentum encoder to obtain the candidate reference feature k;
[0155] Then, the candidate reference features and the historical reference features are used as target reference features, and the first influence values corresponding to each target reference feature are determined. The contrast loss of the encoder (i.e., the above-mentioned model loss value) is determined based on the first influence values corresponding to each target reference feature, and then the contrastive loss is used to adjust the parameters of the encoder, and based on the encoder after the parameter adjustment, the parameters of the momentum encoder are adjusted; and the parameters of the generator G are adjusted based on the contrastive loss, wherein the process of determining the first influence value corresponding to each target reference feature can be found in the relevant description of the first training stage, which will not be repeated here.
[0156] See also Figure 8 One round of iterative training in the multiple rounds of iterative training in the first training stage may include, but is not limited to, the following steps S801 to S810:
[0157] Step S801, obtaining a training video from a training video set.
[0158] The specific method of obtaining the training video can refer to the description of step S501.
[0159] Step S802: The generator G performs interference processing on the acquired training video to obtain a target training video.
[0160] Step S803: Perform a first data augmentation process on the target training video to obtain a basic video, and perform a second data augmentation process on the obtained training video to obtain a candidate reference video.
[0161] Step S804: input the base video into the current first network to obtain the corresponding base features, and input the base candidate reference video into the current second network to obtain the corresponding candidate reference features.
[0162] Step S805, determining a target reference feature set based on the candidate reference features and at least one historical reference feature obtained according to the historical second network, and determining a contribution value corresponding to each target reference feature based on the extraction time of each target reference feature in the target reference feature set.
[0163] Step S806: Based on the contribution values corresponding to the respective target reference features, weighted processing is performed on the similarities between the respective target reference features and the basic features to obtain the first influence values corresponding to the respective target reference features.
[0164] Step S807: determining the model loss value of the first network according to the first influence values corresponding to the above-mentioned respective target reference features, and adjusting the parameters of the current first network based on the above-mentioned model loss values.
[0165] Step S808: Based on the first network after parameter adjustment, adjust the parameters of the current second network.
[0166] Step S809: Based on the first network after parameter adjustment, adjust the parameters of the current generator G.
[0167] Step S810, determining whether the network training end condition is met, if so, the training ends, otherwise, proceeding to step S801.
[0168] In the method provided in the embodiment of the present application, based on the idea of contrastive learning, if for a training video, there is 1 positive example pair (i.e., the candidate reference video mentioned above) and N-1 negative example pairs (i.e., the historical reference video that obtains the above at least one historical reference feature), the model loss value of the first network can be regarded as an N-type problem, which is actually a cross entropy. In the embodiment of the present application, the positive example pair is obtained by performing data augmentation processing on the current training sample, and the negative example pair is obtained by performing data augmentation processing on historical videos different from the current training sample, that is, all samples in the batch queue are considered to be negative example pairs; then the basic video is input into the first network, and the candidate reference candidate video is input into the second network, and the contrastive loss is used to guide the trainable first network to learn the features of the video;
[0169] On the one hand, in the embodiment of the present application, when the trained target video processing network extracts features from the video, the feature representation of the video obtained is independent of the specific auxiliary task and is an identical feature representation. In addition, more video timing information can be retained in the obtained video representation, thereby improving the accuracy of feature extraction of the video by the target video processing network. On the other hand, the generator G is introduced in the training process for adversarial learning, and no additional network is required in the actual testing phase. Therefore, the target video processing network obtained by the method provided in the embodiment of the present application has good universality.
[0170] Please refer to Fig. 9 Based on the same inventive concept, the present application embodiment provides a training device for a video processing network, including:
[0171] The network training unit 9000 is used to: perform multiple rounds of iterative training on the first network and the second network based on the training video, and determine the first network output from the last round of iterative training as the target video processing network; the first network and the second network have the same network structure and different network parameters; the network training unit includes a feature extraction subunit 9001, an influence value determination subunit 9002 and a network adjustment subunit 9003, wherein:
[0172] The feature extraction subunit 9001 is used to input the basic video obtained based on the training video into the current first network in one round of iterative training to obtain the corresponding basic features, and input the candidate reference video obtained based on the training video into the current second network to obtain the corresponding candidate reference features;
[0173] The above-mentioned influence value determination subunit 9002 is used to determine a target reference feature set based on the above-mentioned candidate reference features and at least one historical reference feature in a round of iterative training, and determine the first influence value corresponding to each of the above-mentioned target reference features according to the similarity between each target reference feature in the above-mentioned target reference feature set and the above-mentioned basic features and their respective corresponding contribution values, wherein the contribution value represents the influence degree of the corresponding target reference feature on the above-mentioned round of iterative training, and the above-mentioned contribution value is obtained based on the extraction time of the corresponding target reference feature; when the above-mentioned round of iterative training is the first round of iterative training in the above-mentioned multiple rounds of iterative training, the above-mentioned historical reference features are obtained based on the above-mentioned current second network, and when the above-mentioned round of iterative training is not the first round of iterative training in the above-mentioned multiple rounds of iterative training, the above-mentioned historical reference features are obtained based on the historical second network;
[0174] The network adjustment subunit 9003 is used to adjust the parameters of the current first network in a round of iterative training according to the first influence values corresponding to each of the target reference features to obtain the first network output by the above round of iterative training; and to adjust the parameters of the current second network based on the first network output by the above round of iterative training.
[0175] As an embodiment, a contribution value corresponding to a first target reference feature in the target reference feature set is smaller than a contribution value corresponding to a second target reference feature, wherein an extraction time of the first target reference feature is earlier than an extraction time of the second target reference feature.
[0176] As an embodiment, the above-mentioned basic video is obtained by performing a first data augmentation process on the target training video, and the above-mentioned candidate reference video is obtained by performing a second data augmentation process on the training video; wherein, the above-mentioned target training video is obtained by performing interference process on the above-mentioned training video based on a third network; or the above-mentioned basic video is obtained by performing the first data augmentation process on the above-mentioned training video, and the above-mentioned candidate reference video is obtained by performing the second data augmentation process on the above-mentioned target training video; or the above-mentioned basic video is obtained by performing the first data augmentation process on the above-mentioned target training video, and the above-mentioned candidate reference video is obtained by performing the second data augmentation process on the above-mentioned target training video.
[0177] As an embodiment, the feature extraction subunit 9001 is further used to: before performing interference processing on the training video based on the third network, determine whether the model loss value of the current first network meets a preset loss value stability condition, and the model loss value represents the degree of loss of the current first network in extracting features from the video.
[0178] As an embodiment, the network adjustment subunit 9003 is further used to: adjust parameters of the third network according to the first influence values corresponding to the respective target reference features.
[0179] As an embodiment, the network adjustment subunit 9003 is specifically used to: determine the sum of the first influence values corresponding to each of the above-mentioned target reference features as the negative sample influence value; determine the model loss value of the above-mentioned current first network based on the above-mentioned negative sample influence value and the first influence value corresponding to the above-mentioned target reference feature; the above-mentioned model loss value is negatively correlated with the above-mentioned negative sample influence value, and the above-mentioned model loss value is positively correlated with the first influence value corresponding to the above-mentioned target reference feature; based on the above-mentioned model loss value, adjust the parameters of the above-mentioned current first network; based on the above-mentioned model loss value, adjust the parameters of the above-mentioned third network.
[0180] As an embodiment, the influence value determination subunit 9002 is specifically used to: perform the following operations for each target reference feature in the above-mentioned target reference feature set: determine the contribution value corresponding to the above-mentioned target reference feature according to the extraction time of the above-mentioned target reference feature in the above-mentioned target reference feature set; determine the similarity between the above-mentioned target reference feature and the above-mentioned basic feature; based on the contribution value corresponding to the above-mentioned target reference feature, weight the determined similarity to obtain the first influence value corresponding to the above-mentioned target reference feature.
[0181] As an embodiment, the influence value determination subunit 9002 is specifically used to: determine the contribution value corresponding to the above-mentioned target reference feature based on the arrangement position of the above-mentioned target reference feature in the reference feature sequence; wherein the above-mentioned reference feature sequence is obtained by sorting the above-mentioned each target reference feature based on the extraction time of the above-mentioned each target reference feature.
[0182] As an embodiment, the first network is created based on an R(2+1)D encoder or an S3D-G encoder.
[0183] As an embodiment, the above-mentioned interference processing includes one or any combination of the following: deleting the first type of video frames in the above-mentioned training video; adjusting the position of the second type of video frames in the above-mentioned training video in the above-mentioned training video; modifying the color information of the third type of video frames in the above-mentioned training video.
[0184] As an embodiment, the feature extraction subunit 9001 is further used to: before deleting the first category of video frames in the above-mentioned training video, based on the feature extraction subnetwork in the above-mentioned third network, extract the image features of each of the above-mentioned video frames; input the image features of each of the above-mentioned video frames into the long short-term memory subnetwork in the above-mentioned third network to determine the importance corresponding to each of the above-mentioned video frames, and the above-mentioned importance represents the degree of semantic influence of each video frame on the above-mentioned training video; according to the importance corresponding to each of the above-mentioned video frames, determine the first category of video frames among the above-mentioned video frames.
[0185] As an embodiment, the above-mentioned reference feature sequence is obtained by sorting the above-mentioned target reference features in order from early to late extraction time, and the contribution value corresponding to the above-mentioned one target reference feature is positively correlated with the above-mentioned arrangement position; or the above-mentioned reference feature sequence is obtained by sorting the above-mentioned each target reference feature in order from late to early extraction time, and the contribution value corresponding to the above-mentioned one target reference feature is negatively correlated with the above-mentioned arrangement position.
[0186] As an example, Fig. 9 The device in can be used to implement any video processing network training method discussed above.
[0187] Based on the same inventive concept as the above method embodiment, a computer device is also provided in the embodiment of the present application. The computer device can be used for data processing based on push content. In one embodiment, the computer device can be a server, such as Figure 1 In this embodiment, the structure of the computer device can be as follows: Fig.10 As shown, it includes a memory 1001 , a communication module 1003 and one or more processors 1002 .
[0188] The memory 1001 is used to store computer programs executed by the processor 1002. The memory 1001 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system and programs required for running the instant messaging function, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.
[0189] The memory 1001 may be a volatile memory, such as a random access memory (RAM); the memory 1001 may also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD); or the memory 1001 may be any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1001 may be a combination of the above memories.
[0190] The processor 1002 may include one or more central processing units (CPU) or a digital processing unit, etc. The processor 1002 is used to implement the above-mentioned video processing network training method when calling the computer program stored in the memory 1001 .
[0191] The communication module 1003 is used to communicate with terminal devices and other servers.
[0192] The specific connection medium between the memory 1001, the communication module 1003 and the processor 1002 is not limited in the embodiment of the present application. Fig.10 In the embodiment, the memory 1001 and the processor 1002 are connected via a bus 1004. The bus 1004 is connected to the processor 1002 via a bus 1004. Fig.10 The connections between other components are shown in bold lines, which are only for illustration and are not intended to be limiting. Bus 1004 can be divided into address bus, data bus, control bus, etc. For ease of illustration, Fig.10 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0193] The memory 1001 stores a computer storage medium, and the computer storage medium stores computer executable instructions, and the computer executable instructions are used to implement the method for extracting account features in the embodiment of the present application. The processor 1002 is used to execute the training method of the video processing network.
[0194] A person skilled in the art can understand that: all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), disks or optical disks, etc. Various media that can store program codes.
[0195] Alternatively, if the above-mentioned integrated unit of the invention is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention can essentially or partly contribute to the prior art in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the training method of the video processing network in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.
[0196] Based on the same technical concept, an embodiment of the present application also provides a computer-readable storage medium, which stores computer instructions. When the above-mentioned computer instructions are executed on a computer, the computer executes the training method of the video processing network discussed above.
[0197] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0198] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A training method for a video processing network, It is characterized in that include: Based on the training video, the first network and the second network are trained for multiple rounds of iterations, and the first network output from the last round of iteration training is determined as the target video processing network; The first network and the second network are networks with the same network structure and different network parameters, wherein one round of iterative training includes: Inputting a basic video obtained based on the training video into the current first network to obtain corresponding basic features, and inputting a candidate reference video obtained based on the training video into the current second network to obtain corresponding candidate reference features; Based on the candidate reference features and at least one historical reference feature, a target reference feature set is determined, and according to the similarity between each target reference feature in the target reference feature set and the basic feature and the corresponding contribution value, the first influence value corresponding to each target reference feature is determined, wherein the contribution value represents the influence of the corresponding target reference feature on the one round of iterative training, and the contribution value is obtained based on the extraction time of the corresponding target reference feature; when the one round of iterative training is the first round of iterative training in the multiple rounds of iterative training, the historical reference feature is obtained based on the current second network, and when the one round of iterative training is not the first round of iterative training in the multiple rounds of iterative training, the historical reference feature is obtained based on the historical second network; According to the first influence values corresponding to the respective target reference features, the parameters of the current first network are adjusted to obtain the first network output by the one-round iterative training; and based on the first network output by the one-round iterative training, the parameters of the current second network are adjusted.
2. The method according to claim 1, It is characterized in that A contribution value corresponding to a first target reference feature in the target reference feature set is smaller than a contribution value corresponding to a second target reference feature, wherein an extraction time of the first target reference feature is earlier than an extraction time of the second target reference feature.
3. The method according to claim 1, It is characterized in that The basic video is obtained by performing a first data augmentation process on the target training video, and the candidate reference video is obtained by performing a second data augmentation process on the training video; wherein the target training video is obtained by performing interference processing on the training video based on a third network; or The basic video is obtained by performing the first data augmentation process on the training video, and the candidate reference video is obtained by performing the second data augmentation process on the target training video; or The basic video is obtained by performing the first data augmentation process on the target training video, and the candidate reference video is obtained by performing the second data augmentation process on the target training video.
4. The method according to claim 3, It is characterized in that Before the interference processing is performed on the training video based on the third network, the method further includes: It is determined that the model loss value of the current first network satisfies a preset loss value stability condition, and the model loss value represents the degree of loss of the current first network in extracting features from the video.
5. The method according to claim 3, It is characterized in that The method further comprises: Parameters of the third network are adjusted according to the first influence values corresponding to the respective target reference features.
6. The method according to claim 5, It is characterized in that The step of adjusting parameters of the current first network according to the first influence values corresponding to the respective target reference features includes: Determine the sum of the first influence values corresponding to each of the target reference features as the negative sample influence value; Determine a model loss value of the current first network based on the negative sample influence value and the first influence value corresponding to the target reference feature; the model loss value is negatively correlated with the negative sample influence value, and the model loss value is positively correlated with the first influence value corresponding to the target reference feature; Based on the model loss value, adjusting parameters of the current first network; The step of adjusting parameters of the third network according to the first influence values corresponding to the respective target reference features includes: Based on the model loss value, parameters of the third network are adjusted.
7. The method according to any one of claims 1 to 6, It is characterized in that The determining, according to the similarity between each target reference feature in the target reference feature set and the basic feature and the corresponding contribution value thereof, the first influence value corresponding to each target reference feature respectively comprises: For each target reference feature in the target reference feature set, perform the following operations respectively: Determining a contribution value corresponding to a target reference feature in the target reference feature set according to an extraction time of the target reference feature; Determining a similarity between the one target reference feature and the basic feature; Based on the contribution value corresponding to the one target reference feature, the determined similarity is weighted to obtain a first influence value corresponding to the one target reference feature.
8. The method according to claim 7, It is characterized in that The step of determining, according to the extraction time of a target reference feature in the target reference feature set, a contribution value corresponding to the target reference feature, comprises: Based on the arrangement position of the target reference feature in the reference feature sequence, a contribution value corresponding to the target reference feature is determined; wherein the reference feature sequence is obtained by sorting the target reference features based on the extraction time of the target reference features.
9. The method according to any one of claims 1 to 6, It is characterized in that The first network is created based on an R(2+1)D encoder or an S3D-G encoder.
10. The method according to any one of claims 3 to 6, It is characterized in that The interference processing includes one or any combination of the following: Deleting the first type of video frames in the training video; Adjusting the position of the second type of video frames in the training video; The color information of the third type of video frames in the training video is modified.
11. The method according to claim 10, It is characterized in that Before deleting the first type of video frames in the training video, the method further includes: Extracting image features of each video frame based on the feature extraction subnetwork in the third network; Inputting the image features of each of the video frames into the long short-term memory subnetwork in the third network to determine the importance corresponding to each of the video frames, wherein the importance represents the degree of semantic influence of each video frame on the training video; According to the importance corresponding to each of the video frames, a first type of video frame among the video frames is determined.
12. A training device for a video processing network, It is characterized in that include: A network training unit, the network training unit is used to: perform multiple rounds of iterative training on the first network and the second network based on the training video, and determine the first network output from the last round of iterative training as the target video processing network; The first network and the second network are networks with the same network structure and different network parameters, wherein one round of iterative training includes: Inputting a basic video obtained based on the training video into the current first network to obtain corresponding basic features, and inputting a candidate reference video obtained based on the training video into the current second network to obtain corresponding candidate reference features; Based on the candidate reference features and at least one historical reference feature, a target reference feature set is determined, and according to the similarity between each target reference feature in the target reference feature set and the basic feature and the corresponding contribution value, the first influence value corresponding to each target reference feature is determined, wherein the contribution value represents the influence of the corresponding target reference feature on the one round of iterative training, and the contribution value is obtained based on the extraction time of the corresponding target reference feature; when the one round of iterative training is the first round of iterative training in the multiple rounds of iterative training, the historical reference feature is obtained based on the current second network, and when the one round of iterative training is not the first round of iterative training in the multiple rounds of iterative training, the historical reference feature is obtained based on the historical second network; According to the first influence values corresponding to the respective target reference features, the parameters of the current first network are adjusted to obtain the first network output by the one-round iterative training; and based on the first network output by the one-round iterative training, the parameters of the current second network are adjusted.
13. A computer program product, It is characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method according to any one of claims 1 to 11.
14. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the program, the method according to any one of claims 1 to 11 is implemented.
15. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores computer instructions, and when the computer instructions are executed on a computer, the computer is caused to execute the method according to any one of claims 1 to 11.