Video data processing method and device, storage medium and electronic equipment
Through the training method of the data processing model, the poor picture quality caused by video data processing is solved, and the automatic optimization and detailed enhancement of video quality is achieved.
Patent Information
- Application Number
- CN202510397722.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-03-28
AI Technical Summary
In the prior art, the poor picture quality caused by the processing of video data before playing.
Provide a training method for data processing models, by obtaining the set of images to be trained, converting them to the brightness chromaticity space, selecting the brightness channel image as the training sample, and iteratively training in the data processing model. The loss function includes at least a negative optimization portion determined by the difference between the output image and the input image of the current iteration, and the difference between the portion and the standard image.
It can automatically detect areas with insufficient training optimization, enhance processing, ensure that the detailed areas can be fully enhanced during the model iterative training process, and improve the clarity of video quality.
Smart Images

Figure CN119919862A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of digital media technology, and more particularly to a video data processing method and device. The present application also relates to a data processing model training method and device, as well as a computer storage medium and electronic device. Background Art
[0002] With the popularization of smart phones and mobile networks, video, as an emerging form of communication, has become an important form of entertainment in daily life.
[0003] From uploading to presenting on the user side, videos need to be edited, encoded / decoded, transcoded, compressed, and other related processes. As users' requirements for video viewing experience continue to increase, high-definition and ultra-high-definition videos have become mainstream. During the uploading and playback process, the video quality will be finely processed. For example, on the one hand, by adopting video compression technology, the size of the video file can be reduced as much as possible while maintaining the image quality, thereby reducing storage and transmission costs. On the other hand, according to the device and network environment, the video version or format suitable for the current conditions is selected for playback to ensure the viewing experience. Summary of the invention
[0004] The present application provides a training method for a data processing model to solve the problem in the prior art that video data has poor playback picture quality due to processing before playback.
[0005] The present application provides a method for training a data processing model, comprising: Obtain the image set to be trained; Convert the images in the image set to be trained into a luminance and chrominance space, and select luminance channel images as training samples; The training samples are input into a data processing model for iterative training, wherein the loss function of the iterative training includes at least a component determined in the following manner: Determining a negatively optimized portion of the output image according to a difference between an output image of a current iteration and an input image; The difference between the image including the negative optimization part and the standard image corresponding to the input image is determined as one of the components of the loss function.
[0006] In some embodiments, determining the negative optimization portion of the output image according to the difference between the output image of the current iteration and the input image comprises: Comparing the output image with the input image pixel by pixel, marking a region where the pixel contrast of the output image is smaller than the pixel contrast of the input image with a first mark, and marking a region where the contrast of the output image is greater than or equal to the contrast of the input image with a second mark; determining an image including the first mark and the second mark as a mask image; The area of the first mark in the mask image is determined as a negative optimization part of the output image.
[0007] In some embodiments, determining the difference between the image including the negative optimization part and the standard image corresponding to the input image as one of the components of the loss function comprises: Determine a first pixel value of the negative optimized part in the standard image by comparing the image of the negative optimized part with the standard image pixel by pixel; Determine a second pixel value of the negative optimized portion in the output image by comparing the image of the negative optimized portion with the output image pixel by pixel; An average value of the difference between the first pixel value and the second pixel value is determined as a pixel difference component of the loss function.
[0008] In some embodiments, determining the difference between the image including the negative optimization part and the standard image corresponding to the input image as one of the components of the loss function comprises: Determine a first eigenvalue of the negative optimized portion in the output image by performing feature matching between the image of the negative optimized portion and the output image; Determine a second eigenvalue of the negative optimized part in the standard image by performing feature matching on the image of the negative optimized part and the standard image; An average value of the difference between the first eigenvalue and the second eigenvalue is determined as a feature difference component of the loss function.
[0009] In some embodiments, determining the difference between the image including the negative optimization part and the standard image corresponding to the input image as one of the components of the loss function comprises: Determine a first pixel value of the negative optimized part in the standard image by comparing the image of the negative optimized part with the standard image pixel by pixel; Determine a second pixel value of the negative optimized portion in the output image by comparing the image of the negative optimized portion with the output image pixel by pixel; Determine an average value of the difference between the first pixel value and the second pixel value as a pixel difference component of the loss function; Determine a first eigenvalue of the negative optimized portion in the output image by performing feature matching between the image of the negative optimized portion and the output image; Determine a second eigenvalue of the negative optimized part in the standard image by performing feature matching on the image of the negative optimized part and the standard image; Determine the difference between the first eigenvalue and the second eigenvalue as a characteristic difference component of the loss function; The pixel difference component and the feature difference component are respectively determined as one of the components of the loss function.
[0010] In some embodiments, it also includes: When the component of the loss function is the pixel difference component and / or the feature difference component, the weight proportion of the pixel difference component and / or the feature difference component is greater than the weight proportion of the third component in the loss function.
[0011] In some embodiments, it also includes: The weights of the pixel difference component and the feature difference component are determined according to the proportions of the pixel difference component and the feature difference component in the standard image respectively.
[0012] In some embodiments, when the image set to be trained is derived from reference video data, the step of obtaining the image set to be trained includes: Performing a first degradation process on the reference video data, and determining the video data after the first degradation process as the first video data; performing a second degradation process on the frame image in the first video data, and converting the frame image after the second degradation process into second video data; According to the second video data, acquiring the image set to be trained; When the image set to be trained is derived from a reference image set, the step of obtaining the image set to be trained includes: Performing degradation processing on the reference image set, and determining the image set after the degradation processing as the degraded image set; The to-be-trained image set is acquired according to the degraded image set.
[0013] In some embodiments, determining the to-be-trained image set according to the second video data includes: Randomly extracting a frame image from the second video data; The frame image is determined as the image set to be trained.
[0014] The present application also provides a data processing model training device, comprising: An acquisition unit, used for acquiring an image set to be trained; A conversion unit, used for converting images in the image set to be trained into a luminance and chrominance space, and selecting a luminance channel image as a training sample; A training unit is used to input the training samples into a data processing model for iterative training, wherein the loss function of the iterative training includes at least a component determined in the following manner: A first determining subunit, configured to determine a negative optimization portion of the output image according to a difference between an output image of a current iteration and an input image; The second determining subunit is used to determine the difference between the image including the negative optimization part and the standard image corresponding to the input image as one of the components of the loss function.
[0015] The present application also provides a video data processing method, comprising: Read the video data to be processed frame by frame to determine the video frame image to be processed; Performing color space conversion on the video frame image to be processed, and selecting a brightness channel image as an input frame image; Inputting the input frame image into the model trained according to the training method of the data processing model for processing to obtain an output frame image; Merging channels of the output frame image, and converting the channel-merged image into an image of three primary color channels through the color space; The images are merged into a video sequence according to the video frame sequence of the video data to be processed to obtain target video data.
[0016] The present application also provides a video data processing device, comprising: A determination unit, used to read the video data to be processed frame by frame, and determine the video frame image to be processed; A first conversion unit, configured to perform color space conversion on the video frame image to be processed, and select a brightness channel image as an input frame image; A processing unit, used for inputting the input frame image into a model trained according to the training method of the data processing model for processing to obtain an output frame image; A second conversion unit is used to perform channel merging on the output frame image, and convert the channel merging image into an image of three primary color channels through the color space; The obtaining unit is used to merge the images into a video sequence according to the video frame sequence of the video data to be processed to obtain the target video data.
[0017] The present application also provides a computer storage medium, characterized in that it includes a computer program, and when the computer program is run on an electronic device, the electronic device executes the steps in the training method of the data processing model as described above; or, executes the steps in the video data processing method as described above.
[0018] The present application also provides an electronic device, including: processor; A memory for storing a program for processing data generated by an electronic device, wherein when the program is read and executed by the processor, the program executes the steps in the training method of the data processing model as described above; or, the program executes the steps in the video data processing method as described above.
[0019] Compared with the prior art, this application has the following advantages: The present application provides a data processing model training method, which can determine the difference between the input image and the output image by comparing the two during the iterative training of the data processing model, and determine the part of the output image that needs to be optimized during the iterative process, that is, the negative optimization part during the iterative process based on the difference. The difference obtained by comparing the image including the negative optimization part with the standard image corresponding to the input image is determined as one of the components of the loss function, so that the areas where the training optimization is insufficient can be automatically detected during the iterative training process, so as to facilitate the enhancement processing of these areas, thereby ensuring that the detail areas that are easily overlooked during the iterative training process of the model can be fully enhanced.
[0020] In the training method of a data processing model provided in the present application, various image quality defect problems can be simulated to obtain a rich and diverse set of images to be trained, and complex image quality defect problems can be trained through a separate video data processing model. Since the training link is simple, the training computing cost and computing time can be effectively reduced, and there is no need to use multiple detection models to detect image quality defect types in order to deal with complex image quality problems.
[0021] A video data processing method provided in the present application can obtain a frame image with clearer image quality by inputting the frame image in the video data into a trained data processing model, especially for the details, edges and other areas of the frame image, which can achieve a more natural and significant image quality enhancement, and the processing process link is simple, the computing power requirement is low, and the processing time is short. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a flowchart of a method for training a data processing model provided in this application.
[0023] Figure 2 It is a schematic diagram of an upsampling and downsampling embodiment in a training method for a data processing model provided in the present application.
[0024] Figure 3 It is a schematic diagram of an embodiment of determining difference loss data in a model process in a training method of a data processing model provided in the present application.
[0025] Figure 4It is a structural schematic diagram of a training device for a data processing model provided in this application.
[0026] Figure 5 It is a flow chart of a video data processing method provided by the present application.
[0027] Figure 6 It is a structural schematic diagram of a video data processing device provided by the present application.
[0028] Figure 7 It is a structural schematic diagram of an electronic device provided by this application. DETAILED DESCRIPTION
[0029] Many specific details are described in the following description to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of the present application, so the present application is not limited by the specific implementation disclosed below.
[0030] The terms used in this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The descriptions used in this application and the appended claims, such as "a", "first", and "second", are not limitations on quantity or sequence, but are used to distinguish the same type of information from each other.
[0031] Based on the above background technology, it can be known that the invention of this application is derived from the requirements for the viewing experience of video data. In the field of digital media technology, video is widely used in various scenarios as a means of information dissemination. From uploading to presentation, the video usually needs to go through video editing, video encoding and decoding and other related processing, aiming to enhance the viewing and practicality of the video. In the above processing, there will be processing flows that affect the image quality, and the video will be transmitted multiple times. During this process, the video is compressed multiple times, resulting in a decrease in the quality of the video, resulting in a poor viewing experience for users.
[0032] Although there are some technical means to solve the problem of poor image quality in the prior art, there are still problems such as low processing efficiency, high computing cost, and complex processing links. The present application provides a training method for a data processing model, which can improve the image clarity in practical applications, especially the detail areas that are easily overlooked, and can achieve image quality enhancement without requiring high computing power.
[0033] The following describes a training method for a data processing model provided in this application.
[0034] like Figure 1 As shown, Figure 1 It is a flow chart of a training method of a data processing model provided by the present application, the method comprising: Step S101: obtaining a set of images to be trained; Step S102: converting the images in the image set to be trained into a luminance and chrominance space, and selecting a luminance channel image as a training sample; Step S103: inputting the training samples into a data processing model for iterative training, wherein the loss function of the iterative training includes at least components determined in the following manner: Step S103-1: determining a negative optimization portion of the output image according to a difference between the output image of the current iteration and the input image; Step S103 - 2 : determining the difference between the image including the negatively optimized part and the standard image corresponding to the input image as one of the components of the loss function.
[0035] Before describing the above steps S101 to S103 in detail, a general description of the relevant technical terms involved in the technical solution of the present application is first given.
[0036] Deep learning: An AI technique that uses multi-layered neural networks to perform complex data processing and learning.
[0037] CNN: short for Convolutional Neural Network, is a neural network model in deep learning algorithms. CNN is an artificial neural network that simulates the human visual perception mechanism. It extracts key features from data through stacked convolution, pooling and other operations to complete classification, recognition or prediction tasks. Its core lies in the convolution operation, which extracts local features by sliding a filter (or convolution kernel) on the input data.
[0038] Perceptual loss: A loss function used in computer vision and image processing that aims to evaluate the similarity between two images in the feature domain to improve visual perception.
[0039] GAN loss: A loss function in a generative adversarial network (GAN) that evaluates the difference between generated images and real images.
[0040] VGG: A classic convolutional neural network structure.
[0041] Degradation operation: refers to the operation of reducing the quality of a high-quality image to a low-quality image.
[0042] Real-basicVSR: A video super-resolution solution that provides a complete set of degradation solutions for simulating low-quality video frames in real scenes. In order to improve the generalization performance and training efficiency of the model, Real-BasicVSR introduces a random degradation mechanism. This mechanism generates different combinations of degradation (such as Gaussian blur, Poisson noise, JPEG compression, etc.) for supervised training, so that the model can be generalized to real scenes.
[0043] Crf (Constant Rate Factor): A method of allocating bitrate in ffmpeg encoder, which means to ensure "certain quality" and intelligently allocate bitrate, including allocating bitrate within the same frame and allocating bitrate between frames.
[0044] Regarding step S101: obtaining a set of images to be trained.
[0045] The purpose of step S101 is to obtain a set of images to be trained before training a data processing model (for example, a neural network model, in this embodiment, a convolutional neural network model is used as an example for explanation). In this embodiment, the set of images to be trained can be obtained through video data or image data. In this embodiment, the images in the set of images to be trained are images with quality problems, such as unclear images, etc. Correspondingly, a reference image or a standard image corresponding to the image with quality problems in the training image set is required during the training process. Therefore, an image with quality problems can be obtained by simulating multiple quality defects on the reference image. When the set of images to be trained is obtained from video data, the reference data is reference video data (i.e., a video with clear quality). When the set of images to be trained is obtained from image data, the reference data (also referred to as standard data) is a standard image. Therefore, the specific implementation process of step S102 may include two methods.
[0046] In a first approach, when the image set to be trained is derived from reference video data, obtaining the image set to be trained includes: Step S101-11: performing a first degradation process on the reference video data, and determining the video data after the first degradation process as the first video data; Step S101-12: performing a second degradation process on the frame image in the first video data, and converting the frame image after the second degradation process into second video data; Step S101 - 13: Determine the image set to be trained according to the second video data.
[0047] In this embodiment, the first degradation processing may be compression, scaling, blurring, etc. of the reference video data, and the processed video data is the first video data. The second degradation processing may be degradation operations such as adding noise, sharpening, white edges, etc. to the frame image; of course, the reference video data may be subjected to degradation operations such as adding noise, etc., so as to obtain video data with reduced image quality.
[0048] The specific implementation process of step S101-13 may include: randomly extracting a frame image from the second video data; and determining the frame image as the image set to be trained. The randomly extracted frame image has a corresponding relationship with the reference image (i.e., the standard image) extracted from the reference video, and the two may be in the form of a data pair.
[0049] It can be understood that in the present embodiment, during the degradation processing of the reference video data and / or frame image, the parameters of the related degradation operations involved can also be randomly set, the degradation operation can also be performed multiple times, and the specific operation or process of simulating the image quality problem can be set according to the needs. The above is only an example.
[0050] In a second method, when the image set to be trained is derived from a reference image set, the step of obtaining the image set to be trained includes: Step S101-21: performing degradation processing on the reference image set, and determining the image set after the degradation processing as the degraded image set; Step S101 - 22 : acquiring the to-be-trained image set according to the degraded image set.
[0051] In the second method, the degradation processing of the reference image set may also include using one or more operations such as compression, noise, white edge, sharpening, etc. to reduce the image quality. The parameters of the degradation operation may also be randomly selected, and the degradation operation may also be performed multiple times. The same image may be subjected to multiple different or the same degradation operations. The specific degradation operation process may be combined with actual needs and is not limited to the above examples.
[0052] In addition, when the generation of image quality problems of different video data is related to the type of video data generation, the following operations can also be used in some other implementation methods: Determining, according to the generation type of the video data, a method of simulating multiple image quality defects on the reference video data; Performing degradation processing on the reference video data according to the simulation method, and determining the frame image after the degradation processing as the simulated frame image; The training image set is determined according to the simulated frame images.
[0053] The generation types of the video data may include: original self-shot videos (i.e. original videos shot and produced by users themselves), reposted videos (i.e. videos reposted between platforms), original edited videos (i.e. videos edited based on non-original videos and with original audio), etc. Videos of different generation types may have the same or different image quality defects, for example: original self-shot videos may have shooting technical problems such as out-of-focus blur and camera noise during the shooting process; reposted videos may have compression distortion problems after multiple transcodings on multiple platforms; original edited videos may have problems caused by shooting, problems caused by compression, and problems such as blur and oversharpness introduced during the editing process. Of course, the above three video types are only examples, and in fact, the video generation types are not limited to the above examples. The above examples are only used to illustrate the causes of image quality defects, that is, different video generation types have the same or different image quality defects, whether it is due to image quality defects in the video generation type or image quality types caused by other reasons, they can all be included in the scope of simulation. Image quality defects include the above-mentioned out-of-focus blur, camera noise, compression distortion, editing blur, editing over-sharpness, etc., so as to determine the simulation method of image quality defects that can be performed on the reference video data. Therefore, after performing a degradation operation on the video data, a degradation operation can be performed on the image in the video after the degradation operation to obtain video data with image quality defects.
[0054] The degradation processing or degradation operation in this embodiment can be understood as one or more operations of scaling, blurring, adding noise, adding sharp white edges, compression, etc., which can be respectively operated on the video and the image in the video. In this embodiment, the degradation operation processing can be performed multiple times, randomly ordered, and with random parameters to obtain an image with reduced image quality and image defects after the degradation processing.
[0055] In another implementation scheme, the degradation operation can also be based on the identification of the generation type of the reference video data for degradation processing, for example: when the reference video data is an original self-shot video (i.e., an original video shot and produced by the user himself), blurring, noise and other degradation operations can be performed; when the reference video data is a reprinted video (i.e., a video reprinted between platforms), compression degradation operations can be performed; when the reference video data is an original edited video (i.e., a video edited based on a non-original video and with original audio), blurring, noise, compression, sharpening white edges and other degradation operations can be performed. Of course, the video type is not limited to the above examples, and the specific degradation operation can also be selected according to the video data processing requirements. The purpose of the degradation operation is to obtain an image with quality problems, which is a training image that needs to be trained in this embodiment.
[0056] It is understandable that, in order to improve the generalization of the data processing model, this embodiment may not consider the simulation of image quality defects achieved by the generation type of reference video data, but it does not exclude the possibility of this method being implemented in a specific scenario.
[0057] It should be noted that other possible implementation methods of simulating various image quality defects on reference video data or reference image data include: determining the simulation method of image quality defects through viewing feedback data of video data or image data, and then implementing corresponding processing of image quality defects.
[0058] Based on the above-mentioned degradation processing of the reference video data or reference image data, different image quality defects in the video data or image data in the simulated application scenario can be obtained, so that the image set to be trained can include a variety of rich and complex image quality defect information without relying on additional image quality detection operations, thereby providing a basis for the accuracy of subsequent model training, and also providing a basis for improving the effect and efficiency of video data or image data processing in actual application scenarios.
[0059] When training the convolutional neural network model, the reference image and the image in the image set to be trained are images with a corresponding relationship. The corresponding relationship can be that the two images are completely identical in expression content, and of course, the local content can also be the same. For example, when the training image is mainly trained for the edge, the reference image and the training image can be images with the same edge.
[0060] Based on the above, step S101 can be understood as constructing a set of images to be trained, that is, a data set with image quality defects. The reference image can be understood as an image whose quality meets the requirements (such as a high-quality image). The reference image is used as the standard answer, and the images in the set of images to be trained are used as images that need to be trained and processed.
[0061] Regarding step S102: converting the images in the to-be-trained image set into the luminance and chrominance space, and selecting the luminance channel image as a training sample.
[0062] The purpose of step S102 is to pre-process the images in the image set to be trained.
[0063] In this embodiment, the preprocessing may be color space conversion of the image in the image set to be trained. In this embodiment, the color space conversion may be converting the image in the image set to be trained from RGB color channels to YUV color channels, and splitting the converted YUV color channels into Y channel, U channel and V channel. Among them, Y represents Luma, U and V represent Chroma, which represent blue difference and red difference respectively; RGB color includes three primary colors of red, green and blue, R represents red, G represents green, and B represents blue.
[0064] After the channel conversion, the pixel values of the image are normalized, that is, the pixel range is converted from integers between 0-255 to floating-point numbers between 0-1, which is beneficial to improve calculation accuracy, speed up model training, enhance model performance, etc. Correspondingly, the brightness channel image after the convolutional neural network model training needs to be denormalized, that is, the pixel range is converted from floating-point numbers between 0-1 to integers between 0-255; then the channels are merged, that is, the enhanced Y channel, the U channel after color space conversion, and the V channel after color space conversion are merged, and the color space conversion is performed again, that is, the YUV channel is converted into an RGB channel, so as to obtain the output image.
[0065] There are many ways to convert color space, such as converting through a conversion formula, converting through image processing software, or converting through machine learning. Color space conversion is a prior art, and examples are not given here one by one.
[0066] In order to avoid the problem of color deviation in the iterative processing of the image after conversion, in this embodiment, the Y channel image, that is, the brightness channel image, is used as a training sample of the convolutional neural network model.
[0067] Regarding the step S103: inputting the training sample into a convolutional neural network model for iterative training; The purpose of step S103 is to iteratively train the model according to the training sample images, and the loss function will be calculated during the iterative training process. In this embodiment, the components in the iterative training loss function can be implemented using steps S103-1 to S103-2. The iterative process in step S103 is first described below.
[0068] like Figure 2 As shown, Figure 2 It is a schematic diagram of an upsampling and downsampling embodiment in a training method for a video data processing model provided in the present application.
[0069] Step S103-a: downsampling the training samples according to the set downsampling requirements, determining downsampled images corresponding to the downsampling requirements in sequence, and extracting downsampled image feature data from the downsampled images; Step S103-b: upsampling the downsampled image according to the upsampling requirement corresponding to the downsampled requirement, determining an upsampled image corresponding to the upsampling requirement, and merging the upsampled image feature data extracted from the upsampled image with the matched downsampled image feature to determine a merged image; Step S103-c: Determine the merged image as the output image of the convolutional neural network model.
[0070] Among them, the purpose of the step S103-a is to downsample the training image. In this embodiment, the downsampling can be completed by an encoder in a CNN network model (convolutional neural network model). The encoder is on the left side of the CNN network, which is responsible for extracting image feature information at different resolutions. In this embodiment, two 3×3 convolutional layers are used as an example for image feature extraction, and the extraction process is to extract image features from different downsampled images. The downsampling requirements in this embodiment can be sampling resolution requirements, for example: downsampling from the original resolution to one-half (1 / 2), one-quarter (1 / 4), and one-eighth (1 / 8) resolutions in turn. Of course, it is not limited to the sampling resolution requirements, for example, it can also be based on size requirements, etc.
[0071] like Figure 2 As shown on the left side, the specific implementation process of step S103-a may include: Step S103-a1: downsample the training image according to the set first downsampling requirement, determine the first downsampled image, and extract the first downsampled image feature data; in this embodiment, the first downsampling requirement can be half of the original resolution of the training image, that is, the training image is downsampled according to the requirement of half resolution to obtain the first downsampled image, and the first downsampled image feature data is extracted from the first downsampled image.
[0072] Step S103-a2: downsample the first downsampled image according to the set second downsampled requirement, determine the second downsampled image, and extract the second downsampled image feature data; in this embodiment, the second downsampled requirement may be to downsample the first sampled image to one quarter of the resolution to obtain the second downsampled image, and extract the second downsampled image feature data from the second downsampled image.
[0073] Step S103-a3: downsample the second downsampled image according to the set third downsampled requirement, determine the third downsampled image, and extract the third downsampled image feature data; in this embodiment, the third downsampled requirement may be to downsample the second downsampled image by one eighth to obtain the third downsampled image, and extract the third downsampled image feature data from the third downsampled image.
[0074] It should be noted that downsampling refers to reducing the resolution or sampling rate of data by reducing the number of data samples. In image processing, it is used to reduce the image, reduce the amount of data or perform feature extraction. Downsampling methods include: simple averaging, averaging multiple pixel values to obtain the pixel value after downsampling. Maximum pooling, selecting the maximum value of multiple pixel values as the pixel value after downsampling, is often used for feature extraction. Minimum pooling, selecting the minimum value of multiple pixel values as the pixel value after downsampling, this method may be useful in certain specific applications. Convolution downsampling, downsampling is achieved through convolution operation, and features are extracted at the same time. In this embodiment, downsampling is achieved through convolution, but the specific method used for downsampling is not limited to this method, and other sampling methods that can be performed according to the downsampling requirements may also be included.
[0075] Regarding step S103-b: according to the upsampling requirement corresponding to the downsampling requirement, the downsampled image is upsampled, the upsampled image corresponding to the upsampling requirement is determined, and the upsampled image feature data extracted from the upsampled image is merged with the matching downsampled image feature to determine a merged image.
[0076] The purpose of step S103-b is to upsample the training image. The upsampling can be completed by the decoder in the CNN network model. The encoder is on the right side of the CNN network, which is responsible for restoring images of different resolutions and extracting image feature information. Similar to downsampling, in this embodiment, upsampling and image feature extraction are performed using two 3×3 convolutional layers as an example, and the extraction process is to extract image features from different upsampled images. The upsampling requirements in this embodiment correspond to the downsampling requirements, that is, upsampling from 1 / 8 to 1 / 4, 1 / 2, and original resolution in sequence.
[0077] The specific implementation process of step S103 - b includes at least two implementation methods.
[0078] Method 1 specifically includes: Step S103-b-11: upsample the third downsampled image according to the set first upsampling requirement to determine a first upsampled image corresponding to the second downsampled image; the first upsampling requirement may be upsampling at one quarter of the resolution of the third downsampled image (one eighth of the resolution), that is, upsampling starts from the third downsampled image at a resolution of one quarter to obtain the first upsampled image with a resolution of one quarter, the sampling resolution of the second downsampled image is the same as the sampling resolution of the first upsampled image, and therefore, the first upsampled image and the second downsampled image correspond to each other.
[0079] Step S103-b-12: Merge the first up-sampled image feature data extracted from the first up-sampled image with the second down-sampled image feature data to determine first merged image feature data of a first merged image.
[0080] Step S103-b-13: Upsample the first merged image according to the set second upsampling requirement to determine a second upsampled image corresponding to the first downsampled image; in this embodiment, the second upsampling requirement may be to upsample the first merged image at half the resolution to obtain a second upsampled image, and the upsampling resolution of the second upsampled image is the same as the downsampling resolution of the first downsampled image, and therefore, the second upsampled image corresponds to the first downsampled image.
[0081] Step S103-b-14: Merge the extracted second up-sampled image feature data in the second up-sampled image with the first down-sampled image feature data to determine second merged image feature data of the second merged image.
[0082] Step S103-b-15: Upsample the second merged image according to the set third upsampling requirement to determine a third upsampled image corresponding to the training image; in this embodiment, the resolution of the third upsampling requirement corresponds to the original resolution of the training image, and therefore, the second merged image is upsampled and the third upsampled image is obtained by upsampling with the original resolution of the training image.
[0083] Step S103-b-16: Merge the extracted third up-sampled image feature data in the third up-sampled image with the training image feature data of the training image to determine a target merged image.
[0084] Method 2 specifically includes: Step S103-b-21: upsample the third downsampled image according to the set first upsampling requirement, and determine a first upsampled image corresponding to the second downsampled image; the first upsampling requirement in the second method is the same as the first upsampling requirement in the first method, and the first upsampling requirement may be upsampling according to one-fourth of the resolution of the third downsampled image (one-eighth of the resolution), that is, upsampling starts from the third downsampled image according to the requirement of one-fourth of the resolution, and obtains the first upsampled image with a resolution of one-fourth, and the sampling resolution of the second downsampled image is the same as the sampling resolution of the first upsampled image, therefore, the first upsampled image and the second downsampled image correspond to each other.
[0085] Step S103-b-22: upsample the first upsampled image according to the set second upsampling requirement to determine a second upsampled image corresponding to the first downsampled image; the second upsampling requirement in the second method is the same as the second upsampling requirement in the first method. In this embodiment, the second upsampling requirement can be upsampling the first upsampled image in the step S103-2-21 at half the resolution to obtain a second upsampled image. The upsampling resolution of the second upsampled image is the same as the downsampling resolution of the first downsampled image. Therefore, the second upsampled image corresponds to the first downsampled image.
[0086] Step S103-b-23: upsample the second upsampled image according to the set third upsampled requirement to determine a third upsampled image corresponding to the training image; the third upsampled requirement in the second method is the same as the third upsampled requirement in the first method. In this embodiment, the resolution of the third upsampled requirement corresponds to the original resolution of the training image. Therefore, the second upsampled image is upsampled with the original resolution of the training image to obtain the third upsampled image.
[0087] Step S103-b-24: merging the first up-sampled image feature data extracted from the first up-sampled image with the second down-sampled image feature data extracted from the second down-sampled image to determine first merged image feature data; Step S103-b-25: merging the extracted second up-sampled image feature data of the second up-sampled image with the extracted first down-sampled image feature data of the first down-sampled image to determine second merged image feature data; Step S103-b-26: merging the extracted third up-sampled image feature data of the third up-sampled image with the extracted training image feature data of the training image to determine third merged image feature data; Step S103-b-27: Merge the first merged image feature data, the second merged image feature data, and the third merged image feature data to determine the target merged image.
[0088] Regardless of method one or method two, the merging can be to merge the image feature data based on the number of channels in the feature space, for example: adding the number of channels of two feature maps, that is: merging the three-dimensional feature map of (number of channels C1, height H, width W) and the three-dimensional feature map of (number of channels C2, height H, width W) into a three-dimensional feature map of (C1+C2, H, W).
[0089] It should be noted that upsampling refers to increasing the resolution or sampling rate of data by increasing the number of data samples. In image processing, it is used to enlarge the image or improve the details of the image. Upsampling methods may include: nearest neighbor interpolation: select the known pixel value closest to the point to be interpolated as the interpolation result. Bilinear interpolation: According to the four known pixel values around the point to be interpolated, the interpolation result is obtained by linear interpolation. Bicubic interpolation: Considering the 16 known pixel values around the point to be interpolated, the interpolation result is obtained by cubic polynomial interpolation. Transposed convolution (also called deconvolution): This is a method of achieving upsampling through convolution operation, which is commonly used in image generation tasks in deep learning. In this embodiment, upsampling is achieved by convolution, but the specific method used for upsampling is not limited to this method, and other sampling methods that can be performed according to the upsampling requirements may also be included.
[0090] The above is a description of the process of learning and training the input image during the iteration of step S103, thereby obtaining an output image corresponding to the input image, that is, a merged image (the target merged image in the above specific embodiment). The convolutional neural network model needs to determine the loss function for the merged image (output image) and the standard image, that is, the execution of steps S103-1 and S103-2.
[0091] Regarding step S103 - 1 : according to the difference between the output image of the current iteration and the input image, determine the negative optimization part of the output image.
[0092] The purpose of step S103-1 is to determine the negative optimization part of the current iterative output image of the data processing model. The negative optimization part can be understood as: the area where the current iterative output image has image quality defects compared with the input image, for example: the clarity of the upper right corner of the output image is lower than that of the upper right corner of the input image, indicating that the iteration does not improve the image quality of the upper right corner area but reduces it, so this part of the area belongs to the negative optimization part.
[0093] In this embodiment, the input image can be the image that is first input to the data processing model, that is, the negative optimization part is determined by the difference between the first input image and the output image of the current iteration; or the output image of the previous iteration is used as the input image of the model, and the negative optimization part is determined according to the difference between the input image of the previous iteration and the output image of the current iteration, and so on.
[0094] like Figure 3 As shown, Figure 3 It is a schematic diagram of an embodiment of determining difference loss data in a model process in a training method of a data processing model provided by the present application. The specific implementation process of step S103-1 may include: Step S103-11: Compare the output image and the input image pixel by pixel, make a first mark on the area where the pixel contrast of the output image is smaller than the pixel contrast of the input image, and make a second mark on the area where the contrast of the output image is greater than or equal to the contrast of the input image; for example, the first mark is 1 and the second mark is 0; it can be shown as the following formula:
[0095] in, represents a binary mask image; the area with decreased contrast in the xth iteration compared to the x-1th iteration is set to 1, and the other areas are set to 0; when x=0, is a mask image with all zeros, Represents the weight ratio of iterative loss; Represents a standard image; Represents the output image of the current iteration.
[0096] Step S103-12: determining the image including the first mark and the second mark as a mask image; that is, obtaining a mask image including 0 and 1; Step S103 - 13 : determining the first marked area in the mask image as a negatively optimized part of the output image.
[0097] It is understandable that in this embodiment, the negative optimization part existing in the output image is compared pixel by pixel by comparing the input image and the output image, and the input image and the output image can also be compared by frequency domain analysis or image gradient information. In order to enable the convolutional neural network model to better train the negative optimization part during the iterative training process, step S103-2 can be executed.
[0098] Regarding step S103 - 2 : determining the difference between the image including the negative optimization part and the standard image corresponding to the input image as one of the components of the loss function.
[0099] The purpose of step S103-2 is to determine the components of the loss function, that is, the loss function of the convolutional neural network model in this embodiment may include multiple components of the loss function, which can be determined by the difference between the image including the negative optimization part and the standard image corresponding to the input image. The specific implementation process may include multiple methods.
[0100] Method 1: Step S103-2-11: Determine a first pixel value of the negative optimized part in the standard image by comparing the image of the negative optimized part with the standard image pixel by pixel; Step S103-2-12: Determine a second pixel value of the negative optimized part in the output image by comparing the image of the negative optimized part with the output image pixel by pixel; Step S103-2-13: Determine the average value of the difference between the first pixel value and the second pixel value as the pixel difference component of the loss function.
[0101] When the pixel difference is used as a component of the loss function, the following formula can be used:
[0102] in, represents the pixel difference loss component; represents the output image, represents a standard image, represents the image including the negative optimization part, m represents the total number of pixels, Represents the input image.
[0103] Assuming that the output image of the convolutional neural network model in the x-1th iteration is Ox-1, and the output image in the xth iteration (that is, the current iteration) is Ox, by comparing the pixel contrast of Ox-1 and Ox, it can be determined which areas of Ox have lower clarity than Ox-1, which areas of Ox have lower details than Ox-1, etc. These areas can be marked as 1, that is, if the image quality of Ox is lower than that of Ox-1, then the area is marked as 1, and if the image quality of Ox is higher than that of Ox-1, then the area is marked as 0, and then the area that needs to be optimized in the current iteration output image can be determined, and an image including the negative optimization part (diff image) can be obtained. Assuming that the areas with a value of 1 in the difference image (that is, the areas where the image quality of Ox is not as good as that of Ox-1) are the upper left corner and the lower right corner, the pixel difference loss between the upper left corner of Ox and the upper left corner of the standard image, and the pixel difference loss between the lower right corner of Ox and the lower right corner of the standard image can be calculated. During the training process, the CNN model will learn how to reduce the gap between the output image and the standard image (that is, the smaller the loss between the output image and the standard image, the higher the output image quality and the clearer the image). By comparing pixels, it can be determined that the CNN model needs to focus on the negative optimization part in the xth iteration, and calculate the pixel difference loss of the image including the negative optimization part and the standard image. By making the pixel difference loss value larger than that of other areas, the CNN model autonomously learns that the area with a large loss value is the focus area, and then improves the clarity of these negative optimization parts during the training process. Therefore, it can be ensured that the number of iterations of the CNN model is proportional to the image quality of the output image, ensuring the stability of the output results of the CNN model. The same process can also be applied to the following process related to feature difference loss.
[0104] The above is the pixel difference loss component of the output image at the current iteration determined at the pixel level.
[0105] Method 2: Step S103-2-21: determining a first eigenvalue of the negative optimized part in the output image by performing feature matching on the image of the negative optimized part and the output image; the feature matching may be a pixel-by-pixel comparison to extract image features of the negative optimized part, thereby determining the first eigenvalue of the negative optimized part in the output image; Step S103-2-22: determining a second eigenvalue of the negative optimized part in the standard image by performing feature matching on the image of the negative optimized part and the standard image; the feature matching may be a pixel-by-pixel comparison to extract the image features of the negative optimized part, thereby determining the second eigenvalue of the negative optimized part in the standard image; Step S103-2-23: Determine the average value of the difference between the first eigenvalue and the second eigenvalue as the characteristic difference component of the loss function.
[0106] In this embodiment, the feature difference component of the loss function can be expressed by the following formula:
[0107] in, represents the feature difference loss component; Represents the features in the feature space of the standard image extracted by the VGG network; Represents the features in the feature space of the output image extracted by the VGG network.
[0108] In this embodiment, the calculation process of the feature difference loss (ie, the feature difference loss of the output image) may refer to the calculation process of the pixel difference loss, which will not be described in detail here.
[0109] Method 3: Step S103-2-31: Determine a first pixel value of the negative optimized part in the standard image by comparing the image of the negative optimized part with the standard image pixel by pixel; Step S103-2-32: Determine a second pixel value of the negative optimized part in the output image by comparing the image of the negative optimized part with the output image pixel by pixel; Step S103-2-33: determining an average value of the difference between the first pixel value and the second pixel value as a pixel difference component of the loss function; Step S103-2-34: determining a first eigenvalue of the negative optimized part in the output image by performing feature matching on the image of the negative optimized part and the output image; the feature matching may be a pixel-by-pixel comparison to extract image features of the negative optimized part, thereby determining the first eigenvalue of the negative optimized part in the output image; Step S103-2-35: determining a second eigenvalue of the negative optimized part in the standard image by performing feature matching on the image of the negative optimized part and the standard image; the feature matching may be a pixel-by-pixel comparison to extract the image features of the negative optimized part, thereby determining the second eigenvalue of the negative optimized part in the standard image; Step S103-2-36: Determine the difference between the first eigenvalue and the second eigenvalue as the characteristic difference component of the loss function; Step S103-2-37: Determine the pixel difference component and the feature difference component as components of the loss function.
[0110] The third approach mainly uses both the first approach and the second approach as components of the loss function. In the first approach, the component of the loss function is a pixel difference loss component, and in the second approach, the component of the loss function is a feature difference loss component. However, both approaches can be used as components of the loss function.
[0111] In this embodiment, the components of the loss function may also include: one or more components of pixel loss, feature loss and discrimination loss, which can be obtained by comparing the output image with the standard image. The loss function can be expressed by the following formula: ; ; ; The above formulas are representations of three embodiments, and are not limited to the above representations. The components of the loss function can be adjusted according to requirements.
[0112] in, (pixel loss component), (feature loss component) and (Discrimination loss component) is also a component of the loss function, Represents weight.
[0113] About the pixel loss in the loss function The specific implementation process of the component may include: Comparing the output image with the standard image pixel by pixel to determine the pixel difference between the output image and the standard image; Based on the pixel differences, a pixel loss of the output image is determined.
[0114] In this embodiment, the pixel loss component can be expressed by the following formula: .
[0115] About the feature loss in the loss function The specific implementation process of the component may include: comparing the feature data of the output image with the feature data of the standard image to determine feature differences between the output image and the standard image in the feature data; A feature loss of the output image is determined according to the feature difference.
[0116] In this embodiment, the characteristic loss component can be expressed by the following formula: .
[0117] Regarding the discriminant loss in the loss function The specific implementation process of the component may include: Inputting the output image into a discriminator model for discrimination and determining a discrimination result, wherein the discriminator model is a model that has been trained in advance; Determining the discrimination loss according to the discrimination result; In this embodiment, the discrimination loss component can be expressed by the following formula:
[0118] in, The discriminator model can discriminate images that meet the high quality requirements of the image (high-quality images) as 1, and images that do not meet the image quality requirements (low-quality images) as 0. When training the convolutional neural network model, the output image can be measured by the discriminant loss. Is it a high-quality image? In this embodiment, the discriminator model It can be a trained model. The loss function of the discriminator model can be expressed as follows:
[0119] in, is the discriminant loss function during the discriminator model training process, The discriminator model is required to classify the standard image that meets the image quality requirements as 1. The discriminator model is required to classify the output image processed by CNN as 0 if the image quality is lower than the standard image.
[0120] It should be noted that in this embodiment The discriminant loss function is determined by discriminating the output image based on the trained discriminator model.
[0121] Based on the above content, it can be seen that when determining the loss function according to any of the above loss function components When , any component can be weighted, or each component can be weighted. The specific process may include: When the components of the loss function are the pixel difference component and / or the feature difference component, the weight proportion of the pixel difference component and / or the feature difference component is greater than the weight proportion of the third component in the loss function. The third component may be a pixel loss component, a feature loss component, a discrimination loss component, etc., that is: and / or Can be greater than , , .
[0122] The weight ratio between the pixel difference component and the feature difference component can be determined according to the ratio of the pixel difference component and the feature difference component in the output image, for example, if the ratio of the pixel difference component in the output image is greater than the ratio of the feature difference component in the output image, the weight of the pixel difference component can be set to be greater than the weight of the feature difference component, so that the CNN model pays more attention to the pixel difference area corresponding to the pixel difference component.
[0123] It is understandable that the weight of any of the above loss function components can be adjusted in real time according to the needs of model training, the input image and the output image, and is not limited to the above examples. Therefore, the weight setting can be a dynamic weight that is dynamically adjusted with the number of iterations, output results, etc.
[0124] In this embodiment, after obtaining the components of the above-mentioned loss function and thus determining the loss function value, the CNN model can be adjusted so that the image quality of the output image output by the model training process gradually approaches (or is equal to) the image quality of the standard image.
[0125] The above is a description of an embodiment of a training method for a data processing model provided by the present application. The method performs color space conversion on the images in the training image set, and inputs the converted brightness channel images as training samples into the convolutional neural network model for iterative training, thereby avoiding image color deviation during the training process and ensuring the accuracy of image color. Also, in order to avoid the situation where the image quality does not improve or even decreases with the increase in the number of iterations, causing the model to ignore the negative optimization areas, thereby causing the model output result to fail to meet higher requirements, the method can determine the negative optimization part included in the output image through the difference between the current iteration output image and the current iteration input image, determine the components of the loss function for the image including the negative optimization part and the standard image, so that the model can self-learn the areas of the negative optimization parts, thereby improving the quality of the output image and further enhancing the model's ability to retain image details.
[0126] In addition, this training method can obtain a rich and diverse set of images to be trained by simulating various image quality defect problems, and can complete the training of more complex image quality defect problems through a separate data processing model of this application. Since the training link is simple, the training computing cost and computing time can be effectively reduced, and there is no need to use multiple detection models to detect image quality defect types in order to deal with complex image quality problems.
[0127] The above is a specific description of an embodiment of a training method for a data processing model provided by the present application. Corresponding to the embodiment of a training method for a data processing model provided above, the present application also discloses an embodiment of a training device for a video data processing model. Please refer to Figure 4 Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0128] like Figure 4 As stated, Figure 4 : is a structural diagram of a training device for a data processing model provided by the present application, the device comprising: An acquisition unit 401 is used to acquire a set of images to be trained; A conversion unit 402 is used to convert the images in the to-be-trained image set into a luminance and chrominance space, and select a luminance channel image as a training sample; The training unit 403 is used to input the training samples into the convolutional neural network model for iterative training, wherein the loss function of the iterative training includes at least a component determined in the following manner: A first determining subunit 403-1, configured to determine a negative optimization part of the output image according to a difference between an output image of a current iteration and an input image of the current iteration; The second determining subunit 403-2 is used to determine the difference between the image including the negative optimization part and the standard image corresponding to the input image as one of the components of the loss function.
[0129] When the image set to be trained is derived from reference video data, the acquisition unit 401 includes: A first processing subunit, configured to perform a first degradation process on the reference video data, and determine the video data after the first degradation process as the first video data; A second processing subunit is used to perform a second degradation process on the frame image in the first video data, and convert the frame image after the second degradation process into second video data; The acquisition subunit is used to acquire the image set to be trained based on the second video data; specifically, it may include: an extraction subunit, used to randomly extract frame images from the second video data; and a determination subunit, used to determine the frame image as the image set to be trained.
[0130] When the image set to be trained is derived from a reference image set, the step of obtaining the image set to be trained includes: An exit processing subunit, configured to perform degradation processing on the reference image set, and determine the image set after the degradation processing as a degraded image set; The acquisition subunit is used to acquire the to-be-trained image set according to the degraded image set.
[0131] The first determining subunit 403-1 may include: a marking subunit, configured to compare the output image with the input image pixel by pixel, mark the area where the pixel contrast of the output image is smaller than the pixel contrast of the input image with a first mark, and mark the area where the contrast of the output image is greater than or equal to the contrast of the input image with a second mark; a mask determination subunit, configured to determine the image including the first mark and the second mark as a mask image; The negative optimization determination subunit is used to determine the area of the first mark in the mask image as the negative optimization part of the output image.
[0132] The second determining subunit 403-2 may include at least three of the following methods: Method 1 includes: A first pixel value determining subunit, configured to determine a first pixel value of the negative optimized part in the standard image by comparing the image of the negative optimized part with the standard image pixel by pixel; A second pixel value determining subunit, configured to determine a second pixel value of the negative optimized portion in the output image by comparing the image of the negative optimized portion with the output image pixel by pixel; The pixel difference determination subunit is configured to determine an average value of the difference between the first pixel value and the second pixel value as a pixel difference component of the loss function.
[0133] Method 2 includes: A first eigenvalue determining subunit, configured to determine a first eigenvalue of the negative optimized part in the output image by performing feature matching between the image of the negative optimized part and the output image; A second eigenvalue determining subunit is configured to determine a second eigenvalue of the negative optimized part in the standard image by performing feature matching on the image of the negative optimized part and the standard image; The feature difference determination subunit is used to determine the average value of the difference between the first feature value and the second feature value as the feature difference component of the loss function.
[0134] Method three includes: A first pixel value determining subunit, configured to determine a first pixel value of the negative optimized part in the standard image by comparing the image of the negative optimized part with the standard image pixel by pixel; A second pixel value determining subunit, configured to determine a second pixel value of the negative optimized portion in the output image by comparing the image of the negative optimized portion with the output image pixel by pixel; a pixel difference determination subunit, configured to determine an average value of a difference between the first pixel value and the second pixel value as a pixel difference component of the loss function; A first eigenvalue determining subunit, configured to determine a first eigenvalue of the negative optimized part in the output image by performing feature matching between the image of the negative optimized part and the output image; A second eigenvalue determining subunit is configured to determine a second eigenvalue of the negative optimized part in the standard image by performing feature matching on the image of the negative optimized part and the standard image; a feature difference determination subunit, configured to determine an average value of a difference between the first feature value and the second feature value as a feature difference component of the loss function; The component determination subunit is used to determine the pixel difference component and the feature difference component as one of the components of the loss function respectively.
[0135] The method further includes: a first weight setting unit, which is used to, when the component of the loss function is the pixel difference component and / or the feature difference component, have a weight proportion greater than a weight proportion of the third component in the loss function.
[0136] The method further includes: a second weight setting unit, configured to determine the weights of the pixel difference component and the feature difference component according to the proportions of the pixel difference component and the feature difference component in the standard image, respectively.
[0137] The above is a description of a training device for a data processing model provided in the present application. For the specific content of the device, please refer to the above-mentioned training method embodiment, which will not be described in detail here.
[0138] Based on the above, the present application also provides a video data processing method, such as Figure 5 As shown, Figure 5 This is a flow chart of a video data processing method provided by the present application, the method comprising: Step S501: Read the video data to be processed frame by frame to determine the video frame image to be processed; Step S502: performing color space conversion on the video frame image to be processed, and selecting a brightness channel image as an input frame image; Step S503: inputting the input frame image into the model trained according to the training method of the data processing model for processing to obtain an output frame image; Step S504: performing channel merging on the output frame image, and converting the channel merging image into an image of three primary color channels through the color space; Step S505: merging the images into a video sequence according to the video frame sequence of the video data to be processed to obtain target video data.
[0139] The video data processing method can retain more details, so that the target video data has a clearer and more natural picture enhancement effect.
[0140] The specific contents of the above steps S501 to S505 can refer to the relevant contents in the training method of the data processing model, which will not be described in detail here.
[0141] Based on the above content, the present application also provides a video data processing device, such as Figure 6 As shown, Figure 6 : is a structural schematic diagram of a video data processing device provided by the present application, the device comprising: A determination unit 601 is used to read the video data to be processed frame by frame to determine the video frame image to be processed; A first conversion unit 602 is used to perform color space conversion on the video frame image to be processed, and select a brightness channel image as an input frame image; A processing unit 603 is used to input the input frame image into the model trained according to the training method of the data processing model for processing to obtain an output frame image; A second conversion unit 604 is used to perform channel merging on the output frame image, and convert the channel merging image into an image of three primary color channels through the color space; The obtaining unit 605 is used to merge the images into a video sequence according to the video frame sequence of the video data to be processed to obtain target video data.
[0142] Based on the above content, the present application also provides a computer storage medium, including a computer program. When the computer program is run on an electronic device, the electronic device executes the steps in the training method of the data processing model; or, executes the steps in the video data processing method.
[0143] Based on the above content, the present application also provides an electronic device, such as Figure 7 As shown, Figure 7 The structural diagram of an electronic device provided by the present application includes: Processor 701; The memory 702 is used to store a program for processing data generated by an electronic device. When the program is read and executed by the processor, it executes the steps in the training method of the data processing model; or, it executes the steps in the video data processing method.
[0144] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0145] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0146] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0147] 1. Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include non-transitory media such as modulated data signals and carrier waves.
[0148] 2. Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0149] Although the present application is disclosed as above in the form of a preferred embodiment, it is not intended to limit the present application. Any technical personnel in this field may make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.
Claims
1. A method for training a data processing model, characterized in that: include: Obtain the image set to be trained; Convert the images in the image set to be trained into a luminance and chrominance space, and select luminance channel images as training samples; The training samples are input into a data processing model for iterative training, wherein the loss function of the iterative training includes at least a component determined in the following manner: Determining a negatively optimized portion of the output image according to a difference between an output image of a current iteration and an input image; The difference between the image including the negative optimization part and the standard image corresponding to the input image is determined as one of the components of the loss function.
2. The data processing model training method according to claim 1, characterized in that: The step of determining the negative optimization part of the output image according to the difference between the output image of the current iteration and the input image comprises: Comparing the output image with the input image pixel by pixel, marking a region where the pixel contrast of the output image is smaller than the pixel contrast of the input image with a first mark, and marking a region where the pixel contrast of the output image is greater than or equal to the contrast of the input image with a second mark; determining an image including the first mark and the second mark as a mask image; The area of the first mark in the mask image is determined as a negative optimization part of the output image.
3. The training method of the data processing model according to claim 1, characterized in that: Determining the difference between the image including the negative optimization part and the standard image corresponding to the input image as one of the components of the loss function comprises: Determine a first pixel value of the negative optimized part in the standard image by comparing the image of the negative optimized part with the standard image pixel by pixel; Determine a second pixel value of the negative optimized portion in the output image by comparing the image of the negative optimized portion with the output image pixel by pixel; An average value of the difference between the first pixel value and the second pixel value is determined as a pixel difference component of the loss function.
4. The training method of the data processing model according to claim 1, characterized in that: Determining the difference between the image including the negative optimization part and the standard image corresponding to the input image as one of the components of the loss function comprises: Determine a first eigenvalue of the negative optimized portion in the output image by performing feature matching between the image of the negative optimized portion and the output image; Determine a second eigenvalue of the negative optimized part in the standard image by performing feature matching on the image of the negative optimized part and the standard image; An average value of the difference between the first eigenvalue and the second eigenvalue is determined as a feature difference component of the loss function.
5. The method for training a data processing model according to claim 1, characterized in that: Determining the difference between the image including the negative optimization part and the standard image corresponding to the input image as one of the components of the loss function comprises: Determine a first pixel value of the negative optimized part in the standard image by comparing the image of the negative optimized part with the standard image pixel by pixel; Determine a second pixel value of the negative optimized portion in the output image by comparing the image of the negative optimized portion with the output image pixel by pixel; Determine an average value of the difference between the first pixel value and the second pixel value as a pixel difference component of the loss function; Determine a first eigenvalue of the negative optimized portion in the output image by performing feature matching between the image of the negative optimized portion and the output image; Determine a second eigenvalue of the negative optimized part in the standard image by performing feature matching on the image of the negative optimized part and the standard image; Determine the difference between the first eigenvalue and the second eigenvalue as a characteristic difference component of the loss function; The pixel difference component and the feature difference component are respectively determined as one of the components of the loss function.
6. The training method of the data processing model according to claim 4 or 5, characterized in that: Also includes: When the component of the loss function is the pixel difference component and / or the feature difference component, the weight proportion of the pixel difference component and / or the feature difference component is greater than the weight proportion of the third component in the loss function.
7. The training method of the data processing model according to claim 4 or 5, characterized in that: Also includes: The weights of the pixel difference component and the feature difference component are determined according to the proportions of the pixel difference component and the feature difference component in the standard image respectively.
8. The method for training a data processing model according to claim 1, characterized in that: When the image set to be trained is derived from reference video data, the step of obtaining the image set to be trained includes: Performing a first degradation process on the reference video data, and determining the video data after the first degradation process as the first video data; performing a second degradation process on the frame image in the first video data, and converting the frame image after the second degradation process into second video data; According to the second video data, acquiring the image set to be trained; When the image set to be trained is derived from a reference image set, the step of obtaining the image set to be trained includes: Performing degradation processing on the reference image set, and determining the image set after the degradation processing as the degraded image set; The to-be-trained image set is acquired according to the degraded image set.
9. The method for training a data processing model according to claim 8, characterized in that: The step of determining the image set to be trained according to the second video data includes: Randomly extracting a frame image from the second video data; The frame image is determined as the image set to be trained.
10. A training device for a data processing model, characterized in that: include: An acquisition unit, used for acquiring an image set to be trained; A conversion unit, used for converting images in the image set to be trained into a luminance and chrominance space, and selecting a luminance channel image as a training sample; A training unit is used to input the training samples into a data processing model for iterative training, wherein the loss function of the iterative training includes at least a component determined in the following manner: A first determining subunit, configured to determine a negative optimization portion of the output image according to a difference between an output image of a current iteration and an input image; The second determining subunit is used to determine the difference between the image including the negative optimization part and the standard image corresponding to the input image as one of the components of the loss function.
11. A video data processing method, characterized in that: include: Read the video data to be processed frame by frame to determine the video frame image to be processed; Performing color space conversion on the video frame image to be processed, and selecting a brightness channel image as an input frame image; Inputting the input frame image into a model trained by the data processing model training method according to any one of claims 1 to 9 for processing to obtain an output frame image; Merging channels of the output frame image, and converting the channel-merged image into an image of three primary color channels through the color space; The images are merged into a video sequence according to the video frame sequence of the video data to be processed to obtain target video data.
12. A video data processing device, characterized in that: include: A determination unit, used to read the video data to be processed frame by frame, and determine the video frame image to be processed; A first conversion unit, configured to perform color space conversion on the video frame image to be processed, and select a brightness channel image as an input frame image; A processing unit, used for inputting the input frame image into a model trained by the training method for a data processing model according to any one of claims 1 to 9 for processing to obtain an output frame image; A second conversion unit is used to perform channel merging on the output frame image, and convert the channel merging image into an image of three primary color channels through the color space; The obtaining unit is used to merge the images into a video sequence according to the video frame sequence of the video data to be processed to obtain the target video data.
13. A computer storage medium, characterized in that: The method comprises a computer program, which, when executed on an electronic device, causes the electronic device to execute the steps in the method for training a data processing model according to any one of claims 1 to 9; Alternatively, execute the steps in the video data processing method as claimed in claim 11.
14. An electronic device, characterized in that: include: processor; A memory, for storing a program for processing data generated by an electronic device, wherein when the program is read and executed by the processor, the program executes the steps in the training method of the data processing model according to any one of claims 1 to 9; Alternatively, execute the steps in the video data processing method as claimed in claim 11.
Citation Information
Patent Citations
A joint image restoration and matching method and system based on linear features
CN109903233A
Image stereo matching method and system based on deformable convolutional network
CN112598722A
Video super-resolution processing method and device, and storage medium
CN112700392A
Image super-resolution reconstruction method and device for real-time video stream
CN112785506A
Image correction model training method and device, image correction method and device and computer equipment
CN115115552A