Image enhancement method based on video IP frame

By analyzing and processing the types of video frames in video IP frames and acquiring the transformation matrix using neural network models, the problem of slow video image enhancement speed in the prior art is solved, and more efficient image enhancement processing is achieved.

CN119941591AActive Publication Date: 2025-05-06沐曦科技(成都)有限公司
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202311458810.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-02
Publication Date
2025-05-06
Estimated Expiration
2043-11-02

AI Technical Summary

Technical Problem

In the prior art, the CPU-based image enhancement algorithm is slow to enhance video, making it difficult to meet users' high efficiency requirements for video image enhancement.

Method used

Using an image enhancement method based on video IP frames, the image enhancement process is performed in the order of frame count from small to large by analyzing the keyframes and forward-predicted encoded frames in the video, and the image enhancement process is obtained using the trained target neural network model to accelerate the image enhancement process.

Benefits of technology

By obtaining the conversion matrix of the keyframe in advance and using it as the latest conversion matrix, it is directly used for image enhancement processing of subsequent forward prediction coded frames, eliminating the process of obtaining the conversion matrix, which significantly improves the efficiency of video image enhancement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941591A_ABST
    Figure CN119941591A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image data processing, in particular to an image enhancement method based on a video IP frame. The method comprises the following steps: acquiring a target video V; analyzing the V to obtain a frame number list L corresponding to the key frames in the V; v is traversed in sequence, if q is equal to a certain idm in L, a first sub-model in the trained target neural network model is used for obtaining a conversion matrix Tq of vq, Tq is determined to be the latest conversion matrix Tnew, Tq is used as input of a second sub-model in the trained target neural network model, and if q is equal to a certain idm in L, the conversion matrix Tq is determined to be the latest conversion matrix Tnew; taking the output of a second sub-model in the trained target neural network model as an enhanced image corresponding to the vq; and otherwise, taking the Tnew as the input of a second sub-model in the trained target neural network model, and taking the output of the second sub-model in the trained target neural network model as the enhanced image corresponding to the vq. According to the invention, the efficiency of performing image enhancement on the video is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image data processing, and in particular to an image enhancement method based on video IP frames. Background Art

[0002] There are many image enhancement algorithms disclosed in the prior art. For example, the contrast-constrained adaptive histogram equalization (CLAHE) algorithm is a relatively effective image enhancement algorithm. However, the algorithm runs on the CPU and needs to occupy CPU resources. Moreover, the CPU is a serial processing mechanism. The speed of running the CLAHE algorithm on the CPU to enhance the video is slow, which is difficult to meet the high efficiency requirements of users for video image enhancement. How to improve the efficiency of video image enhancement is a technical problem that needs to be solved urgently. Summary of the invention

[0003] The present invention aims to provide an image enhancement method based on video IP frames to improve the efficiency of image enhancement for videos.

[0004] According to the present invention, there is provided an image enhancement method based on a video IP frame, comprising the following steps:

[0005] S10, obtain the target video V, V=(v1,v2,…,v q ,…,v Q ), v q is the qth frame image included in V, the value range of q is 1 to Q, Q is the number of images included in V, v q v q+1 The previous frame image, v q+1 is the q+1th frame image included in V; v q A key frame or a forward predictive coded frame in V.

[0006] S20, parse V and obtain a list of frame numbers L corresponding to the key frames in V, where L = (id1, id2, ..., id m ,…,id M ), id m is the number of frames corresponding to the mth keyframe in V, the value range of m is 1 to M, and M is the number of keyframes included in V.

[0007] S30, traverse V in order, if q matches an id in L m If they are equal, then go to S40; otherwise, go to S50.

[0008] S40, using the first sub-model in the trained target neural network model to obtain v q The transformation matrix T q , and T qDetermine the latest transformation matrix T new .

[0009] S41, T q As the input of the second sub-model in the trained target neural network model, the output of the second sub-model in the trained target neural network model is used as v q The corresponding enhanced image.

[0010] S50, Get T new .

[0011] S51, T new As the input of the second sub-model in the trained target neural network model, the output of the second sub-model in the trained target neural network model is used as v q The corresponding enhanced image.

[0012] Compared with the prior art, the present invention has at least the following beneficial effects:

[0013] When performing image enhancement processing on each frame image in a target video, the present invention parses the type of each frame in the target video and obtains a list of frame numbers corresponding to key frames in the target video; on this basis, the present invention performs image enhancement processing in the order of the frame numbers of each frame in the target video from small to large, thereby first obtaining a conversion matrix corresponding to the key frame and determining it as the latest conversion matrix, and when subsequently performing image enhancement on a forward prediction coding frame, the conversion matrix corresponding to the corresponding key frame (i.e., the latest conversion matrix) is a known value and can be directly used as an input of a second sub-model, thereby, in the process of performing image enhancement processing on a forward prediction coding frame, the process of obtaining the conversion matrix corresponding to the forward prediction coding frame using the first sub-model can be omitted, which can save time for obtaining an enhanced image, thereby improving the efficiency of image enhancement on the video. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0015] Figure 1 The present invention provides a flowchart of an image enhancement method based on a video IP frame according to an embodiment of the present invention. DETAILED DESCRIPTION

[0016] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0017] Embodiment 1

[0018] According to this embodiment, Figure 1 As shown, a method for image enhancement based on video IP frames is provided, comprising the following steps:

[0019] S10, obtain the target video V, V=(v1,v2,…,v q ,…,v Q ), v q is the qth frame image included in V, the value range of q is 1 to Q, Q is the number of images included in V, v q v q+1 The previous frame image, v q+1 is the q+1th frame image included in V; v q A key frame or a forward predictive coded frame in V.

[0020] In this embodiment, if v q is the first frame of the target video, then q = 1; if v q is the second frame image in the target video, then q=2; and so on. In this embodiment, the frame number of each image in the target video is determined according to the time sequence corresponding to each frame image, that is, the earlier the image corresponds to, the smaller its corresponding frame number.

[0021] S20, parse V and obtain a list of frame numbers L corresponding to the key frames in V, where L = (id1, id2, ..., id m ,…,id M ), id m is the number of frames corresponding to the mth keyframe in V, the value range of m is 1 to M, and M is the number of keyframes included in V.

[0022] Those skilled in the art know that any method of parsing key frames and forward predictive coding frames in a video in the prior art falls within the protection scope of the present invention. Optionally, ffmpeg is used to parse the target video to obtain a list of frame numbers corresponding to the key frames in the target video.

[0023] S30, traverse V in order, if q matches an id in L m If they are equal, then go to S40; otherwise, go to S50.

[0024] S40, using the first sub-model in the trained target neural network model to obtain v q The transformation matrix T q , and T q Determine the latest transformation matrix T new .

[0025] S41, T q As the input of the second sub-model in the trained target neural network model, the output of the second sub-model in the trained target neural network model is used as v q The corresponding enhanced image.

[0026] S50, Get T new .

[0027] S51, T new As the input of the second sub-model in the trained target neural network model, the output of the second sub-model in the trained target neural network model is used as v q The corresponding enhanced image.

[0028] For example, the target video includes 100 frames of images. According to the time sequence of each frame, the 1st, 10th, 20th, 30th, 40th, 50th, 60th, 70th, 80th and 90th frames are key frames, then L = (1, 10, 20, 30, 40, 50, 60, 70, 80, 90); the others are forward prediction coding frames, specifically, the 2nd to 9th frames are forward prediction coding frames corresponding to the 1st frame, the 11th to 19th frames are forward prediction coding frames corresponding to the 10th frame, the 21st to 29th frames are forward prediction coding frames corresponding to the 20th frame, and so on. If v q The number of frames in V is 10, so 10 is equal to 10 in L, then S40 is executed to obtain v using the first sub-model in the trained target neural network model. 10 The transformation matrix T 10 , and T 10 Determine the latest transformation matrix T new , execute S41; if v q The number of frames in V is 11, so 11 is not equal to the number of frames in L, so S50 is executed to obtain T new (i.e. T 10 ), execute S51. q The number of frames in V is 20, so 20 is equal to 20 in L, then S40 is executed to obtain v using the first sub-model in the trained target neural network model. 20 The transformation matrix T 20 , and T 20 Determine the latest transformation matrix T new , execute S41; if v qThe number of frames in V is 25, so 25 is not equal to the number of frames in L. Then execute S50 to obtain T new (i.e. T 20 ), execute S51.

[0029] When performing image enhancement processing on each frame image in a target video, the present invention parses the type of each frame in the target video and obtains a list of frame numbers corresponding to key frames in the target video; on this basis, the present invention performs image enhancement processing in the order of the frame numbers of each frame in the target video from small to large, thereby first obtaining a conversion matrix corresponding to the key frame and determining it as the latest conversion matrix, and when subsequently performing image enhancement on a forward prediction coding frame, the conversion matrix corresponding to the corresponding key frame (i.e., the latest conversion matrix) is a known value and can be directly used as an input of a second sub-model, thereby, in the process of performing image enhancement processing on a forward prediction coding frame, the process of obtaining the conversion matrix corresponding to the forward prediction coding frame using the first sub-model can be omitted, which can save time for obtaining an enhanced image, thereby improving the efficiency of image enhancement on the video.

[0030] Specifically, the target neural network model in this embodiment includes a first sub-model and a second sub-model, the output of the first sub-model is the input of the second sub-model; the first sub-model is used to obtain the transformation matrix of the input image, and the first sub-model includes, in order according to the direction of information transmission: four cascaded convolutional layers, a two-dimensional adaptive average pooling layer, a flatten layer, a first linear layer, a relu activation layer, a second linear layer and a sigmoid activation layer; the output of the sigmoid activation layer is the transformation matrix of the input image; the second sub-model is used to obtain an enhanced image corresponding to the input image according to the output of the first sub-model, and the second sub-model includes a convolution operation module and a summation operation module, the input of the convolution operation module is the input image and the output of the first sub-model, the input of the summation operation module is the input image and the output of the convolution operation module, and the output of the summation operation module is the enhanced image corresponding to the input image.

[0031] Specifically, each convolution layer includes a convolution processing module, a batch normalization operation processing module and a relu activation processing module which are connected in sequence.

[0032] Specifically, when the target neural network model is a neural network model for obtaining an enhanced image of a single-channel image, the convolution processing module includes 1 3×3 convolution kernel, and the transformation matrix of the input image is 1 3×3 matrix; when the target neural network model is a neural network model for obtaining an enhanced image of a three-channel image, the convolution processing module includes 3 3×3 convolution kernels, and the transformation matrix of the input image is 3 3×3 matrices, wherein the first 3×3 matrix is ​​used for convolution with the first channel image of the input image, the second 3×3 matrix is ​​used for convolution with the second channel image of the input image, and the third 3×3 matrix is ​​used for convolution with the third channel image of the input image.

[0033] As a first specific implementation manner, when the target neural network model is a neural network model for obtaining an enhanced image of a single-channel image, according to the direction of information transmission, the convolution processing module of the first convolution layer includes a convolution kernel with parameters of (1, 32, 3, 1), the first 1 indicates that the input is a single channel, the output is 32 channels, the convolution kernel size is 3×3, and the second 1 indicates that the step size is 1; the convolution processing module of the second convolution layer includes a convolution kernel with parameters of (32, 64, 3, 1), 32 indicates that the input is 32 channels, 64 indicates that the output is 64 channels, the convolution kernel size is 3×3, and 1 indicates that the step size is 1; the convolution processing module of the third convolution layer includes a convolution kernel with parameters of (32, 64, 3, 1), 32 indicates that the input is 32 channels, 64 indicates that the output is 64 channels, the convolution kernel size is 3×3, and 1 indicates that the step size is 1. The parameters are (64, 128, 3, 1), 64 means the input is 64 channels, 128 means the output is 128 channels, the convolution kernel size is 3×3, and 1 means the stride is 1; the convolution processing module of the fourth convolution layer includes the convolution kernel parameters (128, 256, 3, 1), 128 means the input is 128 channels, 256 means the output is 256 channels, the convolution kernel size is 3×3, and 1 means the stride is 1; the parameters of the first linear layer are (256, 128), 256 means the input size is 256, and 128 means the output size is 128; the parameters of the second linear layer are (128, 9), 128 means the input size is 128, and 9 means the output size is 9.

[0034] As a second specific implementation manner, when the target neural network model is a neural network model for obtaining an enhanced image of a 3-channel image, according to the direction of information transmission, the convolution processing module of the first convolution layer includes a convolution kernel with parameters of (3, 32, 3, 1), the first 3 indicates that the input is 3 channels and the output is 32 channels, the second 3 indicates that the convolution kernel size is 3×3, and 1 indicates that the step size is 1; the convolution processing module of the second convolution layer includes a convolution kernel with parameters of (32, 64, 3, 1), 32 indicates that the input is 32 channels, 64 indicates that the output is 64 channels, the convolution kernel size is 3×3, and 1 indicates that the step size is 1; the convolution processing module of the third convolution layer includes a convolution kernel with parameters of (3, 32, 3, 1), 32 indicates that the input is 32 channels, 64 indicates that the output is 64 channels, the convolution kernel size is 3×3, and 1 indicates that the step size is 1. The parameters are (64, 128, 3, 1), 64 means the input is 64 channels, 128 means the output is 128 channels, the convolution kernel size is 3×3, and 1 means the stride is 1; the parameters of the convolution kernel included in the convolution processing module of the fourth convolution layer are (128, 256, 3, 1), 128 means the input is 128 channels, 256 means the output is 256 channels, the convolution kernel size is 3×3, and 1 means the stride is 1; the parameters of the first linear layer are (256, 128), 256 means the input size is 256, and 128 means the output size is 128; the parameters of the second linear layer are (128, 27), 128 means the input size is 128, and 27 means the output size is 27.

[0035] In this embodiment, the convolution operation module is used to perform padding processing on the input image, and use the output of the first sub-model to perform a convolution operation on the input image after the padding processing, and the summation operation module is used to perform a summation operation on the input image and the convolution result output by the convolution operation module.

[0036] In this embodiment, when the target neural network model is a neural network model for obtaining an enhanced image of a single-channel image, the transformation matrix of the input image is a 3×3 matrix; when the target neural network model is a neural network model for obtaining an enhanced image of a three-channel image, the transformation matrix of the input image is three 3×3 matrices; accordingly, when padding is performed on the input image, the padding size is 1, specifically, a circle of 0s is filled outside the input image, thereby, the convolution result output by the convolution operation module is equal to the size of the input image, and a summation operation can be performed.

[0037] Specifically, the training process of the target neural network model includes the following steps:

[0038] S1, obtain the image sample list A, A=(a1, a2,…, a n ,…,a N ), a n is the nth image sample, the value range of n is 1 to N, N is the number of image samples; each an The number and type of channels are the same as those corresponding to the images in the target video.

[0039] S2, obtain the image sample label list B = (b1, b2, ..., b n ,…,b N ), b n For a n Image after image enhancement.

[0040] In this embodiment, b n for a n The corresponding label.

[0041] Those skilled in the art know that any method for enhancing an image in the prior art falls within the protection scope of the present invention; optionally, b n To use the clahe algorithm to n Image after image enhancement.

[0042] S3, use A and B to train the target neural network model.

[0043] Those skilled in the art are aware that any supervised training method in the prior art falls within the protection scope of the present invention.

[0044] Optionally, the video IP frame-based image enhancement method of this embodiment is executed on a CPU or a GPU.

[0045] Preferably, the image enhancement method based on the video IP frame of this embodiment is executed on the GPU, so that the operation of performing image enhancement processing on the video no longer occupies CPU resources, which can save more CPU resources for the user; moreover, the GPU is parallel computing, and the time for performing image enhancement processing on multiple videos is shorter, which can improve the efficiency of image enhancement processing on multiple videos and meet the user's high efficiency requirements for image enhancement of videos.

[0046] This embodiment uses a trained target neural network model to implement image enhancement processing. The target neural network model includes a first sub-model and a second sub-model. The second sub-model includes a convolution operation module and a summation operation module, wherein the output of the convolution operation module represents the difference between the input image and the corresponding enhanced image, and the input of the convolution operation module is the input image and the output of the first sub-model (i.e., the transformation matrix of the input image). The transformation matrix of the input image is obtained by the first sub-model, and the first sub-model includes the structure of a neural network. It can be seen that this embodiment does not use a neural network to directly learn how to obtain a corresponding enhanced image based on an input image, but uses a neural network to indirectly learn the difference between the input image and the corresponding enhanced image. Since the variation range of the difference between the input image and the enhanced image is relatively small and relatively easy to learn, the target neural network model is easier to converge during the training of the target neural network model in this embodiment, the corresponding training process is shorter, and the accuracy of the enhanced image obtained using the trained target neural network model is also higher.

[0047] Moreover, the structure of the first sub-model included in the target neural network model in this embodiment is relatively simple. It is a lightweight neural network. The time required to obtain the enhanced image corresponding to the video using the trained target neural network model is short, which improves the speed of image enhancement processing on the video.

[0048] Embodiment 2

[0049] Compared with the first embodiment, the present embodiment has the following differences: the first sub-model of the present embodiment includes, in order of information transmission direction: four cascaded convolutional layers, a two-dimensional adaptive average pooling layer, a flatten layer, a first linear layer, a relu activation layer, a second linear layer, a sigmoid activation layer and an update module; the update module is used to update the transformation matrix of the input image according to the parameter adjustment requirement information input by the user to obtain an updated transformation matrix of the input image.

[0050] The first sub-model in this embodiment also includes an updating module, and the output of the corresponding first sub-model is the updated transformation matrix of the input image, and the input of the corresponding second sub-model is the updated transformation matrix of the input image, that is, v q The transformation matrix T q is the updated transformation matrix.

[0051] In the process of training the target neural network model in this embodiment, the user has no input, the update module does not update the transformation matrix of the input image, the output of the first sub-model is the transformation matrix of the input image, and the input of the second sub-model is the transformation matrix of the input image.

[0052] Specifically, when the target neural network model is a neural network model for obtaining an enhanced image of a single-channel image, the transformation matrix of the input image is a 3×3 matrix; updating the transformation matrix of the input image according to the parameter adjustment requirement information input by the user includes:

[0053] S100, obtain the transformation matrix T0 of the input image, T0 = [e 0 1,1 ,e 0 1,2 ,e 0 1,3 ;e 0 2,1 ,e 0 2,2 ,e 0 2,3 ;e 0 3,1 ,e 0 3,2 ,e 0 3,3 ],e 0 x,y is the element in the xth row and yth column of T0, x=1,2,3, y=1,2,3.

[0054] S200, obtaining a first target adjustment value w1; if the parameter adjustment requirement information input by the user is information indicating increasing VMAF, then w1=e 0 2,2 +Δw1, Δw1 is the adjustment amplitude of the first parameter, 0<Δw1≤f max -e 0 2,2 , f max is the maximum value of the preset conversion matrix value range F'; if the parameter adjustment requirement information input by the user is information indicating increasing PSNR, then w1 = e 0 2,2 -Δw2, Δw2 is the adjustment amplitude of the second parameter, 0<Δw2≤e 0 2,2 -f min , f min is the minimum value of F'.

[0055] Those skilled in the art know that Video Multi-Method Assessment Fusion (VMAF) and Peak Signal-to-Noise Ratio (PSNR) are two common parameters for evaluating video / image quality.

[0056] In this embodiment, Δw1 and Δw2 are empirical values, and Δw1=(f max -e 0 2,2 ) / 2, Δw2=(e0 2,2 -f min ) / 2.

[0057] S300, will [e 0 1,1 ,e 0 1,2 ,e 0 1,3 ;e 0 2,1 ,w1,e 0 2,3 ;e 0 3,1 ,e 0 3,2 ,e 0 3,3 ] is determined as the updated transformation matrix for the input image.

[0058] When the target neural network model is a neural network model for obtaining an enhanced image of a 3-channel image, the conversion matrix of the input image is three 3×3 matrices; updating the conversion matrix of the input image according to the parameter adjustment requirement information input by the user includes:

[0059] S101, obtain the transformation matrix T0 of the input image, T0=(T 1 0,T 2 0,T 3 0), T k 0 is the kth transformation matrix corresponding to the input image, k = 1, 2, 3; T k 0=[e k,0 1,1 ,e k,0 1,2 ,e k,0 1,3 ;e k,0 2,1 ,e k,0 2,2 ,e k,0 2,3 ;e k,0 3,1 ,e k,0 3,2 ,e k ,0 3,3 ],e k,0 x,y T k The element at row x and column y of 0.

[0060] S102, obtaining a second target adjustment value w2, w2=(w 2,1 ,w 2,2 ,w 2,3), w 2,k for e k,0 2,2 corresponding target adjustment value; if the parameter adjustment requirement information input by the user is information indicating an increase in VMAF, then w 2,k =e k,0 2,2 +Δw3, Δw3 is the adjustment amplitude of the third parameter, 0<Δw3≤f max -max(e 1,0 2,2 ,e 2,0 2,2 ,e 3,0 2,2 ), f max is the maximum value of the preset conversion matrix value range F', and max() is the maximum value; if the parameter adjustment requirement information input by the user is information indicating increasing PSNR, then w 2,k =e k,0 2,2 -Δw4, Δw4 is the fourth parameter adjustment range, 0<Δw4≤min(e 1,0 2,2 ,e 2,0 2,2 ,e 3 ,0 2,2 )-f min , f min is the minimum value of F', and min() is the minimum value.

[0061] In this embodiment, Δw3 and Δw4 are empirical values, and Δw3=(f max -max(e 1,0 2,2 ,e 2,0 2,2 ,e 3 ,0 2,2 )) / 2, Δw4=(min(e 1,0 2,2 ,e 2,0 2,2 ,e 3,0 2,2 )-f min ) / 2.

[0062] In this embodiment, if the parameter adjustment requirement information input by the user is information indicating to increase VMAF and f max -max(e 1,0 2,2 ,e 2,0 2,2 ,e 3,0 2,2)=0, then Δw3=0, and the preset information indicating that VMAF is not to be increased is displayed on the user interface; if the parameter adjustment requirement information input by the user is the information indicating the increase of PSNR and min(e 1,0 2,2 ,e 2 ,0 2,2 ,e 3,0 2,2 )-f min =0, then Δw4=0, and preset information indicating that the PSNR increase is not performed is displayed on the user interface.

[0063] S103, ([e 1,0 1,1 ,e 1,0 1,2 ,e 1,0 1,3 ;e 1,0 2,1 ,w 2,1 ,e 1,0 2,3 ;e 1,0 3,1 ,e 1,0 3,2 ,e 1,0 3,3 ],[e 2,0 1,1 ,e 2,0 1,2 ,e 2,0 1,3 ;e 2,0 2,1 ,w 2,2 ,e 2,0 2,3 ;e 2,0 3,1 ,e 2,0 3,2 ,e 2,0 3,3 ],[e 3,0 1,1 ,e 3,0 1,2 ,e 3,0 1,3 ;e 3,0 2,1 ,w 2,3 ,e 3,0 2,3 ;e 3,0 3,1 ,e 3,0 3,2 ,e 3,0 3,3 ]) is determined as the updated transformation matrix for the input image.

[0064] Optionally, the process of obtaining F' includes:

[0065] S110, obtain a test image set C, C = {c1, c2, ..., c r ,…,c R},c r is the rth test image, the value range of r is 1 to R, R is the number of test images included in C, and each c r Same number and type of channels as the images in the target video.

[0066] In this embodiment, each c r The image is in the same application scenario as the image in the target video.

[0067] In this embodiment, R is an empirical value, and R=1000 is optional.

[0068] S210, traverse C, c r Input into the trained target neural network model to get c r The corresponding transformation matrix T' r , and T' r Append to the preset transformation matrix set T', and get T'={T'1,T'2,…,T' r ,…,T' R}, T' is initialized to the empty set.

[0069] S310, obtaining the value range F of the elements in the conversion matrix according to T', F=(f1,f2,…,f h ,…,f H ), f h is the value range of the hth element in the transformation matrix, f h =[f h,1 ,f h,2 ], f h,1 f h The minimum value of f h,2 f h The maximum value of f h,1 =min(D h ), f h,2 =max(D h ), D h To convert each T' in T' r The corresponding h-th element is appended to the preset h-th element set D' h The resulting set, D' h is initialized to an empty set, min() is the minimum value, max() is the maximum value, and the value range of h is 1 to H, where H is the number of elements in the transformation matrix.

[0070] In this embodiment, when the conversion matrix of the input image is a 3×3 matrix, the number of elements in the conversion matrix is ​​9; when the conversion matrix of the input image is three 3×3 matrices, the number of elements in the conversion matrix is ​​27.

[0071] S410, obtain F', F'=[f min ,f max ], f min =min(f 1,1 ,f 2,1 ,…,f h,1 ,…,f H,1 ), f max =max(f 1,2 ,f 2,2 ,…,f h,2 ,…,f H,2 ).

[0072] According to S110-S410 of this embodiment, the overall value range corresponding to all elements in the conversion matrix can be obtained, and the overall value range is used as the value range F' of the conversion matrix to limit the amplitude of adjustment of the intermediate elements of the conversion matrix, which can avoid the situation where the element value of the intermediate element of the conversion matrix after adjustment exceeds the overall value range corresponding to all elements in the conversion matrix before adjustment.

[0073] As a specific implementation, two buttons are set on the user interface, the first button corresponds to increasing VMAF, and the second button corresponds to increasing PSNR; if it is obtained that the user clicks the first button, it is determined that the parameter adjustment requirement information input by the user is information indicating increasing VMAF; if it is obtained that the user clicks the second button, it is determined that the parameter adjustment requirement information input by the user is information indicating increasing PSNR.

[0074] In addition to the advantages of embodiment 1, this embodiment can also obtain a video with a relatively high VMAF or a relatively high PSNR according to user needs; specifically, the update module of this embodiment is used to update the transformation matrix of the input image according to the parameter adjustment requirement information input by the user, wherein the parameter adjustment requirement information input by the user can reflect the user's need to obtain an enhanced image with a relatively high VMAF or a relatively high PSNR. After the update module updates the transformation matrix of the input image, the VMAF or PSNR of the enhanced image output by the second sub-model is improved compared to when the transformation matrix of the input image is not updated, thereby satisfying the user's need to obtain an enhanced image with a relatively high VMAF or a relatively high PSNR.

[0075] Although some specific embodiments of the present invention have been described in detail by way of example, it will be appreciated by those skilled in the art that the above examples are for illustration only and are not intended to limit the scope of the present invention. It will also be appreciated by those skilled in the art that various modifications may be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.

Claims

1. An image enhancement method based on video IP frame, characterized in that: The following steps are involved: S10, obtain the target video V, V=(v1,v2,…,v q ,…,v Q ), v q is the qth frame image included in V, the value range of q is 1 to Q, Q is the number of images included in V, v q v q+1 The previous frame image, v q+1 is the q+1th frame image included in V; v q is a key frame or a forward predictive coding frame in V; S20, parse V and obtain a list of frame numbers L corresponding to the key frames in V, where L = (id1, id2, ..., id m ,…,id M ), id m is the number of frames corresponding to the mth keypin in V, where m ranges from 1 to M, and M is the number of keyframes included in V; S30, traverse V in order, if q matches an id in L m If they are equal, then go to S40; otherwise, go to S50; S40, using the first sub-model in the trained target neural network model to obtain v q The transformation matrix T q , and T q Determine the latest transformation matrix T new ; S41, T q As the input of the second sub-model in the trained target neural network model, the output of the second sub-model in the trained target neural network model is used as v q The corresponding enhanced image; S50, Get T new ; S51, T new As the input of the second sub-model in the trained target neural network model, the output of the second sub-model in the trained target neural network model is used as v q The corresponding enhanced image.

2. The image enhancement method based on video IP frame according to claim 1, characterized in that: Use ffmpeg to parse V.

3. The image enhancement method based on video IP frame according to claim 1, characterized in that: The target neural network model includes a first sub-model and a second sub-model, the output of the first sub-model is the input of the second sub-model; the first sub-model is used to obtain the transformation matrix of the input image, and the first sub-model includes in sequence according to the direction of information transmission: four cascaded convolutional layers, a two-dimensional adaptive average pooling layer, a flatten layer, a first linear layer, a relu activation layer, a second linear layer and a sigmoid activation layer; the output of the sigmoid activation layer is the transformation matrix of the input image; the second sub-model is used to obtain an enhanced image corresponding to the input image according to the output of the first sub-model, the second sub-model includes a convolution operation module and a summation operation module, the input of the convolution operation module is the input image and the output of the first sub-model, the input of the summation operation module is the input image and the output of the convolution operation module, and the output of the summation operation module is the enhanced image corresponding to the input image.

4. The image enhancement method based on video IP frame according to claim 3 is characterized in that: The training process of the target neural network model includes the following steps: S1, obtain the image sample list A, A=(a1, a2,…, a n ,…,a N ), a n is the nth image sample, the value range of n is 1 to N, N is the number of image samples; each a n The number and type of channels are the same as those corresponding to the image in V; S2, obtain the image sample label list B = (b1, b2, ..., b n ,…,b N ), b n For a n Image after image enhancement; S3, use A and B to train the target neural network model.

5. The method for image enhancement based on video IP frame according to claim 3, characterized in that: When each v in V q When it is a single-channel image, the target neural network model is a neural network model for obtaining an enhanced image of the single-channel image, each convolution layer includes a 3×3 convolution kernel, and the transformation matrix of the input image is a 3×3 matrix.

6. The method for image enhancement based on video IP frame according to claim 3, characterized in that: When each v in V q When it is a 3-channel image, the target neural network model is a neural network model for obtaining an enhanced image of the 3-channel image, each convolution layer includes 3 3×3 convolution kernels, and the transformation matrix of the input image is 3 3×3 matrices.

7. The method for image enhancement based on video IP frame according to claim 4, characterized in that: b n To use the clahe algorithm to n Image after image enhancement.

Citation Information

Patent Citations

  • Image processing method and device, equipment and storage medium

    CN110706185A

  • Image enhancement method, system and device based on deep learning and storage medium

    CN112085681A

  • Method and device for synchronously playing instant message and live broadcast video stream

    CN112437316A

  • Model training method and device, electronic equipment, storage medium and product

    CN113936155A

  • Image processing method and device

    CN114972143A