Video processing method and model construction method
By aligning the video data and training the deep convolutional neural network model, the problem of poor high-dynamic range video reconstruction in the existing technology is solved, and high-quality HDR video reconstruction is achieved, avoiding the phenomenon of overlapping shadows.
Patent Information
- Application Number
- CN202010657268.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-09
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-07-09
AI Technical Summary
The prior art is difficult to effectively reconstruct high dynamic range videos, especially when processing complex motion information in videos.
By obtaining video data that meets the first preset condition, performing alignment processing, obtaining video sequences under different exposure times, and inputting them into the deep convolutional neural network model for training, to obtain the target model. This model can effectively avoid the overlapping phenomenon caused by inaccurate alignment of video sequences and improve the reconstruction effect of high dynamic range video.
High-quality high-dynamic range video reconstruction is achieved, avoiding the phenomenon of overlapping shadows, and improving the light and dark levels and detailed richness of the video.
Smart Images

Figure CN113920040B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video processing, and in particular, to a method for processing video and a method for constructing a model. Background Art
[0002] High Dynamic Range (hereinafter referred to as HDR) video, compared with Standard Dynamic Range (hereinafter referred to as SDR) video, has clearer light and dark levels in the image, richer image details, and can reproduce the real scene more vividly. With the development of HDR technology and the gradual popularization of HDR displays, the demand for HDR video is gradually increasing. The production of true HDR video requires the use of high-dynamic-range imaging devices at the acquisition end and also requires the use of HDR non-linear editing software during production. That is to say, the content production of HDR video has high requirements for both shooting equipment and pre-processing technology. Therefore, the current HDR content is still in a relatively scarce state.
[0003] The current HDR reconstruction algorithms are mainly for single-frame pictures and are not suitable for generating HDR videos because there is more complex motion information in videos.
[0004] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] Embodiments of the present invention provide a method for processing video and a method for constructing a model to at least solve the technical problem of poor effect in reconstructing high-dynamic-range video in the prior art.
[0006] According to one aspect of the embodiments of the present invention, a method for processing video is provided, including: receiving a service call request sent by a client, where the service call request carries video data that meets a first preset condition and a video sequence that meets a second preset condition; training the video data that meets the first preset condition and the video sequence that meets the second preset condition through machine learning; and outputting a training result, where the training result is a set of model parameters.
[0007] Further, the method further includes: packing the set of model parameters and sending it to the client.
[0008] According to one aspect of an embodiment of the present invention, there is provided a method for constructing a model, including: obtaining video data that meets a first preset condition; performing alignment processing on video sequences in the video data; processing the aligned video sequences to obtain video sequences that meet a second preset condition under different exposure durations; inputting the video sequences that meet the second preset condition and the video data that meets the first preset condition into a deep convolutional neural network model for training to obtain a target model.
[0009] Further, after obtaining the video sequences that meet the second preset condition under different exposure durations, the method further includes: sampling from the video sequences that meet the second preset condition to obtain training sample data; inputting the video sequences that meet the second preset condition and the video data that meets the first preset condition into a deep convolutional neural network model for training to obtain a target model includes: inputting the training sample data and the video data that meets the first preset condition into a deep convolutional neural network model for training to obtain the target model.
[0010] Further, inputting the training sample data and the video data that meets the first preset condition into a deep convolutional neural network model for training to obtain the target model includes: selecting a first set of three adjacent frames of pictures from the training sample data and inputting them into the deep convolutional neural network model to obtain a first HDR image of the middle frame; comparing the first HDR image of the middle frame with the corresponding first image that meets the first preset condition to obtain a first error; selecting a second set of three adjacent frames of pictures from the training sample data and inputting them into the deep convolutional neural network model to obtain a second HDR image of the middle frame, where the middle frame of the first set of three adjacent frames of pictures is adjacent to the middle frame of the second set of three adjacent frames of pictures; comparing the second HDR image of the middle frame with the corresponding second image that meets the first preset condition to obtain a second error; determining a loss function based on the first error and the second error; transmitting the estimated error calculated by the loss function back to the deep convolutional neural network model through backpropagation algorithm to train the model with gradient descent algorithm to obtain the target model.
[0011] Further, the method further includes: adding a constraint condition of temporal consistency to the first HDR image of the middle frame and the second HDR image of the middle frame.
[0012] Further, aligning the video sequences in the video dataset includes: for three adjacent frames of images in the video data, extracting feature points on each frame of the image, where the three adjacent frames of images include: a first frame of image, a second frame of image, and a third frame of image, and the second frame of image is the middle frame of the three adjacent frames of images; determining a transformation matrix of the first frame of image and a transformation matrix of the third frame of image based on the feature points on each frame of the image; transforming the first frame of image based on the transformation matrix of the first frame of image to align the first frame of image and the second frame of image, and transforming the third frame of image based on the transformation matrix of the third frame of image to align the third frame of image and the second frame of image.
[0013] Further, sampling from the video sequences that meet the second preset condition to obtain training sample data includes: sampling image regions where the optical flow between two adjacent frame image sequences in the video sequences that meet the second preset condition is greater than a preset optical flow to obtain training sample data.
[0014] Further, processing the aligned video sequences to obtain video sequences that meet the second preset condition under different exposure durations includes: using a camera response function and randomly generated exposure times to perform re-exposure processing on the aligned video sequences to obtain video sequences that meet the second preset condition under different exposure durations.
[0015] Further, after obtaining the video sequences that meet the second preset condition under different exposure durations, the method further includes: adding noise signals and applying gamma color changes to the video sequences that meet the second preset condition under different exposure durations to simulate real video sequences that meet the second preset condition.
[0016] Further, after obtaining the target model, the method further includes: acquiring video data that meets the second preset condition, where the video data includes video sequences with different exposure durations; aligning the video sequences; inputting the aligned video sequences into the target model to obtain video data that meets the first preset condition.
[0017] According to one aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, characterized in that the storage medium includes a stored program, wherein when the program runs, it controls the device where the storage medium is located to execute the method described in any one of the above.
[0018] According to one aspect of the embodiments of the present invention, there is provided a processor, wherein the processor is used to run a program, and when the program runs, it executes the method described in any one of the above.
[0019] In an embodiment of the present invention, video data that meets a first preset condition is obtained; the video sequences in the video data are aligned; the aligned video sequences are processed to obtain video sequences that meet a second preset condition under different exposure durations; the video sequences that meet the second preset condition and the video data that meets the first preset condition are input into a deep convolutional neural network model for training to obtain a target model. By inputting the aligned video sequences into the target model, the video data pairs that meet the first preset condition can effectively avoid the ghosting phenomenon in the output video caused by inaccurate alignment of the video sequences, thereby ensuring the effect of reconstructing a high dynamic range video, and further solving the technical problem of poor effect in reconstructing a high dynamic range video in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0021] Figure 1 is a hardware structure block diagram of a computer terminal according to an embodiment of the present invention;
[0022] Figure 2 is a flowchart of a method for constructing a model according to Embodiment 1 of the present invention;
[0023] Figure 3 is a schematic diagram of training a model in the method for constructing a model according to Embodiment 1 of the present invention;
[0024] Figure 4 is a flowchart of an alternative method for constructing a model according to Embodiment 1 of the present invention;
[0025] Figure 5 is a flowchart of a method for processing a video according to Embodiment 2 of the present invention; and
[0026] Figure 6 is a block diagram of an alternative computer terminal according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] In order to enable those skilled in the art of the present technology to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.
[0028] It should be noted that the terms "first", "second", etc. in the description, claims and the above drawings of the present invention are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0029] First, some nouns or terms that appear during the description of the embodiments of the present application are applicable to the following explanations:
[0030] Video sequence with alternating exposure durations or video sequences with different exposure durations: Generally, the exposure times of different frames in a captured video are the same. In a video sequence with alternating exposure durations, the exposure time of each frame can be different. For example, the first frame has a high exposure, the second frame has a low exposure, the third frame has a high exposure, and so on.
[0031] High Dynamic Range (HDR) video reconstruction: Videos captured by ordinary cameras are of low dynamic range (LDR), that is, the signal change is very small in very dark or bright video areas, and the original brightness information in the real scene cannot be restored. High Dynamic Range video reconstruction refers to reconstructing a high dynamic range video that can reflect the brightness of the real scene based on some information (such as a low dynamic range video).
[0032] In at least one embodiment of the present application, High Dynamic Range (HDR) means that the exposure dynamic range reaches a preset threshold, so that the difference between light and dark reaches a preset threshold. Compared with ordinary graphics, it can provide more dynamic range and image details, and can better translate the visual effects in the real environment. On the contrary, in a low dynamic range (LDR) video, the signal change is very small in very dark or bright video areas, and the original brightness information in the real scene cannot be restored. High Dynamic Range (HDR) video reconstruction refers to reconstructing a high dynamic range video that can reflect the brightness of the real scene based on some information (such as a low dynamic range video).
[0033] Convolutional Neural Network: A convolutional neural network is a type of feedforward neural network that contains convolutional calculations and has a deep structure, and is one of the representative algorithms in deep learning.
[0034] Attention Model: The attention model has a wide range of applications in the fields of image and natural language processing. Its core idea is that the attention to the input data is not balanced, but there are certain weight differentiations.
[0035] Example 1
[0036] According to an embodiment of the present invention, an embodiment of a method for constructing a model is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0037] The method embodiment provided by the first embodiment of this application can be executed on a mobile terminal, a computer terminal or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing the method of constructing a model is shown. As Figure 1 shown, the computer terminal 10 (or mobile device 10) may include one or more (shown as 102a, 102b,..., 102n in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than those Figure 1 shown, or have a different configuration from that Figure 1 shown.
[0038] It should be noted that the above one or more processors 102 and / or other data processing circuits are generally referred to as "data processing circuits" in this article. The data processing circuit may be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuit is a kind of processor control (such as the selection of a variable resistance terminal path connected to an interface).
[0039] The memory 104 can be used to store software programs and modules of application software, such as the program instruction / data storage device corresponding to the model construction method in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, realizes the model construction method of the above-mentioned application program. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.
[0040] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include the wireless network provided by the communication provider of the computer terminal 10. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0041] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0042] Under the above operating environment, the present application provides a model construction method as Figure 2 shown. Figure 2 It is a flowchart of the model construction method according to Embodiment 1 of the present invention.
[0043] Step S101, obtain video data that meets the first preset condition.
[0044] The video data that meets the first preset condition can be obtained from historical data, and the above video data that meets the first preset condition is high dynamic range (HDR) video data.
[0045] Step S102, perform alignment processing on the video sequences in the video data.
[0046] Step S103, process the aligned video sequences to obtain video sequences that meet the second preset condition at different exposure durations.
[0047] The above video sequences that meet the second preset condition are low dynamic range video sequences.
[0048] Step S104: Input the video sequences that meet the second preset condition and the video data that meet the first preset condition into the deep convolutional neural network model for training to obtain the target model.
[0049] Optionally, in the model construction method provided in Embodiment 1 of the present application, after obtaining the video sequences that meet the second preset condition at different exposure durations, the method further includes: Sampling from the video sequences that meet the second preset condition to obtain training sample data; Inputting the video sequences that meet the second preset condition and the video data that meet the first preset condition into the deep convolutional neural network model for training to obtain the target model includes: Inputting the training sample data and the video data that meet the first preset condition into the deep convolutional neural network model for training to obtain the target model.
[0050] That is, sampling from the low dynamic range video sequences to obtain training sample data; Inputting the training sample data and the high dynamic range video data set into the deep convolutional neural network model for training to obtain the target model.
[0051] As Figure 3 shown, the target model training process can be as follows: Construct a video image alignment processing module (corresponding to the alignment processing module in Figure 3 ), a random re-exposure module, a random degradation module, a random intelligent sampling module, and construct a convolutional neural network based on the attention mechanism as the HDR video generator. Based on the deep convolutional neural network, iteratively train the HDR generation model (corresponding to the above target model), and then reconstruct the input alternating exposure duration video sequences into HDR videos based on the trained model.
[0052] First, collect the publicly available existing HDR video data sets, covering as many scenarios as possible, that is, the video data sets that meet the first preset condition included in the above historical data.
[0053] Optionally, in the model construction method provided in the first embodiment of the present application, inputting the training sample data and the video data that meets the first preset condition into the deep convolutional neural network model for training to obtain the target model includes: selecting the first group of adjacent three-frame pictures from the training sample data and inputting them into the deep convolutional neural network model to obtain the first HDR image of the middle frame; comparing the first HDR image of the middle frame with the corresponding first image that meets the first preset condition to obtain the first error; selecting the second group of adjacent three-frame pictures from the training sample data and inputting them into the deep convolutional neural network model to obtain the second HDR image of the middle frame, where the middle frame of the first group of adjacent three-frame pictures is adjacent to the middle frame of the second group of adjacent three-frame pictures; comparing the second HDR image of the middle frame with the corresponding second image that meets the first preset condition to obtain the second error; determining the loss function based on the first error and the second error; transmitting the estimated error calculated by the loss function back to the deep convolutional neural network model through the backpropagation algorithm to train the model with the gradient descent algorithm to obtain the target model, where a temporal consistency constraint condition is added to the first HDR image of the middle frame and the second HDR image of the middle frame.
[0054] That is, select the first group of adjacent three-frame pictures from the training sample data and input them into the deep convolutional neural network model to obtain the first HDR image of the middle frame; compare the first HDR image of the middle frame with the corresponding high-dynamic-range first image to obtain the first error; select the second group of adjacent three-frame pictures from the training sample data and input them into the deep convolutional neural network model to obtain the second HDR image of the middle frame, where the middle frame of the first group of adjacent three-frame pictures is adjacent to the middle frame of the second group of adjacent three-frame pictures; compare the second HDR image of the middle frame with the corresponding high-dynamic-range second image to obtain the second error; determine the loss function based on the first error and the second error; transmit the estimated error calculated by the loss function back to the deep convolutional neural network model through the backpropagation algorithm to train the model with the gradient descent algorithm to obtain the target model.
[0055] Specifically, Figure 3 the input of the generator in is adjacent three-frame LDR pictures, and the output is the HDR image of the middle frame. The generator is a model based on a convolutional neural network, which includes a feature extraction network, an attention mechanism-based fusion module, and a reconstruction network. For each training iteration, first use the feature extraction network to extract features from the three input pictures respectively, then use the attention mechanism-based fusion module to obtain the fused features, and finally use the reconstruction network to estimate the HDR image H1 of the middle frame from the fused features and calculate the reconstruction error with the real HDR image. Then take three pictures (where the middle frame is adjacent to the middle frame of the previous input) and input them into the generator to obtain the HDR image H2 of the middle frame and calculate the reconstruction error. At the same time, a temporal consistency constraint L is added to H1 and H2. consistent; Based on the above method, the loss function of the HDR generation model is L HDR = αL rec1 + βL rec2 + γL consistent , where α, β, and γ respectively represent the weights corresponding to the losses. The error is passed back to the generation model through the backpropagation algorithm, and the model is trained with the gradient descent algorithm. After convergence, a well-performing HDR generation model is obtained.
[0056] After training the HDR video generation model according to the above steps, only need to align the video sequence with alternating exposure durations and then input it into the model for reconstruction, and a high-quality HDR video can be generated.
[0057] That is to say, this technical solution uses real HDR video data plus a random re-exposure module, a random degradation module, and an intelligent screening module to generate training data, which can effectively simulate the video sequence with alternating exposure durations shot in a real scene. The generation model based on the attention model can better reduce the ghosting phenomenon in the output video caused by inaccurate alignment. The added adjacent frame temporal consistency constraint can make the output video more stable and effectively reduce the flicker effect.
[0058] Optionally, in the model construction method provided in Embodiment 1 of this application, aligning the video sequences in the video dataset includes: for three adjacent frames of images in the video data, extracting the feature points on each frame of image, where the three adjacent frames of images include: the first frame of image, the second frame of image, and the third frame of image, and the second frame of image is the middle frame among the three adjacent frames of images; determining the transformation matrix of the first frame of image and the transformation matrix of the third frame of image based on the feature points on each frame of image; transforming the first frame of image based on the transformation matrix of the first frame of image to align the first frame of image and the second frame of image, and transforming the third frame of image based on the transformation matrix of the third frame of image to align the third frame of image and the second frame of image. That is to say, for three adjacent frames of images, first extract the feature points, then register and calculate the transformation matrix, and align the left and right frames to the middle frame.
[0059] The above entire process can be completed through the Figure 3 alignment processing module in it, that is, input the video dataset into the global contrast module, and the alignment processing module aligns the video sequences in the video dataset.
[0060] The obtained low-dynamic range image is sampled by a random intelligent sampling module to extract small image patches from the entire image as the input for each generator training. For example, the sampling criterion is to select regions with relatively large motion information. Optionally, in the model construction method provided in the first embodiment of this application, sampling is performed from a video sequence that meets the second preset condition, and the obtained training sample data includes: sampling image regions with an optical flow greater than a preset optical flow between two adjacent image sequences in the video sequence that meets the second preset condition to obtain training sample data.
[0061] That is, sampling image regions with an optical flow greater than a preset optical flow between two adjacent image sequences in a low-dynamic range video sequence to obtain training sample data. It should be noted that there are various choices for the sampling size of image patches during model training. If the training time cost and resource consumption are not considered, larger image patches or even the entire high-quality image can be used as the training input.
[0062] The aligned video data is passed through a random re-exposure module to obtain low-dynamic range images with different exposures. Optionally, in the model construction method provided in the first embodiment of this application, processing the aligned video sequence to obtain a video sequence that meets the second preset condition under different exposure durations includes: using a camera response function and randomly generated exposure times to perform re-exposure processing on the aligned video sequence to obtain a video sequence that meets the second preset condition under different exposure durations.
[0063] The random degradation module randomly adds noise and gamma color changes to the low-dynamic range image to simulate real data. Optionally, in the model construction method provided in the first embodiment of this application, after obtaining the video sequence that meets the second preset condition under different exposure durations, the method further includes: adding noise signals and applying gamma color changes to the video sequence that meets the second preset condition under different exposure durations to simulate a real video sequence that meets the second preset condition.
[0064] Optionally, in the model construction method provided in the first embodiment of this application, after obtaining the target model, the method further includes: obtaining video data that meets the second preset condition, where the video data includes video sequences with different exposure durations; performing alignment processing on the video sequences; inputting the aligned video sequences into the target model to obtain video data that meets the first preset condition.
[0065] As Figure 4 shown, after training the HDR video generation model, the real alternating exposure duration video sequence is aligned through the alignment processing module and then input into the HDR video generation model (corresponding to the above-mentioned target model) to output high-quality HDR video data.
[0066] In summary, in the method for constructing a model provided in the first embodiment of the present application, video data satisfying a first preset condition is obtained; the video sequences in the video data are aligned; the aligned video sequences are processed to obtain video sequences satisfying a second preset condition under different exposure durations; the video sequences satisfying the second preset condition and the video data satisfying the first preset condition are input into a deep convolutional neural network model for training to obtain a target model. By inputting the aligned video sequences into the target model, it is possible to effectively avoid the ghosting phenomenon in the output video caused by inaccurate alignment of the video sequences, thereby ensuring the effect of reconstructing a high-dynamic-range video, and further solving the technical problem of poor effect in reconstructing a high-dynamic-range video in the prior art.
[0067] It should be noted that the method for constructing a model provided in the first embodiment of the present application is applicable to refurbishing old video data, old movies, etc. to obtain high-dynamic-range videos. The method for constructing a model provided in the first embodiment of the present application can also be applied to the monitoring field. For example, if the clarity of the monitored video is low and it is difficult to accurately obtain the information in the monitoring video, the monitoring video can be input into the above-mentioned target model for processing to obtain qualified video data for accurately obtaining the information in the monitoring video.
[0068] In addition, those skilled in the art can also adjust the method for constructing a model provided in the first embodiment based on actual applications. For example, if it is necessary to convert a high-dynamic-range video into a low-dynamic-range video, the parameters of the above-mentioned target model can be adjusted to blur the input high-dynamic-range video and output a low-dynamic-range video. That is, the above-mentioned target model is used in reverse. Whether used forward or in reverse, within the core technical concept of the present application, it belongs to the application scenarios protected by the present application.
[0069] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0070] Through the description of the above embodiments, those skilled in the art can clearly understand that the video processing method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0071] Embodiment 2
[0072] This application provides a video processing method as Figure 5 shown. Figure 5 It is a flowchart of the video processing method according to Embodiment 2 of the present invention.
[0073] Step S501, receive a service call request sent by a client. Among them, the service call request carries video data that meets the first preset condition and a video sequence that meets the second preset condition.
[0074] Step S502, perform machine learning training on the video data that meets the first preset condition and the video sequence that meets the second preset condition.
[0075] The above-mentioned video sequence of the second preset condition is a low dynamic range video sequence.
[0076] Step S503, output a training result, where the training result is a set of model parameters.
[0077] Optionally, the method further includes: packing the set of model parameters and sending it to the client. The parameters in the set of model parameters are used to determine the target model.
[0078] After obtaining the target model, the method further includes: obtaining video data that meets the second preset condition, where the video data includes video sequences with different exposure durations; performing alignment processing on the video sequences; inputting the aligned video sequences into the target model to obtain video data that meets the first preset condition.
[0079] That is, after training the HDR video generation model, the real alternating exposure duration video sequences are aligned through an alignment processing module and then input into the HDR video generation model (corresponding to the above-mentioned target model) to output high-quality HDR video data.
[0080] It should be noted that the set of model parameters can also be displayed on the client.
[0081] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0082] Through the description of the above embodiments, those skilled in the art can clearly understand that the video processing method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present invention.
[0083] Embodiment 3
[0084] An embodiment of the present invention can provide a computer terminal, and this computer terminal can be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the above computer terminal can also be replaced with a terminal device such as a mobile terminal.
[0085] Optionally, in this embodiment, the above computer terminal can be located in at least one of multiple network devices in a computer network.
[0086] In this embodiment, the above computer terminal can execute the program code of the following steps in the method for constructing a model of an application program: obtaining video data that meets a first preset condition; performing alignment processing on the video sequences in the video data; processing the aligned video sequences to obtain video sequences that meet a second preset condition under different exposure durations; inputting the video sequences that meet the second preset condition and the video data that meets the first preset condition into a deep convolutional neural network model for training to obtain a target model.
[0087] The above computer terminal can also execute the program code of the following steps in the method for constructing a model of an application program: After obtaining video sequences that meet the second preset condition under different exposure durations, the method further includes: sampling from the video sequences that meet the second preset condition to obtain training sample data; inputting the video sequences that meet the second preset condition and the video data that meet the first preset condition into a deep convolutional neural network model for training to obtain a target model, including: inputting the training sample data and the video data that meet the first preset condition into a deep convolutional neural network model for training to obtain the target model.
[0088] The above computer terminal can also execute the program code of the following steps in the method for constructing a model of an application program: Inputting the training sample data and the video data that meet the first preset condition into a deep convolutional neural network model for training to obtain the target model, including: selecting a first set of three adjacent frames of pictures from the training sample data and inputting them into the deep convolutional neural network model to obtain a first HDR image of the intermediate frame; comparing the first HDR image of the intermediate frame with the corresponding first image that meets the first preset condition to obtain a first error; selecting a second set of three adjacent frames of pictures from the training sample data and inputting them into the deep convolutional neural network model to obtain a second HDR image of the intermediate frame, where the intermediate frame of the first set of three adjacent frames of pictures is adjacent to the intermediate frame of the second set of three adjacent frames of pictures; comparing the second HDR image of the intermediate frame with the corresponding second image that meets the first preset condition to obtain a second error; determining a loss function based on the first error and the second error; transmitting the estimated error calculated by the loss function back to the deep convolutional neural network model through the backpropagation algorithm to train the model using the gradient descent algorithm to obtain the target model.
[0089] The above computer terminal can also execute the program code of the following steps in the method for constructing a model of an application program: The method further includes: adding a constraint condition of temporal consistency to the first HDR image of the intermediate frame and the second HDR image of the intermediate frame.
[0090] The above computer terminal can also execute the program code for the following steps in the method for constructing a model of an application program: Aligning the video sequences in the video dataset includes: For three adjacent frames of images in the video data, extracting the feature points on each frame of the image, where the three adjacent frames of images include: a first frame of image, a second frame of image, and a third frame of image, and the second frame of image is the middle frame among the three adjacent frames of images; Determining the transformation matrix of the first frame of image and the transformation matrix of the third frame of image based on the feature points on each frame of the image; Transforming the first frame of image based on the transformation matrix of the first frame of image to align the first frame of image and the second frame of image, and transforming the third frame of image based on the transformation matrix of the third frame of image to align the third frame of image and the second frame of image.
[0091] The above computer terminal can also execute the program code for the following steps in the method for constructing a model of an application program: Sampling from the video sequences that meet the second preset condition to obtain training sample data includes: Sampling the image regions where the optical flow between two adjacent frame sequences in the video sequences that meet the second preset condition is greater than the preset optical flow to obtain training sample data.
[0092] The above computer terminal can also execute the program code for the following steps in the method for constructing a model of an application program: Processing the aligned video sequences to obtain video sequences that meet the second preset condition under different exposure durations includes: Using the camera response function and randomly generated exposure times to perform re-exposure processing on the aligned video sequences to obtain video sequences that meet the second preset condition under different exposure durations.
[0093] The above computer terminal can also execute the program code for the following steps in the method for constructing a model of an application program: After obtaining the video sequences that meet the second preset condition under different exposure durations, the method further includes: Adding noise signals and applying gamma color changes to the video sequences that meet the second preset condition under different exposure durations to simulate real video sequences that meet the second preset condition.
[0094] The above computer terminal can also execute the program code for the following steps in the method for constructing a model of an application program: After obtaining the target model, the method further includes: Obtaining video data that meets the second preset condition, where the video data includes video sequences with different exposure durations; Aligning the video sequences; Inputting the aligned video sequences into the target model to obtain video data that meets the first preset condition.
[0095] Optionally, Figure 6 is a structural block diagram of a computer terminal according to an embodiment of the present invention. AsFigure 6 As shown, the computer terminal 10 may include: one or more ( Figure 6 only one is shown in the figure) processors and a memory.
[0096] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the model construction method in the embodiments of the present invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the above-mentioned model construction method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely provided relative to the processor, and these remote memories can be connected to the terminal 10 through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and their combinations.
[0097] The processor can call the information and application programs stored in the memory through a transmission device to execute the following steps: obtaining video data that meets the first preset condition; performing alignment processing on the video sequences in the video data; processing the aligned video sequences to obtain video sequences that meet the second preset condition under different exposure durations; inputting the video sequences that meet the second preset condition and the video data that meets the first preset condition into a deep convolutional neural network model for training to obtain a target model.
[0098] Optionally, the processor can also call the information and application programs stored in the memory through a transmission device to execute the following steps: after obtaining the video sequences that meet the second preset condition under different exposure durations, the method further includes: sampling from the video sequences that meet the second preset condition to obtain training sample data; inputting the video sequences that meet the second preset condition and the video data that meets the first preset condition into a deep convolutional neural network model for training to obtain a target model includes: inputting the training sample data and the video data that meets the first preset condition into a deep convolutional neural network model for training to obtain the target model.
[0099] Optionally, the processor can also call the information and application programs stored in the memory through the transmission device and execute the following steps: input the training sample data and the video data that meet the first preset condition into the deep convolutional neural network model for training to obtain the target model, including: select the first set of three adjacent frames of pictures from the training sample data and input them into the deep convolutional neural network model to obtain the first HDR image of the intermediate frame; compare the first HDR image of the intermediate frame with the corresponding first image that meets the first preset condition to obtain the first error; select the second set of three adjacent frames of pictures from the training sample data and input them into the deep convolutional neural network model to obtain the second HDR image of the intermediate frame, where the intermediate frame of the first set of three adjacent frames of pictures is adjacent to the intermediate frame of the second set of three adjacent frames of pictures; compare the second HDR image of the intermediate frame with the corresponding second image that meets the first preset condition to obtain the second error; determine the loss function based on the first error and the second error; transmit the estimated error calculated by the loss function back to the deep convolutional neural network model through the backpropagation algorithm and train the model with the gradient descent algorithm to obtain the target model.
[0100] Optionally, the processor can also call the information and application programs stored in the memory through the transmission device and execute the following steps: the method further includes: adding a temporal consistency constraint condition to the first HDR image of the intermediate frame and the second HDR image of the intermediate frame.
[0101] Optionally, the processor can also call the information and application programs stored in the memory through the transmission device and execute the following steps: aligning the video sequences in the video dataset includes: for three adjacent frames of images in the video data, extract the feature points on each frame of the image, where the three adjacent frames of images include: the first frame of the image, the second frame of the image, and the third frame of the image, and the second frame of the image is the intermediate frame of the three adjacent frames of images; determine the transformation matrix of the first frame of the image and the transformation matrix of the third frame of the image based on the feature points on each frame of the image; transform the first frame of the image based on the transformation matrix of the first frame of the image to align the first frame of the image and the second frame of the image, and transform the third frame of the image based on the transformation matrix of the third frame of the image to align the third frame of the image and the second frame of the image.
[0102] Optionally, the processor can also call the information and application programs stored in the memory through the transmission device and execute the following steps: sampling from the video sequences that meet the second preset condition to obtain training sample data, including: sampling the image regions where the optical flow between adjacent two-frame image sequences in the video sequences that meet the second preset condition is greater than the preset optical flow to obtain the training sample data.
[0103] Optionally, the processor can also call the information and application programs stored in the memory through the transmission device and execute the following steps: processing the aligned video sequence to obtain video sequences that meet the second preset condition under different exposure durations, including: using the camera response function and randomly generated exposure times to perform re-exposure processing on the aligned video sequence to obtain video sequences that meet the second preset condition under different exposure durations.
[0104] Optionally, the processor can also call the information and application programs stored in the memory through the transmission device and execute the following steps: after obtaining the video sequences that meet the second preset condition under different exposure durations, the method further includes: adding a noise signal and applying gamma color change to the video sequences that meet the second preset condition under different exposure durations to simulate a real video sequence that meets the second preset condition.
[0105] Optionally, the processor can also call the information and application programs stored in the memory through the transmission device and execute the following steps: after obtaining the target model, the method further includes: obtaining video data that meets the second preset condition, where the video data includes video sequences with different exposure durations; performing alignment processing on the video sequences; inputting the aligned video sequences into the target model to obtain video data that meets the first preset condition.
[0106] By adopting the embodiment of the present invention, a method for constructing a model is provided. By obtaining video data that meets the first preset condition; performing alignment processing on the video sequences in the video data; processing the aligned video sequences to obtain video sequences that meet the second preset condition under different exposure durations; inputting the video sequences that meet the second preset condition and the video data that meets the first preset condition into a deep convolutional neural network model for training to obtain a target model, and inputting the aligned video sequences into the target model to obtain video data that meets the first preset condition, it can effectively avoid the ghosting phenomenon in the output video caused by inaccurate alignment of the video sequences, thus ensuring the effect of reconstructing a high dynamic range video, and further solving the technical problem of poor effect of reconstructing a high dynamic range video in the prior art.
[0107] Those of ordinary skill in the art can understand that Figure 6 the structure shown is only for illustration, and the computer terminal can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, and terminal devices such as Mobile Internet Devices (MID), PAD, etc. Figure 6 It does not limit the structure of the above electronic device. For example, the computer terminal 10 may further include more Figure 6more or fewer components (such as network interfaces, display devices, etc.) shown in the figure, or having a configuration different from that Figure 6 shown in the figure.
[0108] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and this program can be stored in a computer-readable storage medium. The storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, etc.
[0109] Embodiment 4
[0110] The embodiments of the present invention also provide a computer-readable storage medium. Optionally, in this embodiment, the above storage medium can be used to store the program code executed by the method for constructing the model provided in the first embodiment above.
[0111] Optionally, in this embodiment, the above storage medium can be located in any one of the computer terminals in a computer terminal group in a computer network, or in any one of the mobile terminals in a mobile terminal group.
[0112] Optionally, in this embodiment, the storage medium is set to store program code for performing the following steps: obtaining video data that meets a first preset condition; performing alignment processing on the video sequences in the video data; processing the aligned video sequences to obtain video sequences that meet a second preset condition under different exposure durations; inputting the video sequences that meet the second preset condition and the video data that meets the first preset condition into a deep convolutional neural network model for training to obtain a target model.
[0113] The storage medium is also set to store program code for performing the following steps: after obtaining the video sequences that meet the second preset condition under different exposure durations, the method further includes: sampling from the video sequences that meet the second preset condition to obtain training sample data; inputting the video sequences that meet the second preset condition and the video data that meets the first preset condition into a deep convolutional neural network model for training to obtain a target model includes: inputting the training sample data and the video data that meets the first preset condition into a deep convolutional neural network model for training to obtain the target model.
[0114] The storage medium is further configured to store program code for performing the following steps: inputting the training sample data and the video data satisfying the first preset condition into a deep convolutional neural network model for training to obtain the target model, including: selecting a first group of three adjacent frames of pictures from the training sample data and inputting them into the deep convolutional neural network model to obtain a first HDR image of the middle frame; comparing the first HDR image of the middle frame with a corresponding first image satisfying the first preset condition to obtain a first error; selecting a second group of three adjacent frames of pictures from the training sample data and inputting them into the deep convolutional neural network model to obtain a second HDR image of the middle frame, wherein the middle frame of the first group of three adjacent frames of pictures is adjacent to the middle frame of the second group of three adjacent frames of pictures; comparing the second HDR image of the middle frame with a corresponding second image satisfying the first preset condition to obtain a second error; determining a loss function based on the first error and the second error; transmitting the estimated error calculated by the loss function back to the deep convolutional neural network model through backpropagation algorithm to train the model using gradient descent algorithm to obtain the target model.
[0115] The storage medium is further configured to store program code for performing the following steps: The method further includes: adding a temporal consistency constraint condition to the first HDR image of the middle frame and the second HDR image of the middle frame.
[0116] The storage medium is further configured to store program code for performing the following steps: Aligning the video sequences in the video dataset includes: for three adjacent frames of images in the video data, extracting feature points on each frame of the image, wherein the three adjacent frames of images include: a first frame of image, a second frame of image, and a third frame of image, and the second frame of image is the middle frame of the three adjacent frames of images; determining a transformation matrix of the first frame of image and a transformation matrix of the third frame of image based on the feature points on each frame of the image; transforming the first frame of image based on the transformation matrix of the first frame of image to align the first frame of image and the second frame of image, and transforming the third frame of image based on the transformation matrix of the third frame of image to align the third frame of image and the second frame of image.
[0117] The storage medium is further configured to store program code for performing the following steps: Sampling from the video sequences satisfying the second preset condition to obtain training sample data includes: sampling the image regions with optical flow greater than a preset optical flow between two adjacent frame sequences in the video sequences satisfying the second preset condition to obtain training sample data.
[0118] The storage medium is also configured to store program code for performing the following steps: processing the aligned video sequence to obtain video sequences satisfying a second preset condition at different exposure durations, including: performing re-exposure processing on the aligned video sequence by using a camera response function and randomly generated exposure times to obtain video sequences satisfying the second preset condition at different exposure durations.
[0119] The storage medium is also configured to store program code for performing the following steps: after obtaining video sequences satisfying the second preset condition at different exposure durations, the method further includes: adding a noise signal to and applying a gamma color transformation to the video sequences satisfying the second preset condition at different exposure durations to simulate real video sequences satisfying the second preset condition.
[0120] The storage medium is also configured to store program code for performing the following steps: after obtaining the target model, the method further includes: acquiring video data satisfying the second preset condition, where the video data includes video sequences at different exposure durations; performing alignment processing on the video sequences; inputting the aligned video sequences into the target model to obtain video data satisfying the first preset condition.
[0121] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0122] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0123] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the units or modules can be in an electrical or other form.
[0124] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0125] In addition, in each embodiment of the present invention, the functional units may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0126] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0127] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A method for processing video, characterized in that, it includes: Receiving a service call request sent by a client, wherein the service call request carries video data that meets a first preset condition and a video sequence that meets a second preset condition, wherein the video data that meets the first preset condition is high dynamic range (HDR) video data, and the video sequence that meets the second preset condition is a low dynamic range video sequence; Training the video data that meets the first preset condition and the video sequence that meets the second preset condition through machine learning training; wherein, sampling from the video sequence that meets the second preset condition to obtain training sample data; obtaining a first HDR image of an intermediate frame and a second HDR image of an intermediate frame from the training sample data; comparing the first HDR image of the intermediate frame with a corresponding first image that meets the first preset condition to obtain an error one; comparing the second HDR image of the intermediate frame with a corresponding second image that meets the first preset condition to obtain an error two; determining a loss function based on the error one and the error two; transmitting the estimated error calculated by the loss function back to a deep convolutional neural network model through backpropagation algorithm to train the machine learning with gradient descent algorithm; Outputting a training result, wherein the training result is a set of model parameters.
2. The method according to claim 1, characterized in that, the method further includes: packing the set of model parameters and sending it to the client.
3. A method for constructing a model, characterized in that, it includes: Obtaining video data that meets a first preset condition, wherein the video data that meets the first preset condition is high dynamic range (HDR) video data; Aligning the video sequences in the video data; Processing the aligned video sequences to obtain video sequences that meet a second preset condition under different exposure durations, wherein the video sequences that meet the second preset condition are low dynamic range video sequences; Inputting the video sequences that meet the second preset condition and the video data that meets the first preset condition into a deep convolutional neural network model for training to obtain a target model; Sampling from the video sequences that meet the second preset condition to obtain training sample data; obtaining a first HDR image of an intermediate frame and a second HDR image of an intermediate frame from the training sample data; comparing the first HDR image of the intermediate frame with a corresponding first image that meets the first preset condition to obtain an error one; comparing the second HDR image of the intermediate frame with a corresponding second image that meets the first preset condition to obtain an error two; determining a loss function based on the error one and the error two; transmitting the estimated error calculated by the loss function back to a deep convolutional neural network model through backpropagation algorithm to train the model with gradient descent algorithm to obtain the target model.
4. The method according to claim 3, characterized in that, After obtaining the video sequences that meet the second preset condition under different exposure durations, the method further includes: inputting the video sequences that meet the second preset condition and the video data that meet the first preset condition into a deep convolutional neural network model for training to obtain a target model, including: inputting the training sample data and the video data that meet the first preset condition into the deep convolutional neural network model for training to obtain the target model.
5. The method according to claim 4, wherein, the method further includes: selecting a first set of three adjacent frames of pictures from the training sample data and inputting them into the deep convolutional neural network model to obtain a first HDR image of the middle frame; selecting a second set of three adjacent frames of pictures from the training sample data and inputting them into the deep convolutional neural network model to obtain a second HDR image of the middle frame, wherein the middle frame of the first set of three adjacent frames of pictures is adjacent to the middle frame of the second set of three adjacent frames of pictures.
6. The method according to any one of claims 3 or 5, wherein, the method further includes: adding a constraint condition of temporal consistency to the first HDR image of the middle frame and the second HDR image of the middle frame.
7. The method according to claim 5, wherein, performing alignment processing on the video sequences in the video dataset includes: extracting feature points on each frame of the adjacent three frames of images in the video data, wherein the adjacent three frames of images include: a first frame of image, a second frame of image, and a third frame of image, and the second frame of image is the middle frame of the adjacent three frames of images; determining a transformation matrix of the first frame of image and a transformation matrix of the third frame of image based on the feature points on each frame of image; transforming the first frame of image based on the transformation matrix of the first frame of image to align the first frame of image and the second frame of image, and transforming the third frame of image based on the transformation matrix of the third frame of image to align the third frame of image and the second frame of image.
8. The method according to claim 3, wherein, sampling from the video sequences that meet the second preset condition to obtain training sample data includes: sampling the image regions with optical flow greater than a preset optical flow between adjacent two-frame image sequences in the video sequences that meet the second preset condition to obtain training sample data.
9. The method according to claim 3, wherein, processing the aligned video sequences to obtain video sequences that meet the second preset condition under different exposure durations includes: performing re-exposure processing on the aligned video sequences by using a camera response function and randomly generated exposure times to obtain video sequences that meet the second preset condition under different exposure durations.
10. The method according to claim 8, wherein, after obtaining the video sequences that meet the second preset condition under different exposure durations, the method further includes: Add a noise signal and apply gamma color transformation to the video sequences that meet the second preset condition under different exposure durations to simulate a real video sequence that meets the second preset condition.
11. The method according to claim 3, wherein, after obtaining the target model, the method further includes: obtaining video data that meets the second preset condition, wherein the video data includes video sequences with different exposure durations; performing alignment processing on the video sequences; inputting the aligned video sequences into the target model to obtain video data that meets the first preset condition.
12. A computer-readable storage medium, wherein, the storage medium includes a stored program, and when the program runs, it controls the device where the storage medium is located to execute the method according to any one of claims 1 to 11.
13. A processor, wherein, the processor is used to run a program, and when the program runs, it executes the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Image high dynamic range reconstruction method based on deep learning
CN111292264A
Generation of high dynamic range visual media
US20190096046A1