Training Method, Device, Equipment and Medium of Super-Resolution Model
By splitting the training video into multiple training samples and arranging them according to specific standards, the problem of slow training speed of super-resolution models is solved, and a more efficient training process and better model accuracy is achieved.
Patent Information
- Application Number
- CN202111596501.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-24
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-12-24
AI Technical Summary
The prior art is slow when training super-resolution models and cannot effectively utilize the information in the training video.
The training video is split into multiple training samples, each of which contains images of the same image size and number of video frames, arranged by the number of video frames and image size, and the super-resolution model is trained in stages using these training samples in sequence.
Through phased training, the training speed of the super-resolution model is improved, the model accuracy is maintained, and the information in the training video is effectively utilized.
Smart Images

Figure CN114332561B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of machine learning, and particularly to a method, device, equipment and medium for training a super-resolution model. Background Art
[0002] Super-resolution is used to increase the resolution of the original image by hardware or software methods. A super-resolution model is a model that obtains a high-resolution image from a low-resolution image.
[0003] In related technologies, when training a super-resolution model, all video frames of a training video are extracted, and all video frames are input into the super-resolution model frame by frame to train the super-resolution model.
[0004] However, the training video contains a lot of information, and the training speed of the super-resolution model is slow. Summary of the Invention
[0005] This application provides a method, device, equipment and medium for training a super-resolution model. This method can improve the training speed of the super-resolution model. The technical solutions are as follows:
[0006] According to one aspect of this application, a method for training a super-resolution model is provided. The method includes:
[0007] Split a training video into p types of training samples. Each type of training sample includes at least f training samples with the same image size and the same number of video frames. The number of video frames of each type of training sample in the p types of training samples is not greater than the number of video frames of the training video, and the image size of each type of training sample in the p types of training samples is not greater than the image size of the training video. p is a positive integer greater than 1, and f is a positive integer;
[0008] Arrange the p types of training samples in ascending order according to at least one arrangement criterion among the number of video frames and the image size;
[0009] Extract training samples from the p types of training samples in the arranged order and train the super-resolution model in turn.
[0010] According to one aspect of this application, a device for training a super-resolution model is provided. The device includes:
[0011] A splitting module, configured to split a training video into p types of training samples, each type of training sample including at least f training samples with the same image size and the same number of video frames. The number of video frames of each type of training sample among the p types of training samples is not greater than the number of video frames of the training video, and the image size of each type of training sample among the p types of training samples is not greater than the image size of the training video. p is a positive integer greater than 1, and f is a positive integer;
[0012] The splitting module is further configured to arrange the p types of training samples in ascending order according to at least one of the number of video frames and the image size;
[0013] A training module, configured to sequentially extract training samples from the p types of training samples according to the arranged order of the p types of training samples to train the super-resolution model.
[0014] According to another aspect of the present application, there is provided a computer device, which includes: a processor and a memory. The memory stores at least one instruction, at least one program, a code set or an instruction set. The at least one instruction, at least one program, the code set or the instruction set is loaded and executed by the processor to implement the training method of the super-resolution model as described in the above aspect.
[0015] According to another aspect of the present application, there is provided a computer storage medium. At least one program code is stored in the computer-readable storage medium. The program code is loaded and executed by the processor to implement the training method of the super-resolution model as described in the above aspect.
[0016] According to another aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions. The computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the training method of the super-resolution model as described in the above aspect.
[0017] The beneficial effects brought by the technical solution provided by the embodiments of the present application at least include:
[0018] Split the training video to obtain training samples. Arrange the training samples in ascending order according to the number of video frames and the image size. Then, according to the arrangement order, use different training samples to train the super-resolution model in stages. Since the smaller the number of video frames or the image size, the less information the training sample contains, which helps to improve the training speed. Moreover, the training of the super-resolution model in the previous stage has a guiding effect, which can guide the training of the super-resolution model in the current stage, enabling the super-resolution model to learn from simple to difficult, effectively improving the training speed while maintaining the model accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] BRIEF DESCRIPTION OF THE DRAWINGS To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0020] Figure 1 is a schematic structural diagram of a computer system provided by an exemplary embodiment of the present application;
[0021] Figure 2 is a schematic diagram of a method for training a super-resolution model provided by an exemplary embodiment of the present application;
[0022] Figure 3 is a schematic flowchart of a method for training a super-resolution model provided by an exemplary embodiment of the present application;
[0023] Figure 4 is a schematic flowchart of a method for training a super-resolution model provided by an exemplary embodiment of the present application;
[0024] Figure 5 is a schematic diagram of a training sample provided by an exemplary embodiment of the present application;
[0025] Figure 6 is a schematic flowchart of a method for training a super-resolution model provided by an exemplary embodiment of the present application;
[0026] Figure 7 is a schematic flowchart of a method for training a super-resolution model provided by an exemplary embodiment of the present application;
[0027] Figure 8 is a schematic diagram of a training stage of a super-resolution model provided by an exemplary embodiment of the present application;
[0028] Figure 9 is a schematic model diagram of a training device for a super-resolution model provided by an exemplary embodiment of the present application;
[0029] Figure 10 It is a schematic structural diagram of a computer device provided by an exemplary embodiment of the present application. Detailed implementation manners
[0030] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0031] First, introduce the nouns involved in the embodiments of the present application:
[0032] Artificial Intelligence (AI): It is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0033] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0034] Machine Learning (ML): An interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and teaching learning.
[0035] Super-resolution model: Used to improve the resolution of images or videos. Optionally, the super-resolution model includes a super-resolution model based on a recurrent network and a super-resolution model based on a sliding window. The types of super-resolution models are not limited in the embodiments of the present application.
[0036] Figure 1The structural schematic diagram of a computer system provided by an exemplary embodiment of the present application is shown. The computer system 100 includes: a terminal 120 and a server 140.
[0037] An application related to the super-resolution model is installed on the terminal 120. The application can be a mini-program in an app (application), a dedicated application, or a web client. The terminal 120 is at least one of a smart phone, a tablet computer, an e-book reader, an MP3 player, an MP4 player, a laptop portable computer, and a desktop computer. Optionally, the super-resolution model is deployed on the terminal 120.
[0038] The terminal 120 is connected to the server 140 through a wireless network or a wired network.
[0039] The server 140 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Optionally, the super-resolution model is deployed on the server 140. Optionally, the server 140 undertakes the main computing work, and the terminal 120 undertakes the secondary computing work; or, the server 140 undertakes the secondary computing work, and the terminal 120 undertakes the main computing work; or, both the server 140 and the terminal 120 adopt a distributed computing architecture for collaborative computing.
[0040] Optionally, in the embodiment of the present application, there may be only the terminal 120, or only the server 140.
[0041] The present application will split the training video into training samples including different numbers of video frames and different image sizes, and use the foregoing training samples to train the super-resolution model in stages.
[0042] Exemplarily, for ease of understanding, please refer to Figure 2, for ease of explanation, the number of each type of training sample is set to 1 here (the number of each type of training sample can be modified by those skilled in the art according to the actual situation). Here, it is assumed that the training video 201 includes 4 video frames, and the image size is 2*2. The training video 201 is split into training samples 202, 203, 204, and 205. Among them, the training sample 202 includes 2 video frames, and the image size is 1*1; the training sample 203 includes 2 video frames, and the image size is 2*2; the training sample 204 includes 4 video frames, and the image size is 1*1; the training sample 205 includes 4 video frames, and the image size is 2*2.
[0043] The training sample 202 is input into the super-resolution model 206 to complete the training of the super-resolution model 206 in the OA training stage. Then, when the super-resolution model 206 completes the training in the OA training stage, the training sample 203 is input into the super-resolution model 206 to complete the training of the super-resolution model 206 in the AB training stage. Then, when the super-resolution model 206 completes the training in the AB training stage, the training sample 204 is input into the super-resolution model 206 to complete the training of the super-resolution model 206 in the BC training stage. Then, when the super-resolution model 206 completes the training in the BC training stage, the training sample 205 is input into the super-resolution model 206 to complete the training of the super-resolution model 206 in the CD training stage. Then, when the super-resolution model 206 completes the training in the CD training stage, it is considered that the training of the training video 201 on the super-resolution model 206 is completed.
[0044] In summary, the method splits the training video to obtain training samples, arranges the training samples in ascending order according to the number of video frames and the image size, and uses different training samples to train the super-resolution model in stages according to the arrangement order. Since the smaller the number of video frames or the image size, the less information the training sample contains, which helps to improve the training speed. Moreover, the training of the super-resolution model in the previous stage has a guiding effect, which can guide the training of the super-resolution model in the current stage, enabling the super-resolution model to learn from simple to difficult, and effectively improving the training speed while maintaining the model accuracy.
[0045] Figure 3 shows a training method for a super-resolution model provided by an embodiment of the present application. This method can be executed by Figure 1 the terminal 120 or the server 140 shown, and this method includes:
[0046] Step 302: Split the training video into p types of training samples. Each type of training sample includes at least f training samples with the same image size and the same number of video frames. The number of video frames of the p types of training samples is not greater than the number of video frames of the training video, and the image sizes of the p types of training samples are not greater than the image size of the training video.
[0047] Among them, p is a positive integer greater than 1.
[0048] The training video includes one or more segments of video. The training video can be a video stored locally, a video downloaded from the network, or a video provided by other computer devices.
[0049] In an alternative embodiment of the present application, the video frames of the training video are processed first, and then the image size is processed. Exemplarily, according to m types of frame extraction strategies, m types of video frame sequences are extracted from the training video; according to n types of cropping strategies, at least one of the m types of video frame sequences is cropped into samples of n image sizes to obtain p types of training samples, where n and m are positive integers. Among them, the frame extraction strategy is used to extract video frame sequences with different numbers of video frames from the training video, and the cropping strategy is used to crop the images of at least one video frame sequence. Exemplarily, as Figure 2 shown, according to 2 types of frame extraction strategies, frame extraction processing is performed on the training video 201 to obtain two types of video frame sequences. One type of video frame sequence includes 2 video frames, and the other type of video frame includes 4 video frames. According to 2 types of cropping strategies, the video frame sequence including 2 video frames is cropped to obtain the training sample 202 and the training sample 203. According to 2 types of cropping strategies, the video frame sequence including 4 video frames is cropped to obtain the training sample 204 and the training sample 205.
[0050] In another alternative embodiment of the present application, the image size of the training video is processed first, and then the video frames are processed. Exemplarily, according to n types of cropping strategies, the training video is cropped into cropped videos of n image sizes; according to m types of frame extraction strategies, frame extraction processing is performed on at least one segment of video in the video set to obtain p types of training samples, and each type of training sample has different numbers of video frames and image sizes.
[0051] Optionally, p = m * n, that is, the above n types of video frame sequences are all subjected to cropping processing.
[0052] Step 304: Arrange the p types of training samples in ascending order according to at least one arrangement criterion among the number of video frames and the image size.
[0053] Exemplarily, if Figure 2As shown, the known training sample 202 includes 2 video frames with an image size of 1*1; the training sample 203 includes 2 video frames with an image size of 2*2; the training sample 204 includes 4 video frames with an image size of 1*1; the training sample 205 includes 4 video frames with an image size of 2*2. Then, taking the number of video frames and the image size as the arrangement criteria, these 4 training samples are arranged, and the obtained arrangement order is "training sample 202 - training sample 203 - training sample 204 - training sample 205".
[0054] Optionally, if there are a first training sample and a second training sample among the p types of training samples with the same number of video frames and the same image size, then randomly arrange the order between the first training sample and the second training sample.
[0055] Step 306: According to the arrangement order of the p types of training samples, sequentially extract training samples from the p types of training samples to train the super-resolution model.
[0056] Optionally, after the i-th type of training sample among the p types of training samples completes the training of the super-resolution model, take the (i + 1)-th type of training sample to train the super-resolution model.
[0057] Exemplarily, if Figure 2 As shown, according to the arrangement order of "training sample 202 - training sample 203 - training sample 204 - training sample 205", sequentially extract training samples from the 4 training samples to train the super-resolution model.
[0058] The embodiments of the present application do not specifically limit the type and training method of the super-resolution model. Exemplarily, the super-resolution model is any one of a super-resolution model based on a recurrent network and a super-resolution model based on a sliding window. Exemplarily, the training method is the error backpropagation algorithm.
[0059] In other optional ways of the present application, this method can also be applied to other models that require video content as training samples, for example, a video classification model based on video content.
[0060] In summary, this embodiment splits the training video to obtain training samples, arranges the training samples from small to large according to the number of video frames and the image size, and according to the arrangement order, uses different training samples to train the super-resolution model in stages. Since the smaller the number of video frames or the image size, the less information the training sample contains, which helps to improve the training speed. Moreover, the training of the super-resolution model in the previous stage has a guiding effect, which can guide the training of the super-resolution model in the current stage, enabling the super-resolution model to learn from simple to difficult, while maintaining the model accuracy, effectively improving the training speed.
[0061] In the following embodiments, an optional frame extraction strategy and a cropping strategy are provided. Different training samples including different numbers of video frames and different image sizes are extracted from the training video through different frame extraction strategies and cropping strategies. Taking the case of performing frame extraction first and then cropping as an example for illustration.
[0062] Figure 4 FIG. 4 shows a method for training a super-resolution model provided by an embodiment of the present application. This method can be executed by Figure 1 the terminal 120 or the server 140 shown in FIG. 4. This method includes:
[0063] Step 401: Extract k i video frames from the training video according to the i-th frame extraction strategy, and obtain a video frame sequence corresponding to the i-th video frame sequence.
[0064] Each video frame sequence in the i-th video frame sequence among the m video frame sequences includes k i video frames. i is a positive integer less than m + 1, the initial value of i is 1, and k i is a positive integer. k i is not greater than the total number of video frames of the training video. Optionally, k i is proportional to i. Exemplarily, k i = f(i), where f represents a linear function.
[0065] Exemplarily, setting the total number of video frames of the training video as T (T is a positive integer), then according to the first frame extraction strategy, T / 2 video frames are extracted from the training video to obtain a video frame sequence corresponding to the first video frame sequence. According to the second frame extraction strategy, 3*T / 4 video frames are extracted from the training video to obtain a video frame sequence corresponding to the second video frame sequence. According to the third frame extraction strategy, T video frames are extracted from the training video to obtain a video frame sequence corresponding to the third video frame sequence.
[0066] It should be noted that since it is necessary to ensure that the input training samples can provide sufficient information during the training process of the super-resolution model, a lower limit on the number of video frames needs to be set for the video frame sequence to prevent the situation where the training of the super-resolution model is affected due to too few video frames included in the video frame sequence. Optionally, the number of video frames in the i-th video frame sequence is not less than the lower limit of the number of video frames. The lower limit of the number of video frames can be a constant or a value determined based on the total number of video frames of the training video. Exemplarily, the lower limit of the number of video frames is the constant 3. Exemplarily, the lower limit of the number of video frames is 0.5*T. The lower limit of the number of video frames can be set by those skilled in the art according to actual needs.
[0067] Here, k iA video frame includes, but is not limited to, the following three methods:
[0068] (1) Randomly extract k consecutive video frames from the training video to obtain a video frame sequence corresponding to the i-th type of video frame sequence. i A video frame sequence corresponding to the i-th type of video frame sequence is obtained by extracting k consecutive video frames from the training video.
[0069] Exemplarily, the training video includes 16 video frames. The video frames from the 5th frame to the 12th frame are extracted from the training video to obtain a video frame sequence.
[0070] (2) Determine the arrangement rule of k video frames corresponding to the i-th frame extraction strategy; according to the arrangement rule, extract k video frames from the training video to obtain a video frame sequence corresponding to the i-th type of video frame sequence. i The arrangement rule of k video frames is used to represent the arrangement rule of k video frames in the training video. Exemplarily, k video frames are arranged continuously in the training video, or alternatively, k video frames are arranged at intervals in the training video. i A video frame sequence corresponding to the i-th type of video frame sequence is obtained by extracting k video frames from the training video according to the arrangement rule.
[0071] k i The arrangement rule of k video frames is used to represent the arrangement rule of k video frames in the training video. Exemplarily, k video frames are arranged continuously in the training video, or alternatively, k video frames are arranged at intervals in the training video. i Exemplarily, k video frames are arranged continuously in the training video, or alternatively, k video frames are arranged at intervals in the training video. i Exemplarily, k video frames are arranged continuously in the training video, or alternatively, k video frames are arranged at intervals in the training video. i Exemplarily, k video frames are arranged continuously in the training video, or alternatively, k video frames are arranged at intervals in the training video.
[0072] Exemplarily, the training video includes 16 video frames. It is determined that the arrangement rule of 4 video frames corresponding to the 2nd frame extraction strategy is arranged at intervals. Then, 4 video frames arranged at intervals are randomly extracted from the training video. For example, the 1st video frame, the 3rd video frame, the 5th video frame, and the 7th video frame in the training video are extracted. Another example is to extract the 1st video frame, the 4th video frame, the 7th video frame, and the 10th video frame in the training video.
[0073] Optionally, the arrangement rules corresponding to the m frame extraction strategies are the same.
[0074] (3) Randomly extract k video frames from the training video to obtain a video frame sequence corresponding to the i-th type of video frame sequence. i A video frame sequence corresponding to the i-th type of video frame sequence is obtained by extracting k video frames from the training video.
[0075] The above 3 methods for extracting video frame sequences are only used for illustrative purposes. Those skilled in the art can modify the extraction method according to actual needs and the type of super-resolution model.
[0076] Step 402: Repeat the above steps to obtain multiple video frame sequences corresponding to the i-th type of video frame sequence.
[0077] The i-th type of video frame sequence includes multiple video frame sequences, and the number of video frames in each video frame sequence is the same.
[0078] The number of times of repeating the above steps can be adjusted by those skilled in the art according to actual needs.
[0079] Step 403: Update i to i + 1, and repeat the above two steps to obtain m video frame sequences.
[0080] Steps 401 to 403 need to be repeated m times to obtain m video frame sequences, and each video frame sequence includes a different number of video frames. Exemplarily, the total number of video frames of the training video is set to T (T is a positive integer), m is set to 3, the first video frame sequence includes T / 2 video frames, the second video frame sequence includes 3*T / 4 video frames, and the third video frame sequence includes T video frames.
[0081] Step 404: According to the a-th cropping strategy, crop the i-th video frame sequence among the m video frame sequences to the a-th image size to obtain the b-th training sample among the p training samples.
[0082] Wherein, i is a positive integer less than m + 1, a is a positive integer less than n + 1, the initial values of a and i are 1, and b is a positive integer less than p + 1.
[0083] Exemplarily, the image size of the training video is , H represents the height of the training video, W represents the width of the training video, then the image size of the i-th video frame sequence is also , according to the first cropping strategy, crop the image size of the i-th video frame sequence to , according to the second cropping strategy, crop the image size of the i-th video frame sequence to .
[0084] Optionally, if the height of the training video is H and the width is W, then the a-th image size is .
[0085] It should be noted that during the training process of the super-resolution model, it is necessary to ensure that the input training samples can provide sufficient information. Therefore, it is necessary to set a lower limit for the image size to prevent the situation that the training of the super-resolution model is affected due to the too small image size included in the video frame sequence. Optionally, the image size of the b-th training sample is not less than the lower limit of the image size, and the lower limit of the image size can be a constant or a value determined based on the image size of the training video. Exemplarily, the lower limit of the image size is the constant 256×192. Exemplarily, the image size of the training sample is , and the lower limit of the image size is . The lower limit of the image size can be set by those skilled in the art according to actual needs.
[0086] The method of cropping the i-th video frame sequence here includes but is not limited to the following two ways:
[0087] (1) Determine the cropping region corresponding to the a-th cropping strategy, where the size of the cropping region is the same as the a-th image size; according to the cropping region, crop the i-th video frame sequence among the m video frame sequences to obtain the b-th training sample among the p training samples.
[0088] Exemplarily, the cropping region corresponding to the first cropping strategy is located at the lower left corner of the image, and the cropping region corresponding to the second cropping strategy is located at the upper right corner of the image. Then, it is necessary to crop the i-th video frame sequence according to different cropping regions to obtain the training samples.
[0089] (2) Randomly crop the i-th video frame sequence among the m video frame sequences to the a-th image size to obtain the b-th training sample among the p training samples.
[0090] The above two methods for cropping video frame sequences are only used for illustration, and those skilled in the art can modify the cropping method according to actual needs and the types of super-resolution models.
[0091] Step 405: Update a to a + 1, and repeat the above steps until n training samples corresponding to the i-th video frame sequence are obtained.
[0092] It is necessary to repeat step 403 n times to obtain n training samples corresponding to the i-th video frame sequence. Exemplarily, the image size of the training video is , then the image size of the i-th video frame sequence is also , and the i-th video frame sequence corresponds to 2 training samples, where the image size of one training sample is , and the image size of the other training sample is .
[0093] Step 406: Update i to i + 1, initialize a, and repeat the above two steps to obtain p training samples.
[0094] In steps 403 and 404, only n training samples corresponding to the i-th video frame sequence are obtained. However, there are a total of m video frame sequences. To obtain the training samples corresponding to all m video frame sequences, it is necessary to repeat steps 403 and 404 n times to obtain p training samples, where p = m * n.
[0095] Step 407: Arrange the p training samples in ascending order according to at least one arrangement criterion among the number of video frames and the image size.
[0096] Optionally, if the number of video frames and the image size of the first training sample and the second training sample among the p training samples are the same, then randomly arrange the order between the first training sample and the second training sample.
[0097] Step 408: According to the arrangement order of p kinds of training samples, extract training samples from the p kinds of training samples in turn to train the super-resolution model.
[0098] Exemplarily, as Figure 5 shown, when training the super-resolution model, the training samples used increase gradually according to the number of video frames and the image size. From the perspective of the image size, with the iterative training of the super-resolution model, the image size of the training samples will increase with the number of iterations. The image size of the training sample set 501 is smaller than that of the training sample set 502, and the image size of the training sample set 502 is smaller than that of the training sample set 503. From the perspective of the number of video frames, with the iterative training of the super-resolution model, the number of video frames of the training samples will increase with the number of iterations. The number of video frames of the training sample set 504 is smaller than that of the training sample set 505, and the number of video frames of the training sample set 505 is smaller than that of the training sample set 506.
[0099] The embodiments of the present application do not specifically limit the types and training methods of the super-resolution model. Exemplarily, the super-resolution model is any one of a super-resolution model based on a recurrent network and a super-resolution model based on a sliding window. Exemplarily, the training method is the error backpropagation algorithm.
[0100] In other alternative ways of the present application, the method can also be applied to other models that require video content as training samples, for example, a video classification model based on video content.
[0101] In summary, in this embodiment, the training video is split to obtain training samples, the training samples are arranged from small to large according to the number of video frames and the image size, and according to the arrangement order, different training samples are used to train the super-resolution model in stages. Since the smaller the number of video frames or the image size, the less information the training samples contain, which helps to improve the training speed. Moreover, the training of the super-resolution model in the previous stage has a guiding effect, which can guide the training of the super-resolution model in the current stage, so that the super-resolution model can learn from simple to difficult, and while maintaining the model accuracy, effectively improve the training speed.
[0102] In another implementation manner of the present application, the image size of the training video can be cropped first and then frame extraction is performed.
[0103] Figure 6 shows a training method for a super-resolution model provided by an embodiment of the present application. This method can be executed by Figure 1 the terminal 120 or the server 140 shown, and this method includes:
[0104] Step 601: According to the a-th cropping strategy, crop the training video to the a-th image size to obtain a cropped video corresponding to the a-th cropped video.
[0105] Where a is a positive integer less than n + 1, and the initial value of a is 1.
[0106] Optionally, if the height of the training video is H and the width is W, then the a-th image size is .
[0107] It should be noted that during the training process of the super-resolution model, it is necessary to ensure that the input training samples can provide sufficient information. Therefore, it is necessary to set a lower limit for the image size to prevent the situation where the training of the super-resolution model is affected due to the too small image size included in the video frame sequence. Optionally, the image size of the b-th training sample is not less than the lower limit of the image size, and the lower limit of the image size can be a constant or a value determined based on the image size of the training video.
[0108] The methods for cropping the training video here include but are not limited to the following two ways:
[0109] (1) Determine the cropping area corresponding to the a-th cropping strategy, and the size of the cropping area is the same as the a-th image size; according to the cropping area, crop the training video to obtain the a-th cropped video.
[0110] (2) Randomly crop the training video to the a-th image size to obtain the a-th cropped video.
[0111] The above two methods for cropping the training video are only used for illustrative purposes, and technicians can modify the cropping method according to actual needs and the types of super-resolution models.
[0112] Step 602: Repeat the above steps to obtain multiple cropped videos corresponding to the a-th cropped video.
[0113] Each type of cropped video includes multiple cropped videos, and the image sizes of each cropped video are the same.
[0114] The number of times of repeating the above steps can be adjusted by technicians according to actual needs.
[0115] Step 603: Update a to a + 1, and repeat the above two steps to obtain n types of cropped videos.
[0116] It is necessary to repeat Step 601 and Step 602 n times to obtain n types of training samples. Exemplarily, the image size of the first type of cropped video is , and the image size of the second type of cropped video is .
[0117] Step 604: Extract k video frames from the c-th cropped video among the n cropped videos according to the i-th frame extraction strategy, to obtain the b-th training sample among the p training samples. i The i is a positive integer less than m + 1, the initial value of i is 1, k
[0118] is a positive integer, k i is not greater than the total number of video frames of the training video, b is a positive integer less than p, and c is a positive integer less than n. Optionally, k i is proportional to i. Exemplarily, k i = f(i), where f represents a linear function. i
[0119] It should be noted that since during the training process of the super-resolution model, it is necessary to ensure that the input training samples can provide sufficient information, therefore, a lower limit on the number of video frames needs to be set for the training samples to prevent the situation where the training of the super-resolution model is affected due to too few video frames included in the training samples. Optionally, the number of video frames of the b-th training sample is not less than the lower limit of the number of video frames, and the lower limit of the number of video frames can be a constant or a value determined based on the total number of video frames of the training video.
[0120] The extraction of k i video frames includes but is not limited to the following three methods:
[0121] (1) Randomly extract k i continuous video frames from the training video to obtain a training sample corresponding to the b-th training video frame sequence.
[0122] (2) Determine the arrangement rule of k i video frames corresponding to the i-th frame extraction strategy; according to the arrangement rule, extract k i video frames from the training video to obtain a training sample corresponding to the b-th video frame sequence.
[0123] (3) Randomly extract k i video frames from the training video to obtain a training sample corresponding to the b-th video frame sequence.
[0124] The above three methods for extracting video frame sequences are only used for illustrative purposes, and those skilled in the art can modify the extraction method according to actual needs and the type of super-resolution model.
[0125] Step 605: Update i to i + 1, and repeat the above steps to obtain m training samples corresponding to the i-th cropped video.
[0126] Step 604 needs to be repeated m times to obtain m types of training samples, each type of training sample including different numbers of video frames. Exemplarily, if the total number of video frames of the training video is set to T (T is a positive integer), then according to the first frame extraction strategy, T / 2 video frames are extracted from the c-th cropped video to obtain the first type of training sample. According to the second frame extraction strategy, 3*T / 4 video frames are extracted from the c-th cropped video to obtain the second type of training sample. According to the third frame extraction strategy, T video frames are extracted from the c-th cropped video to obtain the third type of training sample.
[0127] Step 606: Update c to c + 1, initialize i, and repeat the above two steps to obtain p types of training samples.
[0128] In steps 603 and 604, only m types of training samples corresponding to the c-th cropped video are obtained. However, there are a total of n types of cropped videos. To obtain the training samples corresponding to all n types of cropped videos, steps 603 and 604 need to be repeated m times to obtain p types of training samples, where p = m * n.
[0129] Step 607: Arrange the p types of training samples in ascending order according to at least one arrangement criterion among the number of video frames and the image size.
[0130] Optionally, if the number of video frames and the image size of the first training sample and the second training sample are the same among the p types of training samples, then randomly arrange the order between the first training sample and the second training sample.
[0131] Step 608: Extract training samples from the p types of training samples in sequence according to the arrangement order of the p types of training samples to train the super-resolution model.
[0132] The embodiments of the present application do not specifically limit the types and training methods of the super-resolution model. Exemplarily, the super-resolution model is any one of a super-resolution model based on a recurrent network and a super-resolution model based on a sliding window. Exemplarily, the training method is the error backpropagation algorithm.
[0133] In other alternative ways of the present application, the method can also be applied to other models that require video content as training samples, such as a video classification model based on video content.
[0134] In summary, in this embodiment, the training video is split to obtain training samples, and the training samples are arranged from small to large according to the number of video frames and the image size. According to the arrangement order, different training samples are used to train the super-resolution model in stages. Since the smaller the number of video frames or the image size, the less information the training sample contains, which helps to improve the training speed. Moreover, the training of the super-resolution model in the previous stage has a guiding effect, which can guide the training of the super-resolution model in the current stage, enabling the super-resolution model to learn from simple to difficult, and effectively improving the training speed while maintaining the model accuracy.
[0135] In the following embodiments, as the training of the super-resolution model progresses, the learning rate of the super-resolution model will continuously decay. This results in a relatively small learning rate when switching to large image sizes and the number of video frames, which hinders the learning ability of the super-resolution model training. Therefore, it is necessary to update the learning rate of the super-resolution model.
[0136] Figure 7 shows a training method for a super-resolution model provided by an embodiment of the present application. This method can be executed by Figure 1 the terminal 120 or the server 140 shown, and this method includes:
[0137] Step 701: Determine the j-th training sample from p types of training samples according to the arrangement order of the p types of training samples.
[0138] j is a positive integer less than p, and the initial value of j is 1.
[0139] The p types of training samples are arranged according to at least one arrangement criterion among the number of video frames and the image size.
[0140] Step 702: Use the j-th training sample to train the super-resolution model until the training of the j-th training stage is completed.
[0141] Optionally, the j-th training sample includes multiple training samples, and the training of the j-th training stage includes the training process of the training samples corresponding to the j-th training sample.
[0142] Optionally, for the training of the j-th training stage, it is necessary to use the j-th training sample to perform iterative training for times, which is used to represent the number of training iterations in the j-th training stage.
[0143] Since there are a total of p types of training samples, in the embodiments of the present application, the training process of the super-resolution model is also divided into p stages, and each stage uses training samples with different numbers of video frames and image sizes for training.
[0144] Step 703: Update j to j + 1, and repeat the above two steps until the super-resolution model is trained using p types of training samples.
[0145] It should be noted that when j is updated to j + 1, the learning rate of the super-resolution model is updated.
[0146] Among them, the learning rate of the (j + 1)-th training stage after update is greater than the learning rate of the j-th training stage before update.
[0147] Optionally, the learning rate has the following formula:
[0148] ;
[0149] Among them, represents the learning rate used in the -th iteration, represents the initial learning rate used in the benchmark training method. refers to the training stage where the -th iteration is located, . is used to represent the number of training iterations in the j-th training stage. represents the total number of iterations required to train the super-resolution model. Therefore, . Exemplarily, as Figure 8 shows, the total number of training times , for the t-th iteration, the training stage it belongs to satisfies: t < .
[0150] It can be obtained from the above formula that for the training stage , when , that is, when just switching to the -th training stage, , that is to say , the learning rate is re-initialized to a relatively large value. Since the total number of iterations is always larger than , so for the training stage , the learning rate will not drop to zero, avoiding the waste of training time caused by too small a learning rate.
[0151] To sum up, when training the super-resolution model in this application, the learning rate of the super-resolution model will be dynamically adjusted. When switching training stages each time, a relatively large value is used to re-initialize the learning rate to improve the training speed of the super-resolution model.
[0152] Taking the BasicVSR (Basic Visual Super Resolution) model and the EDVR-M (Enhanced Deformable Video Restoration-Middle) model as examples, the method provided by the embodiments of the present application can accelerate the training speed of the video super-resolution model without losing the accuracy of the output result. PSNR (Peak Signal to Noise Ratio) and SSIM (Structural Similarity) are used as test parameters. REDS4 (an existing test set) is used as the test set. Among them, since the EDVR model is a sliding window-based model and the number of input video frames is fixed, the present application only changes its image size to obtain Table 1:
[0153] Table 1 Comparison table of training methods based on different video super-resolution
[0154]
[0155] "*" represents the data of the super-resolution model in the original literature. Large-Batch is also introduced to better utilize the parallelization of the GPU (Graphics Processing Unit) to accelerate training.
[0156] According to Table 1, it can be obtained that after applying the training method of the super-resolution model provided by the present application to the BasicVSR model and the EDVR-M model, the accuracy of the output result is comparable to that of the related technology. And the training method of the super-resolution model provided by the embodiments of the present application can effectively reduce the training time and improve the training efficiency.
[0157] Figure 9 The block diagram of a super-resolution training device provided by an exemplary embodiment of the present application is shown. The device 900 can be used to implement the functions of the above-mentioned super-resolution training method. The device includes:
[0158] A splitting module 901, configured to split a training video into p types of training samples, each type of training sample includes at least f training samples with the same image size and the same number of video frames, the number of video frames of each type of training sample in the p types of training samples is not greater than the number of video frames of the training video, the image size of each type of training sample in the p types of training samples is not greater than the image size of the training video, p is a positive integer greater than 1, and f is a positive integer;
[0159] The splitting module 901 is further configured to arrange the p types of training samples in ascending order according to at least one of the number of video frames and the image size.
[0160] The training module 902 is configured to sequentially extract training samples from the p types of training samples according to the arrangement order of the p types of training samples to train the super-resolution model.
[0161] In an alternative design of the present application, the splitting module 901 is further configured to extract m types of video frame sequences from the training video according to m types of frame extraction strategies, where the frame extraction strategies are used to extract video frame sequences with different numbers of video frames from the training video, and the m types of training samples correspond one-to-one to the m types of numbers of video frames; according to n types of cropping strategies, at least one of the m types of video frame sequences is cropped into samples of n types of image sizes to obtain the p types of training samples, where the cropping strategies are used to crop the images of the at least one type of video frame sequence, and n and m are positive integers.
[0162] In an alternative design of the present application, the i-th type of video frame sequence among the m types of video frame sequences includes k i video frames, where i is a positive integer less than m + 1, the initial value of i is 1, and k i is a positive integer, and k i is not greater than the total number of video frames of the training video; the splitting module 901 is further configured to extract the k i video frames from the training video according to the i-th type of frame extraction strategy to obtain a video frame sequence corresponding to the i-th type of video frame sequence; repeat the above steps to obtain multiple video frame sequences corresponding to the i-th type of video frame sequence; update i to i + 1, and repeat the above two steps to obtain the m types of video frame sequences.
[0163] In an alternative design of the present application, the splitting module 901 is further configured to randomly extract consecutive k i video frames from the training video to obtain a video frame sequence corresponding to the i-th type of video frame sequence; or, determine the arrangement rule of the k i video frames corresponding to the i-th type of frame extraction strategy; according to the arrangement rule, extract the k i video frames from the training video to obtain a video frame sequence corresponding to the i-th type of video frame sequence; or, randomly extract the k i video frames from the training video to obtain a video frame sequence corresponding to the i-th type of video frame sequence.
[0164] In an alternative design of the present application, the splitting module 901 is further configured to crop the i-th video frame sequence among the m video frame sequences into the a-th image size according to the a-th cropping strategy, to obtain the b-th training sample among the p training samples, where a is a positive integer less than m + 1, a is a positive integer less than n + 1, the initial values of a and i are 1, and b is a positive integer less than p + 1; update a to a + 1, and repeat the above steps until n training samples corresponding to the i-th video frame sequence are obtained; update i to i + 1, initialize a, and repeat the above two steps to obtain the p training samples.
[0165] In an alternative design of the present application, the splitting module 901 is further configured to determine a cropping region corresponding to the a-th cropping strategy, where the size of the cropping region is the same as the a-th image size; crop the i-th video frame sequence among the m video frame sequences according to the cropping region to obtain the b-th training sample among the p training samples; or randomly crop the i-th video frame sequence among the m video frame sequences into the a-th image size to obtain the b-th training sample among the p training samples.
[0166] In an alternative design of the present application, the training module 902 is further configured to determine the j-th training sample from the p training samples according to the arrangement order of the p training samples, where j is a positive integer less than p and the initial value of j is 1; use the j-th training sample to train the super-resolution model until the training of the j-th training stage is completed; update j to j + 1, and repeat the above two steps until the super-resolution model is trained using the p training samples.
[0167] In an alternative design of the present application, the training module 902 is further configured to update the learning rate of the super-resolution model when j is updated to j + 1; wherein, the learning rate of the (j + 1)-th training stage after update is greater than the learning rate of the j-th training stage before update.
[0168] In summary, in this embodiment, the training video is split to obtain training samples, the training samples are arranged in ascending order according to the number of video frames and the image size, and according to the arrangement order, different training samples are used to train the super-resolution model in stages. Since the smaller the number of video frames or the image size, the less information the training sample contains, which helps to improve the training speed. Moreover, the training of the super-resolution model in the previous stage has a guiding effect, which can guide the training of the super-resolution model in the current stage, enabling the super-resolution model to learn from simple to difficult, and effectively improving the training speed while maintaining the model accuracy.
[0169] Figure 10 It is a schematic structural diagram of a server shown according to an exemplary embodiment. The computer device 1000 includes a Central Processing Unit (CPU) 1001, a system memory 1004 including a Random Access Memory (RAM) 1002 and a Read-Only Memory (ROM) 1003, and a system bus 1005 connecting the system memory 1004 and the central processing unit 1001. The computer device 1000 further includes a Basic Input / Output System (Input / Output, I / O system) 1006 for facilitating information transmission between various components within the computer device, and a mass storage device 1007 for storing an operating system 1013, application programs 1014, and other program modules 1015.
[0170] The basic input / output system 1006 includes a display 1008 for displaying information and input devices 1009 such as a mouse and a keyboard for user input of information. Among them, both the display 1008 and the input devices 1009 are connected to the central processing unit 1001 through an input / output controller 1010 connected to the system bus 1005. The basic input / output system 1006 may further include an input / output controller 1010 for receiving and processing inputs from multiple other devices such as a keyboard, a mouse, or an electronic stylus. Similarly, the input / output controller 1010 also provides output to a display screen, a printer, or other types of output devices.
[0171] The mass storage device 1007 is connected to the central processing unit 1001 through a mass storage controller (not shown) connected to the system bus 1005. The mass storage device 1007 and its associated computer-readable medium provide non-volatile storage for the computer device 1000. That is to say, the mass storage device 1007 may include computer-readable media (not shown) such as a hard disk or a Compact Disc Read-Only Memory (CD-ROM) drive.
[0172] Without loss of generality, the computer device-readable medium may include a computer device storage medium and a communication medium. The computer device storage medium includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer device-readable instructions, data structures, program modules, or other data. The computer device storage medium includes RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM, digital video disc (DVD), or other optical storage, magnetic tape cartridges, tapes, disk storage, or other magnetic storage devices. Of course, those skilled in the art will know that the computer device storage medium is not limited to the above several types. The above system memory 1004 and mass storage device 1007 can be collectively referred to as memory.
[0173] According to various embodiments of the present disclosure, the computer device 1000 can also run on a remote computer device on the network through a network such as the Internet. That is, the computer device 1000 can be connected to the network 1011 through the network interface unit 1012 connected to the system bus 1005. Or rather, the network interface unit 1012 can also be used to connect to other types of networks or remote computer device systems (not shown).
[0174] The memory further includes one or more programs. The one or more programs are stored in the memory, and the central processing unit 1001 implements all or part of the steps of the above method for training the super-resolution model by executing the one or more programs.
[0175] This application also provides a computer-readable storage medium, in which at least one instruction, at least one segment of program, code set, or instruction set is stored. The at least one instruction, at least one segment of program, code set, or instruction set is loaded and executed by a processor to implement the method for training the super-resolution model provided in the above method embodiments.
[0176] This application also provides a computer program product or a computer program. The above computer program product or computer program includes computer instructions, and the above computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the above computer instructions from the computer-readable storage medium, and the processor executes the above computer instructions, so that the computer device executes the method for training the super-resolution model provided in the above aspect embodiments.
[0177] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.
[0178] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disc, or the like.
[0179] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A training method for a super-resolution model, characterized in that, The method includes: Splitting the training video into p types of training samples, each type of training sample including at least f training samples with the same image size and the same number of video frames. The number of video frames of each type of training sample in the p types of training samples is not greater than the number of video frames of the training video, and the image size of each type of training sample in the p types of training samples is not greater than the image size of the training video. p is a positive integer greater than 1, and f is a positive integer; Arranging the p types of training samples in ascending order according to at least one of the arrangement criteria of the number of video frames and the image size; Extracting training samples from the p types of training samples in sequence according to the arrangement order of the p types of training samples to perform staged training on the super-resolution model, so that the super-resolution model learns from simple to difficult.
2. The method according to claim 1, wherein The splitting of the training video into p types of training samples includes: Extracting m types of video frame sequences from the training video according to m types of frame extraction strategies. The frame extraction strategies are used to extract video frame sequences with different numbers of video frames from the training video. The m types of training samples correspond one-to-one with the m types of numbers of video frames; Cropping at least one of the m types of video frame sequences into samples of n image sizes according to n types of cropping strategies to obtain the p types of training samples. The cropping strategies are used to crop the images of the at least one type of video frame sequence. n and m are positive integers.
3. The method according to claim 2, wherein The i-th video frame sequence among the m video frame sequences includes k i video frames, where i is a positive integer less than m + 1, the initial value of i is 1, and k i is a positive integer, and k i is not greater than the total number of video frames of the training video; The extracting of m types of video frame sequences from the training video according to m types of frame extraction strategies includes: According to the i-th frame extraction strategy, extract the k i video frames from the training video to obtain a video frame sequence corresponding to the i-th video frame sequence; Repeating the above steps to obtain multiple video frame sequences corresponding to the i-th type of video frame sequence; Updating i to i + 1 and repeating the above two steps to obtain the m types of video frame sequences.
4. The method according to claim 3, wherein According to the i-th frame extraction strategy, extract the k i video frames from the training video to obtain a video frame sequence corresponding to the i-th video frame sequence, including: Randomly extract consecutive k of the training videos to obtain a video frame sequence corresponding to the i-th video frame sequence; i video frames, to obtain a video frame sequence corresponding to the i-th video frame sequence; Alternatively, determine the arrangement rule of the k i video frames corresponding to the i-th frame extraction strategy; according to the arrangement rule, extract the k i video frames from the training video to obtain a video frame sequence corresponding to the i-th video frame sequence; Alternatively, randomly extract the k i video frames from the training video to obtain a video frame sequence corresponding to the i-th video frame sequence.
5. The method according to claim 2, characterized in that The cropping of at least one of the m types of video frame sequences into samples of n image sizes according to n types of cropping strategies to obtain the p types of training samples includes: Cropping the i-th type of video frame sequence among the m types of video frame sequences into the a-th image size according to the a-th cropping strategy to obtain the b-th training sample among the p types of training samples. i is a positive integer less than m + 1, a is a positive integer less than n + 1, the initial values of a and i are 1, and b is a positive integer less than p + 1; Updating a to a + 1 and repeating the above steps until the training samples corresponding to the i-th video frame sequence are obtained; Updating i to i + 1, initializing a, and repeating the above two steps to obtain the p types of training samples.
6. The method according to claim 5, wherein The cropping of the video sequence corresponding to the i-th type of video frame sequence among the m types of video frame sequences into the a-th image size according to the a-th cropping strategy to obtain the b-th training sample among the p types of training samples includes: Determining a cropping area corresponding to the a-th cropping strategy, where the size of the cropping area is the same as the a-th image size; and cropping the i-th type of video frame sequence among the m types of video frame sequences according to the cropping area to obtain the b-th training sample among the p types of training samples. Alternatively, randomly crop the i-th video frame sequence among the m video frame sequences to the a-th image size to obtain the b-th training sample among the p training samples.
7. The method according to any one of claims 1 to 6, characterized in that, The step of sequentially extracting training samples from the p training samples according to the arrangement order of the p training samples to perform staged training on the super-resolution model includes: Determine the j-th training sample from the p training samples according to the arrangement order of the p training samples, where j is a positive integer less than p, and the initial value of j is 1; Use the j-th training sample to train the super-resolution model until the training of the j-th training stage is completed; Update j to j + 1, and repeat the above two steps until the super-resolution model is trained using the p training samples.
8. The method according to claim 7, wherein The method further includes: When j is updated to j + 1, update the learning rate of the super-resolution model; Wherein, the learning rate of the (j + 1)-th training stage after update is greater than the learning rate of the j-th training stage before update.
9. A training device for a super-resolution model, characterized in that, The device includes: A splitting module, configured to split a training video into p training samples, each training sample includes at least f training samples with the same image size and the same number of video frames, the number of video frames of each training sample among the p training samples is not greater than the number of video frames of the training video, the image size of each training sample among the p training samples is not greater than the image size of the training video, p is a positive integer greater than 1, and f is a positive integer; The splitting module is further configured to arrange the p training samples in ascending order according to at least one arrangement criterion of the number of video frames and the image size; A training module, configured to sequentially extract training samples from the p training samples according to the arrangement order of the p training samples to perform staged training on the super-resolution model, so that the super-resolution model learns from simple to difficult.
10. The device according to claim 9, wherein The splitting module is further configured to extract m video frame sequences from the training video according to m frame extraction strategies, the frame extraction strategies are used to extract video frame sequences with different numbers of video frames from the training video, and the m training samples correspond one-to-one with the m numbers of video frames; according to n cropping strategies, crop at least one of the m video frame sequences into samples of n image sizes to obtain the p training samples, the cropping strategies are used to crop the images of the at least one video frame sequence, and n and m are positive integers.
11. The device according to claim 10, characterized in that, The i-th video frame sequence among the m video frame sequences includes k i video frames, where i is a positive integer less than m + 1, the initial value of i is 1, and k i is a positive integer, and k i is not greater than the total number of video frames of the training video; The splitting module is further configured to extract the k video frames from the training video according to the i-th frame extraction strategy, so as to obtain a video frame sequence corresponding to the i-th video frame sequence; repeat the above steps to obtain multiple video frame sequences corresponding to the i-th video frame sequence; update i to i + 1, and repeat the above two steps to obtain the m video frame sequences. i video frames, obtaining a video frame sequence corresponding to the i-th video frame sequence; repeating the above steps to obtain multiple video frame sequences corresponding to the i-th video frame sequence; updating i to i + 1, and repeating the above two steps to obtain the m video frame sequences.
12. The device according to claim 10, wherein The splitting module is further configured to crop the i-th video frame sequence among the m video frame sequences into the a-th image size according to the a-th cropping strategy to obtain the b-th training sample among the p training samples, where i is a positive integer less than m + 1, a is a positive integer less than n + 1, the initial values of a and i are 1, and b is a positive integer less than p + 1; Update the a to a + 1, and repeat the above steps until the training sample corresponding to the i-th video frame sequence is obtained; update the i to i + 1, initialize the a, and repeat the above two steps to obtain the p training samples.
13. A computer device, characterized in that, The computer device includes: a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the training method of the super-resolution model according to any one of claims 1 to 8.
14. A computer-readable storage medium, characterized in that, At least one program code is stored in the computer-readable storage medium, and the program code is loaded and executed by the processor to implement the training method of the super-resolution model according to any one of claims 1 to 8.
15. A computer program product comprising a computer program or instructions, characterized in that, When the computer program or instruction is executed by the processor, it implements the training method of the super-resolution model according to any one of claims 1 to 8.
Citation Information
Patent Citations
Model prediction method, device and equipment and storage medium
CN110009109A
Image shadow detection method based on deep unsupervised learning
CN113436115A