An image processing method, an image display method, a model training method and an apparatus
A larger magnification is achieved in image processing by cascading, which solves the problem that the model can only be suitable for magnification within a fixed range in the prior art, and improves the flexibility and use range of image processing.
Patent Information
- Application Number
- CN201910667659.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-07-23
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2039-07-23
AI Technical Summary
The Meta-SR model in existing super-resolution technology can only be applied to magnifications within a fixed range, resulting in reduced flexibility in image processing and reduced usage range.
A larger magnification is achieved through cascading. Using the magnification of the power relationship only requires training one model, reducing the number of models that need to be trained and saved.
It improves the flexibility of image processing, increases the scope of use of image amplification, realizes image super-resolution with any magnification, and ensures performance.
Smart Images

Figure CN110363709B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and particularly to an image processing method, an image display method, a model training method and an apparatus. Background Art
[0002] In various application scenarios, limited by the cost of image acquisition devices, or the video image transmission bandwidth, or the technical bottlenecks of the imaging modality itself, it is not always possible to obtain high-definition images. Therefore, the super-resolution technology has emerged. The super-resolution technology can reconstruct the corresponding high-resolution image from the observed low-resolution image, and has important application values in fields such as monitoring devices, satellite images and medical images.
[0003] Currently, in order to better implement the super-resolution technology, a super-resolution arbitrary magnification network (AMagnification-Arbitrary Network for Super-Resolution, Meta-SR) has been designed. Meta-SR realizes the image super-resolution of an arbitrary magnification factor of a single model by adding a filter generation module based on meta-learning and modifying the upsampling method.
[0004] However, since the input information of Meta-SR includes the magnification factor, the trained Meta-SR can only be applicable to magnification factors within a fixed range, and corresponding magnification processing cannot be achieved for some magnification factors. Therefore, the flexibility of image processing is reduced, and the usage range of image magnification is reduced. Summary of the Invention
[0005] Embodiments of the present application provide an image processing method, an image display method, a model training method and an apparatus, which do not need to train different models for different magnification factors. The image processing model at a certain magnification factor can achieve a larger magnification factor through a cascading method, and only one model needs to be trained for magnification factors with a power relationship, thereby reducing the number of models that need to be trained and saved. Therefore, the flexibility is improved, and the usage range of image magnification is increased.
[0006] In view of this, a first aspect of the present application provides an image processing method, including:
[0007] Obtain an image to be processed, where the image to be processed corresponds to a first magnification factor;
[0008] Obtain the scaling factor information corresponding to the image to be processed, where the scaling factor information is used to indicate the magnification factor for magnifying the image to be processed, or the reduction factor for reducing the image to be processed;
[0009] Determine the number of cascades according to the scaling factor information, where the number of cascades is an integer greater than or equal to 1, and the number of cascades represents the number of times of processing the image to be processed using the same image processing model;
[0010] According to the number of cascades, obtain the target image corresponding to the image to be processed through the image processing model, where the target image corresponds to a second magnification factor, the second magnification factor has an associated relationship with the number of cascades, and the second magnification factor is different from the first magnification factor.
[0011] The second aspect of the present application provides an image display method, including:
[0012] Obtain an image to be processed, where the image to be processed corresponds to a first magnification factor;
[0013] Receive an image adjustment instruction, where the image adjustment instruction carries an image magnification parameter, and the image magnification parameter is used to indicate the magnification factor for magnifying the image to be processed;
[0014] In response to the image adjustment instruction, determine the number of cascades according to the image magnification parameter, where the number of cascades is an integer greater than or equal to 1, and the number of cascades represents the number of times of processing the image to be processed using the same image processing model;
[0015] According to the number of cascades, obtain the target image corresponding to the image to be processed through the image processing model, where the target image corresponds to a second magnification factor, the second magnification factor has an associated relationship with the number of cascades, and the second magnification factor is greater than the first magnification factor;
[0016] Display the target image.
[0017] The third aspect of the present application provides a model training method, including:
[0018] Obtain a set of training image samples, where the set of training image samples belongs to a set of training image data, the set of training image samples includes at least one training image sample, each training image sample includes a first image, a second image, and a third image, the first image and the second image have a preset sampling magnification factor, and the second image and the third image have the preset sampling magnification factor;
[0019] Generate a random value, where the random value is greater than or equal to 0 and less than or equal to 1;
[0020] Determine a set of objects to be trained from the set of image samples to be trained according to the random numerical value and the ratio numerical value, where the set of objects to be trained includes at least one object to be trained, and each object to be trained includes the second image and the third image, or each object to be trained includes a first predicted image and the third image, and the first predicted image is obtained after the first image passes through the image processing model to be trained;
[0021] Train the image processing model to be trained using the set of objects to be trained to obtain an image processing model.
[0022] A fourth aspect of the present application provides an image processing apparatus, including:
[0023] An acquisition module, configured to acquire an image to be processed, where the image to be processed corresponds to a first magnification;
[0024] The acquisition module is further configured to acquire scaling factor information corresponding to the image to be processed, where the scaling factor information is used to indicate a magnification factor for magnifying the image to be processed or a reduction factor for reducing the image to be processed;
[0025] A determination module, configured to determine a cascade number according to the scaling factor information acquired by the acquisition module, where the cascade number is an integer greater than or equal to 1, and the cascade number represents the number of times of processing the image to be processed using the same image processing model;
[0026] The acquisition module is further configured to, according to the cascade number determined by the determination module, acquire a target image corresponding to the image to be processed through the image processing model, where the target image corresponds to a second magnification, the second magnification has an associated relationship with the cascade number, and the second magnification is different from the first magnification.
[0027] A fifth aspect of the present application provides an image display apparatus, including:
[0028] An acquisition module, configured to acquire an image to be processed, where the image to be processed corresponds to a first magnification;
[0029] A receiving module, configured to receive an image adjustment instruction, where the image adjustment instruction carries an image magnification parameter, and the image magnification parameter is used to indicate a magnification factor for magnifying the image to be processed;
[0030] A determination module, configured to, in response to the image adjustment instruction received by the receiving module, determine a cascade number according to the image magnification parameter, where the cascade number is an integer greater than or equal to 1, and the cascade number represents the number of times of processing the image to be processed using the same image processing model;
[0031] The obtaining module is further configured to obtain a target image corresponding to the image to be processed through the image processing model according to the cascading times determined by the determining module, where the target image corresponds to a second magnification factor, the second magnification factor has an associated relationship with the cascading times, and the second magnification factor is greater than the first magnification factor;
[0032] The display module is configured to display the target image obtained by the obtaining module.
[0033] A sixth aspect of the present application provides an image processing model training device, including:
[0034] An obtaining module, configured to obtain a set of image samples to be trained, where the set of image samples to be trained belongs to a set of image data to be trained, the set of image samples to be trained includes at least one image sample to be trained, each image sample to be trained includes a first image, a second image, and a third image, the first image and the second image have a preset sampling magnification factor, and the second image and the third image have the preset sampling magnification factor;
[0035] A generating module, configured to generate a random value, where the random value is greater than or equal to 0 and less than or equal to 1;
[0036] A determining module, configured to determine a set of objects to be trained from the set of image samples to be trained according to the random value generated by the generating module and a ratio value, where the set of objects to be trained includes at least one object to be trained, each object to be trained includes the second image and the third image, or each object to be trained includes a first predicted image and the third image, and the first predicted image is obtained after the first image passes through the image processing model to be trained;
[0037] A training module, configured to train the image processing model to be trained by using the set of objects to be trained determined by the determining module to obtain an image processing model.
[0038] In a possible design, in the first implementation manner of the sixth aspect of the embodiments of the present application,
[0039] The determining module is specifically configured to determine whether the random value is greater than the ratio value;
[0040] If the random value is greater than the ratio value, obtain the first image and the third image corresponding to the image sample to be trained from the set of image samples to be trained;
[0041] Obtain the first predicted image corresponding to the first image through the image processing model to be trained;
[0042] Generate the objects to be trained in the set of objects to be trained based on the first predicted image and the third image.
[0043] In a possible design, in the second implementation manner of the sixth aspect of the embodiments of the present application,
[0044] The determining module is specifically configured to determine whether the random value is greater than the ratio value;
[0045] If the random value is less than or equal to the ratio value, obtain the second image and the third image corresponding to the image sample to be trained from the set of image samples to be trained;
[0046] Generate the objects to be trained in the set of objects to be trained based on the second image and the third image.
[0047] In a possible design, in the third implementation manner of the sixth aspect of the embodiments of the present application,
[0048] The obtaining module is specifically configured to obtain a first image sample to be trained, where the first image sample to be trained includes the first image, the second image, and the third image;
[0049] Obtain a second image sample to be trained, where the second image sample to be trained includes a fourth image, the first image, and the second image, and the fourth image and the first image have a preset sampling magnification;
[0050] Obtain a third image sample to be trained, where the third image sample to be trained includes a fifth image, the fourth image, and the first image, and the fifth image and the fourth image have a preset sampling magnification;
[0051] Generate an image sample to be trained based on the first image sample to be trained, the second image sample to be trained, and the third image sample to be trained.
[0052] In a possible design, in the fourth implementation manner of the sixth aspect of the embodiments of the present application,
[0053] The obtaining module is further configured to obtain an offset value and a slope value before the determining module determines the set of objects to be trained from the set of image samples to be trained according to the random value and the ratio value;
[0054] The obtaining module is further configured to obtain the number of iterations corresponding to the set of image samples to be trained;
[0055] The determining module is further configured to determine the ratio value corresponding to the set of to-be-trained image samples according to the offset value, the slope value, and the iteration number corresponding to the set of to-be-trained image samples obtained by the obtaining module.
[0056] In a possible design, in the fifth implementation manner of the sixth aspect of the embodiments of the present application,
[0057] The training module is specifically configured to obtain, through the to-be-trained image processing model, a second predicted image corresponding to each to-be-trained object in the set of to-be-trained objects;
[0058] According to the second predicted image corresponding to each to-be-trained object and the expected image corresponding to each to-be-trained object, use a target loss function to determine network model parameters;
[0059] Use the network model parameters to train the to-be-trained image processing model to obtain the image processing model;
[0060] The training module specifically determines the network model parameters in the following manner:
[0061]
[0062] Wherein, the L(θ) represents the target loss function, the θ represents the network model parameters, the n represents the total number of to-be-trained image samples in the set of to-be-trained image samples, and the x i represents the i-th to-be-trained object in the set of to-be-trained objects, and the net(x i ,θ) represents the second predicted image corresponding to the i-th to-be-trained object, and the y i represents the expected image corresponding to the i-th to-be-trained object.
[0063] A seventh aspect of the present application provides a network device, including: a memory, a transceiver, a processor, and a bus system;
[0064] Wherein, the memory is used to store a program;
[0065] The processor is configured to execute the program in the memory, including the following steps:
[0066] Obtain a to-be-processed image, where the to-be-processed image corresponds to a first magnification;
[0067] Obtain the scaling factor information corresponding to the to-be-processed image, where the scaling factor information is used to indicate the multiple of magnifying the to-be-processed image or the multiple of reducing the to-be-processed image;
[0068] Determine the number of cascades according to the scaling factor information, where the number of cascades is an integer greater than or equal to 1, and the number of cascades represents the number of times of processing the image to be processed using the same image processing model;
[0069] According to the number of cascades, obtain the target image corresponding to the image to be processed through the image processing model, where the target image corresponds to a second magnification factor, the second magnification factor has an associated relationship with the number of cascades, and the second magnification factor is different from the first magnification factor;
[0070] The bus system is used to connect the memory and the processor so that the memory and the processor can communicate.
[0071] The ninth aspect of the present application provides a terminal device, including: a memory, a transceiver, a processor, and a bus system;
[0072] Wherein, the memory is used to store programs;
[0073] The processor is used to execute the programs in the memory, including the following steps:
[0074] Obtain an image to be processed, where the image to be processed corresponds to a first magnification factor;
[0075] Receive an image adjustment instruction, where the image adjustment instruction carries an image magnification parameter, and the image magnification parameter is used to indicate the magnification factor for magnifying the image to be processed;
[0076] In response to the image adjustment instruction, determine the number of cascades according to the image magnification parameter, where the number of cascades is an integer greater than or equal to 1, and the number of cascades represents the number of times of processing the image to be processed using the same image processing model;
[0077] According to the number of cascades, obtain the target image corresponding to the image to be processed through the image processing model, where the target image corresponds to a second magnification factor, the second magnification factor has an associated relationship with the number of cascades, and the second magnification factor is greater than the first magnification factor;
[0078] Display the target image;
[0079] The bus system is used to connect the memory and the processor so that the memory and the processor can communicate.
[0080] The ninth aspect of the present application provides a server, including: a memory, a transceiver, a processor, and a bus system;
[0081] Wherein, the memory is used to store programs;
[0082] The processor is used to execute the program in the memory, including the following steps:
[0083] Obtain a set of image samples to be trained, where the set of image samples to be trained belongs to a set of image data to be trained, the set of image samples to be trained includes at least one image sample to be trained, each image sample to be trained includes a first image, a second image, and a third image, the first image and the second image have a preset sampling magnification, and the second image and the third image have the preset sampling magnification;
[0084] Generate a random value, where the random value is greater than or equal to 0 and less than or equal to 1;
[0085] Determine a set of objects to be trained from the set of image samples to be trained according to the random value and a ratio value, where the set of objects to be trained includes at least one object to be trained, each object to be trained includes the second image and the third image, or each object to be trained includes a first predicted image and the third image, and the first predicted image is obtained after the first image passes through an image processing model to be trained;
[0086] Use the set of objects to be trained to train the image processing model to be trained to obtain an image processing model;
[0087] The bus system is used to connect the memory and the processor to enable the memory and the processor to communicate with each other.
[0088] The tenth aspect of the present application provides a computer-readable storage medium, in which instructions are stored, and when they run on a computer, the computer is enabled to execute the methods described in the above aspects.
[0089] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0090] In an embodiment of the present application, an image processing method is provided. First, an image to be processed is obtained, where the image to be processed corresponds to a first magnification factor. Then, magnification factor information corresponding to the image to be processed is obtained, where the magnification factor information is used to indicate the magnification factor for magnifying the image to be processed or the reduction factor for reducing the image to be processed. Next, the number of cascades is determined according to the magnification factor information, where the number of cascades represents the number of times the image to be processed passes through the image processing model. Finally, according to the number of cascades, the target image corresponding to the image to be processed is obtained through the image processing model, where the target image corresponds to a second magnification factor, the second magnification factor has an associated relationship with the number of cascades, and the second magnification factor is different from the first magnification factor. Through the above method, image scaling processing can be realized according to the number of cascades, without the need to train additional models for different magnification factors. Instead, the image processing model truly realizes image super-resolution with any magnification factor, and realizes image super-resolution with any magnification factor while ensuring performance. Thus, the flexibility of image processing is improved, and the application range of image magnification is increased. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] Figure 1 It is a schematic architecture diagram of an image display system in an embodiment of the present application;
[0092] Figure 2 It is a schematic diagram of an image display interface based on a client in an embodiment of the present application;
[0093] Figure 3 It is a schematic diagram of an embodiment of the method for image processing in an embodiment of the present application;
[0094] Figure 4A It is a schematic diagram of an embodiment of realizing 2-fold image magnification processing based on the number of cascades in an embodiment of the present application;
[0095] Figure 4B It is a schematic diagram of an embodiment of realizing 4-fold image magnification processing based on the number of cascades in an embodiment of the present application;
[0096] Figure 4C It is a schematic diagram of an embodiment of realizing 8-fold image magnification processing based on the number of cascades in an embodiment of the present application;
[0097] Figure 5 It is a schematic diagram of an embodiment of the method for image display in an embodiment of the present application;
[0098] Figure 6 It is a schematic diagram of an embodiment of the method for model training in an embodiment of the present application;
[0099] Figure 7 It is a schematic structural diagram of an image processing model in an embodiment of the present application;
[0100] Figure 8 It is a schematic flowchart of a method for model training in an embodiment of the present application;
[0101] Figure 9 It is a schematic diagram of an embodiment of an image processing apparatus in an embodiment of the present application;
[0102] Figure 10 It is a schematic diagram of an embodiment of an image display apparatus in an embodiment of the present application;
[0103] Figure 11 It is a schematic diagram of an embodiment of an image processing model training apparatus in an embodiment of the present application;
[0104] Figure 12 It is a schematic structural diagram of a network device in an embodiment of the present application;
[0105] Figure 13 It is a schematic structural diagram of a terminal device in an embodiment of the present application;
[0106] Figure 14 It is a schematic structural diagram of a server in an embodiment of the present application. Detailed implementation manners
[0107] The embodiments of the present application provide an image processing method, an image display method, a model training method and apparatus, which do not need to train different models for different magnification factors. The image processing model at a certain magnification factor can achieve a larger magnification factor through a cascading method, and only one model needs to be trained for magnification factors with a power relationship, thereby reducing the number of models that need to be trained and saved. Thus, the flexibility is improved and the usage range of image magnification is increased.
[0108] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0109] It should be understood that the image processing method and the image display method provided in this application can be applied to the image super-resolution scenario. In the image super-resolution scenario, more information than a single image can be obtained based on training samples, and then through a specific algorithm, this additional information can be incorporated into the original image to obtain a high-quality and high-definition image. The application of the image processing method is becoming increasingly widespread, covering many fields such as medicine, banking, and exploration. For old photos stored for many years, the image display method provided in this application can make their details come to life; in the face of the bandwidth pressure of network transmission, the image display method provided in this application can first compress and transmit the image and then restore it with an ultra-clear algorithm, which can greatly reduce the amount of transmitted data.
[0110] For the sake of easy understanding, this application proposes an image display method, which is applicable to scenarios related to image transmission or image quality enhancement, such as instant messaging applications (APPs), video playback APPs, comic applications, and interactive applications (such as multiplayer online battle arena games). This method is applied to Figure 1 the image display system shown in Figure 1 , Figure 1 which is a schematic diagram of the architecture of the image display system in an embodiment of this application. As shown in the figure, taking a comic APP as an example, when a user views comics through the comic APP, they need to continuously request high-definition images from the server. However, downloading high-definition images not only consumes the user's mobile phone traffic but also the latency during the transmission process will cause a poor experience of image loading lag when the user flips through the pages. Therefore, in the image display method provided in this application, when the user's traffic warning is triggered or the network status is poor, after the user requests the server to send a low-resolution image through the comic APP, the comic APP locally uses an image processing model to process the low-resolution image. The comic APP sends the low-resolution image, and the comic APP uses the image processing model to magnify the low-resolution image to obtain the corresponding high-resolution image. In this way, it is not necessary for the server to directly send high-resolution images to the comic APP, thus saving traffic for users using this comic APP.
[0111] It should be noted that the comic APP is deployed on a terminal device, and similar APPs (such as instant messaging APPs, video playback APPs, and interactive applications) can all be deployed on the terminal device as clients. Among them, the terminal device includes but is not limited to tablet computers, laptop computers, personal digital assistants, mobile phones, voice interaction devices, and personal computers (PCs), and no limitation is made here.
[0112] Specifically, the image display method provided in this application will be described below in combination with an application scenario. Please refer toFigure 2 , Figure 2 This is a schematic diagram of an image display interface based on a client in an embodiment of the present application. As shown in the figure, taking the client as a comic APP for example, the server sends a low-resolution comic image to the comic APP. The comic APP provides a continuously adjustable status bar for the user. The value of the status bar has a corresponding relationship with the local magnification factor. The user locally magnifies the low-resolution image by adjusting the value of the status bar. The larger the value of the status bar, the higher the resolution at which the user can choose to view the comic. Figure 2 The left image in has a magnification rate of 20%. If the user hopes to see a clearer image, they can drag the status bar. When dragged to 100%, the right image as shown in Figure 2 will be displayed.
[0113] Higher magnification factors require more cascading times, and the running time of image super-resolution will increase accordingly. Some users think the smoothness during page turning is more important when viewing comics, and these users can choose a lower magnification factor. While some users think the image quality is more important when viewing comics, and these users can choose a higher magnification factor. Since the solution provided in the present application can achieve any magnification factor, users can choose an appropriate magnification factor by adjusting the status bar and find an optimal trade-off between viewing smoothness and image quality.
[0114] It can be understood that in practical applications, the magnification rate of the image can be set as a multiple of 2, such as 2 times magnification rate, 4 times magnification rate, 8 times magnification rate, 16 times magnification rate, etc. It can also be 3 times magnification rate, or 5 times magnification rate, or other magnification rates. The present application can achieve a 16 times magnification rate of the image, however, this should not be construed as a limitation to the present application.
[0115] Combined with the above introduction, the method for image processing in the present application will be introduced below. Please refer to Figure 3 , an embodiment of the method for image processing in an embodiment of the present application includes:
[0116] 101. Obtain an image to be processed, where the image to be processed corresponds to a first magnification factor;
[0117] In this embodiment, the image processing device obtains the image to be processed. It can be understood that the image processing device can be deployed on a network device, and the network device can be a terminal device or a server, which is not limited here. The image to be processed is at a first magnification factor, for example, the first magnification factor is 10%.
[0118] It should be noted that the format of the image to be processed includes, but is not limited to, bitmap (BMP) format, Personal Computer Exchange (PCX) format, Tag Image File Format (TIFF), Graphics Interchange Format (GIF), Joint Photographic Expert Group (JPEG) format, Tagged Graphics (TGA) format, Exchangeable Image File Format (EXIF), Portable Network Graphics (PNG) format, Scalable Vector Graphics (SVG) format, Drawing Exchange Format (DXF), and Encapsulated Post Script (EPS) format.
[0119] 102. Obtain the scaling factor information corresponding to the image to be processed, where the scaling factor information is used to indicate the magnification factor for magnifying the image to be processed or the reduction factor for reducing the image to be processed.
[0120] In this embodiment, the image processing device obtains the scaling factor information corresponding to the image to be processed. The scaling factor information may refer to the magnification factor for magnifying the image to be processed or the reduction factor for reducing the image to be processed. For example, the scaling factor information is to magnify by 8 times, or the scaling information is to reduce by 4 times.
[0121] 103. Determine the number of cascades according to the scaling factor information, where the number of cascades is an integer greater than or equal to 1, and the number of cascades represents the number of times of processing the image to be processed using the same image processing model.
[0122] In this embodiment, the image processing device determines the number of cascades according to the scaling factor information. Specifically, assume that an image processing model is used to magnify (or reduce) an image by N times, where N can be a non-zero positive number, such as 1.2 or 2, etc. If the scaling factor information is to magnify by 2N times, then two magnifying processes of the image processing model are required, that is, the number of cascades is 2. Thus, it can be seen that the number of cascades represents the number of times the image to be processed passes through the image processing model.
[0123] It can be understood that the image processing device can pre-store multiple image processing models to be selected corresponding to different magnifications, and then select a suitable image processing model from the image processing models to be selected according to the zoom factor information, and determine the number of cascades. If the image processing device only pre-stores one image processing model, then this image processing model can be directly used.
[0124] 104. According to the number of cascades, obtain the target image corresponding to the image to be processed through the image processing model, where the target image corresponds to the second magnification, the second magnification has an associated relationship with the number of cascades, and the second magnification is different from the first magnification.
[0125] In this embodiment, the image processing device obtains the target image corresponding to the image to be processed through the image processing model according to the number of cascades. Specifically, assume that an image processing model A is used to magnify an image by 2 times. For ease of understanding, please refer to Figure 4A , Figure 4A FIG. is a schematic diagram of an embodiment for realizing 2-fold image magnification processing based on the number of cascades in the embodiment of the present application. As shown in the figure, when magnifying the image to be processed by 2 times, the zoom factor information is determined to be 2 times. Therefore, it is necessary to pass through the image processing model A once, that is, the number of cascades is 1, and input the image to be processed with X times (i.e., the first magnification) into the image processing model A, so as to obtain the target image with 2X times (i.e., the second magnification).
[0126] Please refer to Figure 4B , Figure 4B FIG. is a schematic diagram of an embodiment for realizing 4-fold image magnification processing based on the number of cascades in the embodiment of the present application. As shown in the figure, when magnifying the image to be processed by 4 times, the zoom factor information is determined to be 4 times. Therefore, it is necessary to pass through the image processing model A twice, that is, the number of cascades is 2, and input the image to be processed with X times (i.e., the first magnification) into the image processing model, so as to obtain the target image with 4X times (i.e., the second magnification).
[0127] Please refer to Figure 4C , Figure 4C FIG. is a schematic diagram of an embodiment for realizing 8-fold image magnification processing based on the number of cascades in the embodiment of the present application. As shown in the figure, when magnifying the image to be processed by 8 times, the zoom factor information is determined to be 8 times. Therefore, it is necessary to pass through the image processing model A three times, that is, the number of cascades is 1, and input the image to be processed with X times (i.e., the first magnification) into the image processing model A, so as to obtain the target image with 2X times (i.e., the second magnification).
[0128] For ease of introduction, please refer to Table 1. Table 1 shows the associated relationship between the magnification of the image to be processed and the number of cascades. Assume that an image processing model is used to magnify an image by 2 times.
[0129] Table 1
[0130] First magnification factor Magnification factor information Cascade times Second magnification factor X 2 1 2X X 4 2 4X X 8 3 8X X 2^N N (2^N)X
[0131] For ease of introduction, please refer to Table 2, which shows the correlation between the reduction factor of the image to be processed and the number of cascades. Assume that an image processing model is used to reduce an image by a factor of 2. It can be understood that, assuming an image processing model is used to enlarge an image by a factor of 1.2, if the magnification information is 1.2 times, then the number of cascades is 1, and if the magnification information is 1.44 times, then the number of cascades is 2.
[0132] Table 2
[0133] First magnification factor Magnification factor information Cascade times Second magnification factor X 1 / 2 1 1 / 2X X 1 / 4 2 1 / 4X X 1 / 8 3 1 / 8X X 1 / (2^N) N 1 / (2^N)X
[0134] Thus, there is a corresponding relationship between the first magnification of the image to be processed, the second magnification of the target image, the scaling factor information (magnification information or reduction factor information), and the number of cascades. In this application, an image processing model capable of implementing a cascade structure is designed, expanding a model that can only be applied to a single magnification to an image processing model that can achieve a larger magnification through cascading, and using the image processing model to achieve image super-resolution at any magnification.
[0135] In an embodiment of this application, an image processing method is provided. First, an image to be processed is obtained, where the image to be processed corresponds to a first magnification. Then, the scaling factor information corresponding to the image to be processed is obtained, where the scaling factor information is used to indicate the magnification factor for enlarging the image to be processed or the reduction factor for reducing the image to be processed. Next, the number of cascades is determined according to the scaling factor information, where the number of cascades represents the number of times the image to be processed passes through the image processing model. Finally, according to the number of cascades, the target image corresponding to the image to be processed is obtained through the image processing model, where the target image corresponds to a second magnification, the second magnification has a correlation with the number of cascades, and the second magnification is different from the first magnification. By the above method, image scaling processing can be realized according to the number of cascades, without the need to train additional models for different magnifications, but using the image processing model to truly achieve image super-resolution at any magnification, and achieving image super-resolution at any magnification while ensuring performance. Thus, the flexibility of image processing is improved, and the application range of image enlargement is increased.
[0136] Combined with the above introduction, the method for image display in this application will be introduced below. Please refer to Figure 5 , an embodiment of the method for image display in an embodiment of this application includes:
[0137] 201. Obtain an image to be processed, where the image to be processed corresponds to a first magnification;
[0138] In this embodiment, an image display device obtains an image to be processed. It can be understood that the image display device can be deployed on a terminal device, specifically a client. The image to be processed is at a first magnification factor, for example, the first magnification factor is 10%.
[0139] It should be noted that the format of the image to be processed includes but is not limited to BMP format, PCX format, TIFF, GIF, JPEG format, TGA format, EXIF, PNG format, SVG format, DXF, and EPS format.
[0140] 202. Receive an image adjustment instruction, where the image adjustment instruction carries an image magnification parameter, and the image magnification parameter is used to indicate the magnification multiple for magnifying the image to be processed;
[0141] In this embodiment, the image display device receives an image adjustment instruction. The image adjustment instruction can be a sliding instruction or an input instruction. The image magnification parameter is carried in the image adjustment instruction, such as magnifying 8 times or 16 times, etc. The image magnification parameter is used to indicate the magnification multiple for magnifying the image to be processed.
[0142] 203. In response to the image adjustment instruction, determine the number of cascades according to the image magnification parameter, where the number of cascades represents the number of times the image to be processed passes through the image processing model;
[0143] In this embodiment, the image display device responds to the image adjustment instruction, determines the image magnification parameter according to the image adjustment instruction, and then determines the number of cascades. Specifically, assume that an image processing model is used to magnify an image by N times, where N can be a non-zero positive number, such as 1.2 or 2, etc. If the image magnification parameter is 2N times, then it needs to go through two magnifications by the image processing model, that is, the number of cascades is 2. That is, the number of cascades represents the number of times the image to be processed passes through the image processing model.
[0144] 204. According to the number of cascades, obtain the target image corresponding to the image to be processed through the image processing model, where the target image corresponds to a second magnification factor, the second magnification factor has an associated relationship with the number of cascades, and the second magnification factor is greater than the first magnification factor;
[0145] In this embodiment, the image display device obtains the target image corresponding to the image to be processed through the image processing model according to the number of cascades. Specifically, assume that an image processing model is used to magnify an image by 2 times. When magnifying the image to be processed by 2^N times, the image magnification parameter is determined to be 2^N times. Therefore, it is necessary to pass through the image processing model N times, that is, the number of cascades is N. The image to be processed at X times (i.e., the first magnification factor) is input into the image processing model to obtain the target image at (2^N)X times (i.e., the second magnification factor).
[0146] 205. Display the target image.
[0147] In this embodiment, after the image display device obtains the target image, it can display the target image.
[0148] It can be understood that in practical applications, the image processing model can also include multiple models with different magnification factors. Suppose an image processing model A is used to magnify an image by 2 times, and an image processing model B is used to magnify an image by 3 times. Then, after passing through an image processing model A and an image processing model B, the image to be processed can be magnified by 6 times. Suppose an image processing model A is used to magnify an image by 1.2 times, and an image processing model B is used to magnify an image by 2 times. Then, after passing through an image processing model A and an image processing model B, the image to be processed can be magnified by 2.4 times. Thus, it can be seen that by training an image processing model that magnifies by N times, this image processing model can magnify any resolution image by N times. Whether the input of this image processing model is a real image or the output of other image processing models, the trained N - times image processing model can achieve magnification multiples of powers of N through multiple cascades of itself, and can achieve magnification multiples of N through cascades with other magnification - factor image processing models.
[0149] In order to achieve magnification processing with arbitrary magnification factors, in practical applications, only some common magnification factors need to be saved, such as the model parameters for 2 times, 3 times, 5 times, and 1.2 times.
[0150] In the embodiment of the present application, an image display method is provided. First, an image to be processed is obtained, and then an image adjustment instruction is received. The image adjustment instruction carries an image magnification parameter. In response to the image adjustment instruction, the number of cascades is determined according to the image magnification parameter. Next, according to the number of cascades, the target image corresponding to the image to be processed is obtained through the image processing model. Finally, the target image is displayed. Through the above method, magnification processing of the image can be achieved according to the number of cascades, without the need to train additional models for different magnification factors. Instead, the image processing model truly realizes image super - resolution with arbitrary magnification factors. For image compression and transmission, the traffic bandwidth of the forwarding server can be greatly saved during the transmission process. The client decodes to obtain a relatively low - definition image, and a high - definition image is obtained through the solution provided by the present application, thereby saving traffic, improving the transmission rate, and improving the picture quality as needed.
[0151] Combined with the above introduction, the method for model training in the present application will be introduced below. Please refer to Figure 6 , an embodiment of the method for model training in the embodiment of the present application includes:
[0152] 301. Obtain a set of training image samples. Among them, the set of training image samples belongs to the set of training image data. The set of training image samples includes at least one training image sample. Each training image sample includes a first image, a second image, and a third image. The first image and the second image have a preset sampling ratio, and the second image and the third image have a preset sampling ratio;
[0153] In this embodiment, the model training device obtains a set of training image samples. It can be understood that the model training device can be deployed on a server. The set of training image samples belongs to the set of training image data. The set of training image data can include at least one set of training image samples. Among them, the number of samples (batchsize) used for each training in a set of training image samples can be 16, 32, 64, or 128, or other integers. In theory, the larger the value of batchsize, the better the training performance. However, the larger the batchsize, the higher the requirement for the video memory of the graphics card (that is, the higher the requirement for the hardware). Therefore, a reasonable batchsize needs to be selected in actual training, such as 128.
[0154] A set of training image samples includes at least one training image sample. A training image sample can be expressed as [x_pre, x, y], where x_pre represents the image obtained by downsampling x by 1 / r, x represents the image obtained by downsampling y by 1 / r. That is, the first image is represented as x_pre, the second image is represented as x, and the third image is represented as y. The preset sampling ratio is 1 / r. For example, the format of a training image sample can be [1 / 4, 1 / 2, 1]. Among them, the "1" in [1 / 4, 1 / 2, 1] represents the resolution of the original image in the training image sample, "1 / 2" represents the image resolution obtained by downsampling the original image by 1 / 2, and "1 / 4" represents the image resolution obtained by downsampling the original image by 1 / 4.
[0155] It can be understood that the spatial resolution of the first image (x_pre) can be 16*16, the spatial resolution of the second image (x) can be 32*32, and the spatial resolution of the third image (y) can be 64*64. And random vertical flipping and horizontal flipping can be performed on each training image sample in the set of training image samples.
[0156] 302. Generate a random value, where the random value is greater than or equal to 0 and less than or equal to 1;
[0157] In this embodiment, the model training device generates uniformly distributed random values. Assume that the distribution function of the random variable X is F(X), and {Xi, i = 1, 2...} are independent and identically distributed as F(X). Then, the first-order observers {X1, X2, X3,...} of {Xi, i = 1, 2...} are called the random number sequence of the distribution F(X), simply referred to as random data. The range of the random values in this application is greater than or equal to 0 and less than 1.
[0158] The generators that generate these random values include, but are not limited to, linear congruential generators, feedback shift register generators, and combined generators.
[0159] 303. Determine the set of objects to be trained from the set of image samples to be trained according to the random values and ratio values. Among them, the set of objects to be trained includes at least one object to be trained. Each object to be trained includes a second image and a third image, or each object to be trained includes a first predicted image and a third image. The first predicted image is obtained after the first image passes through the image processing model to be trained.
[0160] In this embodiment, when the model training device performs training within one epoch, it is necessary to first select the objects to be trained according to the generated random values and ratio values. Specifically, if the random value is greater than the ratio value, obtain the set of objects to be trained from the set of image samples to be trained. The set of objects to be trained includes the first predicted image (x') and the third image (y). The first predicted image (x') is obtained after the first image (x_pre) passes through the image processing model to be trained. Conversely, if the random value is less than or equal to the ratio value, obtain the set of objects to be trained from the set of image samples to be trained. The set of objects to be trained includes the second image (x) and the third image (y).
[0161] After completing one iteration of training, the model training device will regenerate a random value, and then judge the size relationship between the random value and the ratio value again, and select the set of objects to be trained for training again based on the above introduction.
[0162] 304. Use the set of objects to be trained to train the image processing model to be trained, and obtain the image processing model.
[0163] In this embodiment, the model training device uses the set of objects to be trained to train the image processing model to be trained. Specifically, taking the object to be trained including the second image and the third image as an example for introduction, the input of the image processing model to be trained is the second image (x), the output of the image processing model to be trained is net(x), and the ideal output is the third image (y). The process of updating the network model parameters and reducing the loss function is the training of the network, so as to obtain the image processing model.
[0164] For ease of understanding, please refer to Figure 7 , Figure 7 which is a schematic structural diagram of the image processing model in the embodiment of the present application. As shown in the figure, in the present application, the Residual Dense Network (RDN) is used as an example for introduction. The input of the RDN is a low-resolution image (LR), and the output is a high-resolution image (HR). Among them, the RDN mainly includes 4 modules, namely the shallow feature extraction network (SFENet) module, the residual dense block (RDB), the dense feature fusion (DFF) module, and the up-sampling network (UPNet) module.
[0165] The SFENet module includes the first 2 convolutional layers. The RDB module mainly integrates the residual block and the dense block, and combines the two to form the RDB. The DFF module includes two parts: global residual learning and global feature fusion. The UPNet represents the final up-sampling and convolutional operations of the network, realizing the magnification operation of the input image.
[0166] In the embodiment of the present application, a method for model training is provided. First, a set of training image samples is obtained, where the set of training image samples belongs to the set of training image data. Then, random values are generated, and then, according to the random values and the ratio values, a set of training objects is determined from the set of training image samples. Finally, the set of training objects is used to train the image processing model to obtain the image processing model. By the above method, an image processing model with a fixed magnification can be trained, without the need to train different models for different magnification factors. The image processing model at a certain magnification factor can achieve a larger magnification factor through a cascading method. For magnification factors with a power relationship, only one model needs to be trained, thereby reducing the number of models that need to be trained and saved. Therefore, the flexibility is improved, the usage range of image magnification is increased, and the super-resolution of images with any magnification factor is realized by using the cascading times of the image processing model. Moreover, the super-resolution of images with any magnification factor is achieved while ensuring the performance. Thus, the flexibility of image processing is improved, and the usage range of image magnification is increased.
[0167] Optionally, in the above Figure 6Based on the corresponding embodiments, in an alternative embodiment of the model training method provided in the embodiments of the present application, determining a set of objects to be trained from a set of image samples to be trained according to a random value and a ratio value may include:
[0168] Determine whether the random value is greater than the ratio value;
[0169] If the random value is greater than the ratio value, obtain the first image and the third image corresponding to the image sample to be trained from the set of image samples to be trained;
[0170] Obtain the first predicted image corresponding to the first image through the image processing model to be trained;
[0171] Generate the object to be trained in the set of objects to be trained according to the first predicted image and the third image.
[0172] In this embodiment, a method for obtaining a set of objects to be trained is introduced. For the convenience of introduction, please refer to Figure 8 , Figure 8 which is a schematic flowchart of the model training method in the embodiments of the present application. As shown in the figure, specifically:
[0173] In step S1, prepare a set of image data to be trained, and the set of image data to be trained includes multiple sets of image samples to be trained;
[0174] In step S2, build an image processing model, which is the image processing model to be trained subsequently;
[0175] In step S3, when the iteration number is i, randomly obtain a set of image samples to be trained with a quantity of batchsize from the set of image data to be trained. Among them, the format of the set of image samples to be trained is [x_pre, x, y], where x_pre represents the first image, x represents the second image, and y represents the third image. The purpose of using x_pre is to train an image processing model with a magnification factor of r. It can not only magnify the real images (x_pre and x) by r times, but also magnify the output image (x' = net(x_pre)) by r times. Therefore, both the second image (x) and the first predicted image (x') obtained from the first image (x_pre) through the network are required during the training process;
[0176] In step S4, generate a random value t, and t follows a uniform distribution on [0, 1];
[0177] In step S5, calculate the ratio value ratio according to the iteration number i, and determine whether the random value t is greater than the ratio value ratio. If the random value t is greater than the ratio value ratio, jump to step S7; if the random value t is less than or equal to the ratio value ratio, enter step S6;
[0178] In step S6, if the random value t is less than or equal to the ratio value ratio, the second image (x) and the third image (y) corresponding to the to-be-trained image sample are obtained from the set of to-be-trained image samples, and the second image (x) and the third image (y) are used as the to-be-trained objects to train the to-be-trained image processing model.
[0179] In step S7, if the random value t is greater than the ratio value ratio, the first image (x_pre) and the third image (y) corresponding to the to-be-trained image sample are obtained from the set of to-be-trained image samples, the first image (x_pre) is input into the to-be-trained image processing model, and the to-be-trained image processing model outputs the first predicted image (x’). The first predicted image (x’) and the third image (y) are used as the to-be-trained objects.
[0180] In step S8, the to-be-trained objects obtained in step S7 are used to train the to-be-trained image processing model. The training objective is that the first predicted image (x’) can obtain the expected third image (y) after passing through the to-be-trained image processing model. The third image (y) is not limited to the original image resolution 1, and images with resolutions of 1 / 2 or 1 / 4 can also be used. The purpose of this design is that it is hoped that the network input with a magnification factor of r after training can be an image with a resolution of 1 / r, or an image with a resolution of 1 / r*r or 1 / r*r*r, etc.
[0181] In step S9, it is judged whether the iteration number i has ended. If the iteration number i has ended, it means that one use of each to-be-trained image sample in the set of to-be-trained image samples has been completed, and then step S10 is entered. Otherwise, if the iteration number i has not ended, it jumps to step S3.
[0182] In step S10, when the iteration number i has ended, it is judged whether the current model training has reached the termination condition. If the termination condition is reached, it jumps to step S12. Otherwise, if the termination condition is not reached, it enters step S11. It can be understood that the termination condition determines the timing of the end of training, including but not limited to the following termination conditions:
[0183] The first termination condition is that the iteration number i reaches the maximum iteration number, and the maximum iteration number is the maximum value of the iteration number set in advance.
[0184] The second termination condition is that the value of the target loss function does not decrease for a period of time during the training process.
[0185] The third termination condition is that the performance on the validation set does not improve for a period of time.
[0186] In step S11, if the termination condition is not met, the number of iterations is increased by one, i.e., i = i + 1, and the ratio value ratio = function(i).
[0187] In step S12, if the termination condition is met, the trained image processing model is output.
[0188] Secondly, in the embodiments of the present application, a method for selecting a set of objects to be trained is provided. That is, if the random value is greater than the ratio value, the first image and the third image corresponding to the image sample to be trained are obtained from the set of image samples to be trained. Then, the first predicted image corresponding to the first image is obtained through the image processing model to be trained, and the object to be trained in the set of objects to be trained is generated based on the first predicted image and the third image. Through the above method, the situation of continuous adjustment of network model parameters is fully considered, and based on the uncertainty of the random value, different samples will be selected during the training process, thereby improving the reliability of training and the diversity of sample types.
[0189] Optionally, based on the above Figure 6 In an optional embodiment of the model training method provided in the embodiments of the present application corresponding to each of the above embodiments, determining the set of objects to be trained from the set of image samples to be trained according to the random value and the ratio value may include:
[0190] Determine whether the random value is greater than the ratio value;
[0191] If the random value is less than or equal to the ratio value, the second image and the third image corresponding to the image sample to be trained are obtained from the set of image samples to be trained;
[0192] The object to be trained in the set of objects to be trained is generated according to the second image and the third image.
[0193] In this embodiment, another method for obtaining the set of objects to be trained is introduced. For the convenience of introduction, please refer to Figure 8 , Figure 8 is a schematic flowchart of the model training method in the embodiments of the present application. As shown in the figure, specifically, in step S6, if the random value t is less than or equal to the ratio value ratio, the second image (x) and the third image (y) corresponding to the image sample to be trained are obtained from the set of image samples to be trained, and the second image (x) and the third image (y) are used as the objects to be trained to train the image processing model to be trained.
[0194] The training objective is to enable the second image (x) to obtain the desired third image (y) through the image processing model to be trained. The third image (y) is not limited to the original image resolution 1, and images with resolutions of 1 / 2 or 1 / 4 can also be used. The purpose of this design is that it is hoped that the network input with a magnification factor of r after training can be an image with a resolution of 1 / r, or an image with a resolution of 1 / r*r or 1 / r*r*r, etc.
[0195] Secondly, in the embodiments of the present application, a method for selecting a set of objects to be trained is provided. That is, if the random value is less than or equal to the ratio value, the second image and the third image corresponding to the image sample to be trained are obtained from the set of image samples to be trained, and then the object to be trained in the set of objects to be trained is generated according to the second image and the third image. Through the above method, the situation of continuous adjustment of the network model parameters is fully considered, and based on the uncertainty of the random value, different samples will be selected during the training process, thereby improving the reliability of training and the diversity of sample types.
[0196] Optionally, on the basis of the above Figure 6 corresponding embodiments, in an optional embodiment of the model training method provided by the embodiments of the present application, obtaining a set of image samples to be trained may include:
[0197] Obtain a first image sample to be trained, where the first image sample to be trained includes a first image, a second image, and a third image;
[0198] Obtain a second image sample to be trained, where the second image sample to be trained includes a fourth image, a first image, and a second image, and the fourth image and the first image have a preset sampling magnification;
[0199] Obtain a third image sample to be trained, where the third image sample to be trained includes a fifth image, a fourth image, and a first image, and the fifth image and the fourth image have a preset sampling magnification;
[0200] Generate an image sample to be trained according to the first image sample to be trained, the second image sample to be trained, and the third image sample to be trained.
[0201] In this embodiment, the content included in a set of training image samples is introduced. If only the first set of training image samples including the first image, the second image, and the third image are prepared, it may not achieve the expected effect during the training process. For example, if only the first set of training image samples in the format of [1 / 4, 1 / 2, 1] are prepared, and the magnification factor of the image processing model is 2 times, then during the training process, only the original image (the third image), the 1 / 2 downsampled image (the second image), and the 1 / 4 downsampled image (the first image) can be seen. It is feasible to use the image processing model once to magnify the 1 / 2 downsampled image (the second image), or cascade the image processing model twice to magnify the 1 / 4 downsampled image (the first image). However, if the 1 / 8 downsampled image (the fourth image) is magnified by cascading the image processing model three times, a performance degradation may occur because the 1 / 8 downsampled image (the fourth image) does not appear in the set of training image samples.
[0202] To improve the training performance, it is desired that the trained image processing model can be cascaded three or four times. Thus, the second set of training image samples and the third set of training image samples are added. If the format of the first set of training image samples is [1 / 4, 1 / 2, 1], the format of the second set of training image samples can be [1 / 8, 1 / 4, 1 / 2], and the format of the third set of training image samples can be [1 / 16, 1 / 8, 1 / 4]. Therefore, the set of training image samples is a concept of a union.
[0203] During the actual training process, the set of training image samples includes at least one of the first set of training image samples, the second set of training image samples, and the third set of training image samples. It can be understood that in practice, more diverse sets of training image samples can be selected according to the number of cascades of the image processing model. The first set of training image samples, the second set of training image samples, and the third set of training image samples selected here are only for illustration and should not be construed as a limitation to this application.
[0204] Secondly, in the embodiment of this application, the content included in a set of training image samples is provided, that is, the first set of training image samples, the second set of training image samples, and the third set of training image samples are obtained. The first set of training image samples includes the first image, the second image, and the third image. The second set of training image samples includes the fourth image, the first image, and the second image. The third set of training image samples includes the fifth image, the fourth image, and the first image. Finally, the training image samples are generated according to the first set of training image samples, the second set of training image samples, and the third set of training image samples. Through the above method, more types of samples can be obtained, so as to obtain more input distributions during the training process, and the diversity of the training set is increased in the form of a sample union.
[0205] Optionally, in the aboveFigure 6 Based on the corresponding various embodiments, in an alternative embodiment of the method for model training provided by the embodiments of the present application, before determining the set of objects to be trained from the set of image samples to be trained according to the random value and the ratio value, it may further include:
[0206] Obtain the offset value and the slope value;
[0207] Obtain the number of iterations corresponding to the set of image samples to be trained;
[0208] Determine the ratio value corresponding to the set of image samples to be trained according to the offset value, the slope value, and the number of iterations corresponding to the set of image samples to be trained.
[0209] In this embodiment, a method for calculating the ratio value is introduced. The ratio value (ratio) is determined according to the number of iterations, and the ratio value is a monotonically non-increasing function of the number of iterations, that is:
[0210] ratio = function(i);
[0211] where 0 < ratio_min < 1, and three common functions will be introduced below.
[0212] The first is a linear function, that is;
[0213] ratio = max(ratio_min, k - c·epoch);
[0214] where k represents the offset value and c represents the slope value.
[0215] The second is an inverse sigmoid function, that is;
[0216] ratio = max(ratio_min, 1 / (1 + exp(c·(epoch - k))));
[0217] where k represents the offset value and c represents the slope value.
[0218] The third is an exponential function, that is;
[0219] ratio = max(ratio_min, c epoch ), c < 1;
[0220] where c represents the slope value.
[0221] In this application, an inverse sigmoid function can be used for experiments. The ratio_min is set to 0.8, the offset value k is set to 50, and the slope value c is set to 0.1. The offset value k and the slope value c determine the rate of decrease as the number of iterations decreases, and the ratio_min determines the minimum value that the number of iterations can take.
[0222] In the initial stage of model training, the value of the number of iterations is relatively large, and the randomly generated random numerical values have a high probability of being less than the number of iterations. Most of the time, the second image is used for training. At this time, it is desired that the image processing model has an effect of magnifying the second image by two times. As the training progresses, the value of the number of iterations gradually becomes smaller, and the proportion of using the first predicted image output by the image processing model for training gradually increases. The image processing model gradually transitions from being able to only magnify the second image to being able to magnify the image output by the image processing model. At this time, the super-resolution network gradually has the characteristics of cascading and becomes a general image processing model that magnifies by two times.
[0223] Secondly, in the embodiments of the present application, a method for determining the ratio value is provided, that is, first obtain the offset value and the slope value, then obtain the number of iterations corresponding to the set of training image samples to be trained, and finally determine the ratio value corresponding to the set of training image samples to be trained according to the offset value, the slope value, and the number of iterations corresponding to the set of training image samples to be trained. Through the above method, the ratio value can be updated based on the change of the number of iterations, thereby increasing the diversity of training samples and further improving the reliability of model training.
[0224] Optionally, on the basis of the corresponding embodiments above Figure 6 In an optional embodiment of the method for training a model provided by the embodiments of the present application, training the image processing model to be trained with a set of objects to be trained to obtain an image processing model may include:
[0225] Obtain the second predicted image corresponding to each object to be trained in the set of objects to be trained through the image processing model to be trained;
[0226] Determine the network model parameters by using the target loss function according to the second predicted image corresponding to each object to be trained and the expected image corresponding to each object to be trained;
[0227] Train the image processing model to be trained by using the network model parameters to obtain an image processing model;
[0228] Determining the network model parameters by using the target loss function according to the second predicted image corresponding to each object to be trained and the expected image corresponding to each object to be trained may include:
[0229] Determine the network model parameters in the following manner:
[0230]
[0231] Among them, L(θ) represents the target loss function, θ represents the network model parameters, n represents the total number of training image samples in the set of training image samples to be trained, and x i represents the i-th training object to be trained in the set of training objects to be trained, and net(x i , θ) represents the second predicted image corresponding to the i-th training object to be trained, and y i is the expected image corresponding to the i-th training object to be trained.
[0232] In this embodiment, how to train the image processing model to be trained will be introduced. Specifically, for each iteration number, there is a ratio value. During the training at this iteration number, each time a set of training objects to be trained is obtained. Among them, the set of training objects to be trained may only include the second image and the third image, or only include the first predicted image and the third image. The images included in the set of training objects to be trained depend on whether the random value is greater than the ratio value. Assume that there are 10,000 sets of training objects to be trained and the ratio value is 0.8. Since each set of training objects to be trained generates a random value uniformly distributed between [0, 1], 10,000 random values will be generated. Based on the above assumption, generally speaking, 80% of the 10,000 random values within [0, 1] are less than 0.8, and 20% are greater than 0.8. Therefore, in this epoch, 80% of the samples will select the second image (x) and the third image (y) to train the network, and 20% of the samples will select the first predicted image (x') and the third image (y) to train the network.
[0233] In the initial stage of training, the ratio value is 1. At this time, the set of training objects to be trained all uses the second image (x) and the third image (y). The subsequent ratio value will decrease as the number of iterations increases, so that there are gradually more samples to train the image processing model, that is, the first predicted image (x') and the third image (y) are used to train the network. It can be seen that the training of the image processing model can only be magnified from the second image (x) to the third image (y) in the initial stage, and gradually transitions to being able to magnify the first image (x_pre) to the first predicted image (x'), and then from the first predicted image (x') to the third image (y), so that the image processing model has a cascading property.
[0234] It is understandable that the second image (x) needs to be enlarged by a factor of r to obtain the third image (y), and the image processing model can only achieve an additional magnification of r times on the basis of the magnification of r times. Therefore, it is necessary to first use the second image (x) and the third image (y) to enable the image processing model to magnify by r times, and then consider magnifying the first image (x_pre) by r times to obtain the first predicted image (x'), and then magnifying the first predicted image (x') by r times to obtain the third image (y).
[0235] Specifically, the i-th object to be trained x i (i.e., the second image or the first predicted image) can be input into the image processing model to be trained, and the second predicted image net(x i , θ) is output by the image processing model to be trained, and the expected real output image is y i (i.e., the third image). Therefore, the following objective loss function is used for calculation:
[0236]
[0237] Among them, L(θ) represents the objective loss function, which can specifically be the mean square error function, θ represents the network model parameters, n represents the total number of image samples to be trained in the set of image samples to be trained,
[0238] When the objective loss function reaches the minimum value, the network model parameters are output. This application can use the adaptive moment estimation (Adam) optimization algorithm to optimize the network model parameters θ, and the initial value of the learning rate is 1e -4 , and the learning rate is multiplied by 0.1 every 100 iteration times.
[0239] Furthermore, in the embodiments of this application, a method for model training is provided. First, the second predicted image corresponding to each object to be trained in the set of objects to be trained is obtained through the image processing model to be trained, and then, according to the second predicted image corresponding to each object to be trained and the expected image corresponding to each object to be trained, the network model parameters are determined using the objective loss function, and then the image processing model to be trained is trained using the network model parameters to obtain the image processing model. Through the above method, on the one hand, a specific implementation basis is provided for the training of the model, thereby improving the reliability of model training. On the other hand, based on the above training method, there is no need to modify the network structure of the image processing model to be trained. Only the feature extraction module is retained and the upsampling module based on meta-learning is added to achieve multiple magnification factors with a single model, thereby improving the compatibility of model training.
[0240] For the sake of convenience in introduction, the technical solution provided by this application will be described below in combination with experimental results. Please refer to Table 3. Table 3 shows the Peak Signal to Noise Ratio (PSNR) values when using different methods at magnification factors of 2 and 4 on the common test set B100. The larger the PSNR value, the better the performance.
[0241] Table 3
[0242] Magnification factor Bicubic SRCNN MemNet RDN Meta-SR This solution X2 29.56 31.36 32.08 32.34 32.35 32.17 X4 25.96 26.90 27.40 27.72 27.75 27.61
[0243] Among them, the schemes used for comparison with this scheme in the experiment are the bicubic interpolation algorithm (Bicubic), the super-resolution algorithm based on convolutional neural network (Image Super-Resolution Using Deep Convolutional Networks, SRCNN), the deep persistent memory network (MemNet), RDN, and the super-resolution network with arbitrary magnification factor (A Magnification-Arbitrary Network for Super-Resolution, Meta-SR). Except for Meta-SR and this scheme, for the magnification factors of X2 and X4, 2 independent network models are trained respectively for other schemes, while this scheme only trains one image processing model for X2, and the X2 image processing model can be used twice to achieve the magnification factor of X4.
[0244] As can be seen from Table 3, the training method proposed in this scheme is comparable to other methods in terms of performance. Therefore, this experiment also proves the feasibility of the cascade structure.
[0245] It can be understood that Meta-SR realizes that one model is applicable to magnification factors within a certain range. For example, Meta-SR realizes magnification factors at intervals of 0.1 from 1 to 4 times. The training method proposed in this scheme enables the image processing model to have the characteristics of cascade. Meta-SR has realized continuous magnification from 1 to 4 times. Cascade of 2 Meta-SRs can realize continuous magnification from 1 to 16 times, and cascade of 3 Meta-SRs can realize continuous magnification from 1 to 64 times. Therefore, by combining the training method provided by this application with the network structure of Meta-SR, a single network model can be made applicable to any magnification factor.
[0246] The image processing device in this application will be described in detail below. Please refer to Figure 9 , Figure 9Schematic diagram of an embodiment of an image processing apparatus in an embodiment of the present application. The image processing apparatus 40 includes:
[0247] An acquisition module 401, configured to acquire an image to be processed, where the image to be processed corresponds to a first magnification;
[0248] The acquisition module 401 is further configured to acquire scaling factor information corresponding to the image to be processed, where the scaling factor information is used to indicate a magnification factor for magnifying the image to be processed or a reduction factor for reducing the image to be processed;
[0249] A determination module 402, configured to determine the number of cascades according to the scaling factor information acquired by the acquisition module 401, where the number of cascades is an integer greater than or equal to 1, and the number of cascades represents the number of times of processing the image to be processed using the same image processing model;
[0250] The acquisition module 401 is further configured to obtain a target image corresponding to the image to be processed through the image processing model according to the number of cascades determined by the determination module 402, where the target image corresponds to a second magnification, the second magnification has an association relationship with the number of cascades, and the second magnification is different from the first magnification.
[0251] In this embodiment, the acquisition module 401 acquires an image to be processed, where the image to be processed corresponds to a first magnification, the acquisition module 401 acquires scaling factor information corresponding to the image to be processed, where the scaling factor information is used to indicate a magnification factor for magnifying the image to be processed or a reduction factor for reducing the image to be processed, the determination module 402 determines the number of cascades according to the scaling factor information acquired by the acquisition module 401, where the number of cascades is an integer greater than or equal to 1, and the number of cascades represents the number of times of processing the image to be processed using the same image processing model, the acquisition module 401 obtains a target image corresponding to the image to be processed through the image processing model according to the number of cascades determined by the determination module 402, where the target image corresponds to a second magnification, the second magnification has an association relationship with the number of cascades, and the second magnification is different from the first magnification.
[0252] In an embodiment of the present application, an image processing device is provided. First, an image to be processed is obtained, where the image to be processed corresponds to a first magnification factor. Then, magnification factor information corresponding to the image to be processed is obtained, where the magnification factor information is used to indicate the magnification factor for magnifying the image to be processed or the reduction factor for reducing the image to be processed. Next, the number of cascades is determined according to the magnification factor information, where the number of cascades represents the number of times the image to be processed passes through the image processing model. Finally, according to the number of cascades, the target image corresponding to the image to be processed is obtained through the image processing model, where the target image corresponds to a second magnification factor, the second magnification factor has an associated relationship with the number of cascades, and the second magnification factor is different from the first magnification factor. Through the above method, image scaling processing can be realized according to the number of cascades, without the need to train additional models for different magnification factors, but the image processing model is used to truly realize image super-resolution with any magnification factor, and image super-resolution with any magnification factor is achieved while ensuring performance. Thereby, the flexibility of image processing is improved, and the usage range of image magnification is increased.
[0253] The image display device in the present application will be described in detail below. Please refer to Figure 10 , Figure 10 which is a schematic diagram of an embodiment of the image display device in an embodiment of the present application. The image display device 50 includes:
[0254] An acquisition module 501, configured to acquire an image to be processed, where the image to be processed corresponds to a first magnification factor;
[0255] A receiving module 502, configured to receive an image adjustment instruction, where the image adjustment instruction carries an image magnification parameter, and the image magnification parameter is used to indicate the magnification factor for magnifying the image to be processed;
[0256] A determination module 503, configured to respond to the image adjustment instruction received by the receiving module 502 and determine the number of cascades according to the image magnification parameter, where the number of cascades is an integer greater than or equal to 1, and the number of cascades represents the number of times of processing the image to be processed using the same image processing model;
[0257] The acquisition module 501 is further configured to obtain the target image corresponding to the image to be processed through the image processing model according to the number of cascades determined by the determination module 503, where the target image corresponds to a second magnification factor, the second magnification factor has an associated relationship with the number of cascades, and the second magnification factor is greater than the first magnification factor;
[0258] A display module 504, configured to display the target image acquired by the acquisition module 501.
[0259] In this embodiment, an acquisition module 501 acquires an image to be processed, where the image to be processed corresponds to a first magnification factor. A reception module 502 receives an image adjustment instruction, where the image adjustment instruction carries an image magnification parameter, and the image magnification parameter is used to indicate the magnification multiple for magnifying the image to be processed. A determination module 503, in response to the image adjustment instruction received by the reception module 502, determines the number of cascades according to the image magnification parameter, where the number of cascades is an integer greater than or equal to 1, and the number of cascades represents the number of times of processing the image to be processed using the same image processing model. The acquisition module 501 acquires a target image corresponding to the image to be processed through the image processing model according to the number of cascades determined by the determination module 503, where the target image corresponds to a second magnification factor, the second magnification factor has an associated relationship with the number of cascades, and the second magnification factor is greater than the first magnification factor. A display module 504 displays the target image acquired by the acquisition module 501.
[0260] In an embodiment of the present application, an image display method is provided. First, an image to be processed is acquired, and then an image adjustment instruction is received. The image adjustment instruction carries an image magnification parameter. In response to the image adjustment instruction, the number of cascades is determined according to the image magnification parameter. Next, according to the number of cascades, a target image corresponding to the image to be processed is acquired through an image processing model. Finally, the target image is displayed. Through the above method, image magnification processing can be achieved according to the number of cascades. There is no need to train an additional model for different magnification factors, but the image processing model is used to truly achieve image super-resolution with any magnification factor. For image compression and transmission, the traffic bandwidth of the forwarding server can be greatly saved during the transmission process. A relatively low-resolution image is decoded at the client side, and a high-resolution image is obtained through the solution provided by the present application, thereby saving traffic, improving the transmission rate, and improving the picture quality as needed.
[0261] The image processing model training device in the present application will be described in detail below. Please refer to Figure 11 , Figure 11 which is a schematic diagram of an embodiment of the image processing model training device in an embodiment of the present application. The image processing model training device 60 includes:
[0262] An acquisition module 601 is configured to acquire a set of image samples to be trained, where the set of image samples to be trained belongs to a set of image data to be trained. The set of image samples to be trained includes at least one image sample to be trained, and each image sample to be trained includes a first image, a second image, and a third image. The first image and the second image have a preset sampling magnification factor, and the second image and the third image have the preset sampling magnification factor;
[0263] A generation module 602, configured to generate a random value, where the random value is greater than or equal to 0 and less than or equal to 1;
[0264] A determination module 603, configured to determine a set of objects to be trained from the set of image samples to be trained obtained by the acquisition module 601 according to the random value and the ratio value generated by the generation module 602, where the set of objects to be trained includes at least one object to be trained, and each object to be trained includes the second image and the third image, or each object to be trained includes a first predicted image and the third image, and the first predicted image is obtained after the first image passes through the image processing model to be trained;
[0265] A training module 604, configured to train the image processing model to be trained by using the set of objects to be trained determined by the determination module 603 to obtain an image processing model.
[0266] In this embodiment, the acquisition module 601 acquires a set of image samples to be trained, where the set of image samples to be trained belongs to a set of image data to be trained, the set of image samples to be trained includes at least one image sample to be trained, each image sample to be trained includes a first image, a second image, and a third image, the first image and the second image have a preset sampling magnification, and the second image and the third image have the preset sampling magnification. The generation module 602 generates a random value, where the random value is greater than or equal to 0 and less than or equal to 1. The determination module 603 determines a set of objects to be trained from the set of image samples to be trained obtained by the acquisition module 601 according to the random value and the ratio value generated by the generation module 602, where the set of objects to be trained includes at least one object to be trained, and each object to be trained includes the second image and the third image, or each object to be trained includes a first predicted image and the third image, and the first predicted image is obtained after the first image passes through the image processing model to be trained. The training module 604 trains the image processing model to be trained by using the set of objects to be trained determined by the determination module 603 to obtain an image processing model.
[0267] In an embodiment of the present application, a method for model training is provided. First, a set of training image samples is obtained, where the set of training image samples belongs to a set of training image data. Then, a random value is generated. Next, according to the random value and a ratio value, a set of training objects is determined from the set of training image samples. Finally, the set of training objects is used to train a processing model for the training images, and an image processing model is obtained. By the above method, an image processing model with a fixed magnification can be trained without additional training for models with different magnifications. The cascade times of the image processing model are used to achieve image super-resolution with any magnification factor, and image super-resolution with any magnification factor is achieved while ensuring performance. Thus, the flexibility of image processing is improved, and the range of image magnification usage is increased.
[0268] Optionally, based on the corresponding embodiment above, Figure 11 in another embodiment of the image processing model training apparatus 60 provided in the embodiment of the present application,
[0269] The determining module 603 is specifically configured to determine whether the random value is greater than the ratio value;
[0270] If the random value is greater than the ratio value, the first image and the third image corresponding to the training image sample are obtained from the set of training image samples;
[0271] The first predicted image corresponding to the first image is obtained through the processing model for the training images;
[0272] The training object in the set of training objects is generated according to the first predicted image and the third image.
[0273] Secondly, in an embodiment of the present application, a method for selecting a set of training objects is provided, that is, if the random value is greater than the ratio value, the first image and the third image corresponding to the training image sample are obtained from the set of training image samples. Then, the first predicted image corresponding to the first image is obtained through the processing model for the training images, and the training object in the set of training objects is generated according to the first predicted image and the third image. By the above method, the situation of continuous adjustment of network model parameters is fully considered, and based on the uncertainty of the random value, different samples will be selected during the training process, thereby improving the reliability of training and the diversity of sample types.
[0274] Optionally, based on the corresponding embodiment above, Figure 11 in another embodiment of the image processing model training apparatus 60 provided in the embodiment of the present application,
[0275] The determining module 603 is specifically configured to determine whether the random value is greater than the ratio value;
[0276] If the random value is less than or equal to the ratio value, obtain the second image and the third image corresponding to the to-be-trained image sample from the to-be-trained image sample set;
[0277] Generate a to-be-trained object in the to-be-trained object set according to the second image and the third image.
[0278] Secondly, in the embodiment of the present application, a method for selecting a to-be-trained object set is provided, that is, if the random value is less than or equal to the ratio value, obtain the second image and the third image corresponding to the to-be-trained image sample from the to-be-trained image sample set, and then generate the to-be-trained object in the to-be-trained object set according to the second image and the third image. By the above method, the situation that the network model parameters are continuously adjusted is fully considered, and based on the uncertainty of the random value, different samples will be selected during the training process, so as to improve the reliability of training and the diversity of sample types.
[0279] Optionally, on the basis of the corresponding embodiment above, in another embodiment of the image processing model training device 60 provided in the embodiment of the present application, Figure 11 The obtaining module 601 is specifically configured to obtain a first to-be-trained image sample, where the first to-be-trained image sample includes the first image, the second image, and the third image;
[0280] Obtain a second to-be-trained image sample, where the second to-be-trained image sample includes a fourth image, the first image, and the second image, and the fourth image and the first image have a preset sampling magnification;
[0281] Obtain a third to-be-trained image sample, where the third to-be-trained image sample includes a fifth image, the fourth image, and the first image, and the fifth image and the fourth image have a preset sampling magnification;
[0282] Generate a to-be-trained image sample according to the first to-be-trained image sample, the second to-be-trained image sample, and the third to-be-trained image sample.
[0283] Generate a to-be-trained image sample according to the first to-be-trained image sample, the second to-be-trained image sample, and the third to-be-trained image sample.
[0284] Secondly, in the embodiments of the present application, the content included in the set of image samples to be trained is provided, that is, the first image sample to be trained, the second image sample to be trained, and the third image sample to be trained are obtained. The first image sample to be trained includes the first image, the second image, and the third image. The second image sample to be trained includes the fourth image, the first image, and the second image. The third image sample to be trained includes the fifth image, the fourth image, and the first image. Finally, according to the first image sample to be trained, the second image sample to be trained, and the third image sample to be trained, the image sample to be trained is generated. In the above manner, more types of samples can be obtained, so that more input distributions can be obtained during the training process, and the diversity of the training set is increased in the form of the union of samples.
[0285] Optionally, based on the corresponding embodiment above, Figure 11 in another embodiment of the image processing model training apparatus 60 provided in the embodiments of the present application,
[0286] the obtaining module 601 is further configured to obtain an offset value and a slope value before the determining module 603 determines the set of objects to be trained from the set of image samples to be trained according to the random value and the ratio value;
[0287] the obtaining module 601 is further configured to obtain the number of iterations corresponding to the set of image samples to be trained;
[0288] the determining module 603 is further configured to determine the ratio value corresponding to the set of image samples to be trained according to the offset value, the slope value, and the number of iterations corresponding to the set of image samples to be trained obtained by the obtaining module 601.
[0289] Secondly, in the embodiments of the present application, a method for determining the ratio value is provided, that is, first an offset value and a slope value are obtained, then the number of iterations corresponding to the set of image samples to be trained is obtained, and finally, according to the offset value, the slope value, and the number of iterations corresponding to the set of image samples to be trained, the ratio value corresponding to the set of image samples to be trained is determined. In the above manner, the ratio value can be updated based on the change of the number of iterations, so as to increase the diversity of the training samples, and further improve the reliability of model training.
[0290] Optionally, based on the corresponding embodiment above, Figure 11 in another embodiment of the image processing model training apparatus 60 provided in the embodiments of the present application,
[0291] the training module 604 is specifically configured to obtain, through the image processing model to be trained, a second predicted image corresponding to each object to be trained in the set of objects to be trained;
[0292] Determine network model parameters using an objective loss function based on the second predicted image corresponding to each object to be trained and the desired image corresponding to each object to be trained;
[0293] Train the image processing model to be trained using the network model parameters to obtain the image processing model;
[0294] The training module 604 specifically determines the network model parameters in the following manner:
[0295]
[0296] Among them, L(θ) represents the objective loss function, θ represents the network model parameters, n represents the total number of training image samples in the training image sample set, and x i represents the i-th object to be trained in the object to be trained set, and net(x i , θ) represents the second predicted image corresponding to the i-th object to be trained, and y i represents the desired image corresponding to the i-th object to be trained.
[0297] Again, in the embodiments of the present application, a method for model training is provided. First, obtain the second predicted image corresponding to each object to be trained in the object to be trained image processing model, then determine the network model parameters using the objective loss function based on the second predicted image corresponding to each object to be trained and the desired image corresponding to each object to be trained, and then train the image processing model to be trained using the network model parameters to obtain the image processing model. Through the above method, on the one hand, it provides a specific implementation basis for model training, thereby improving the reliability of model training. On the other hand, based on the above training method, there is no need to modify the network structure of the image processing model to be trained. Only the feature extraction module is retained and the upsampling module based on meta-learning is added to achieve multiple magnification factors for a single model, thereby improving the compatibility of model training.
[0298] Figure 12 is a schematic structural diagram of the network device 70 in the embodiments of the present application. The network device 70 may include an input device 710, an output device 720, a processor 730, and a memory 740. The output device in the embodiments of the present application may be a display device.
[0299] The memory 740 may include a read-only memory and a random access memory, and provide instructions and data to the processor 730. A part of the memory 740 may also include a non-volatile random access memory (Non-Volatile Random Access Memory, NVRAM).
[0300] The memory 740 stores the following elements, executable modules or data structures, or subsets thereof, or extended sets thereof:
[0301] Operation instructions: including various operation instructions for implementing various operations.
[0302] Operating system: including various system programs for implementing various basic services and processing hardware-based tasks.
[0303] In the embodiments of the present application, the processor 730 is used for:
[0304] Obtain an image to be processed, where the image to be processed corresponds to a first magnification;
[0305] Obtain the zoom factor information corresponding to the image to be processed, where the zoom factor information is used to indicate the magnification factor for magnifying the image to be processed or the reduction factor for reducing the image to be processed;
[0306] Determine the number of cascades according to the zoom factor information, where the number of cascades is an integer greater than or equal to 1, and the number of cascades represents the number of times of processing the image to be processed using the same image processing model;
[0307] According to the number of cascades, obtain the target image corresponding to the image to be processed through the image processing model, where the target image corresponds to a second magnification, the second magnification has an associated relationship with the number of cascades, and the second magnification is different from the first magnification.
[0308] The processor 730 controls the operation of the network device 70. The processor 730 may also be referred to as a Central Processing Unit (CPU). The memory 740 may include a read-only memory and a random access memory, and provide instructions and data to the processor 730. A part of the memory 740 may also include NVRAM. In a specific application, each component of the network device 70 is coupled together through a bus system 750. The bus system 750 may include a power bus, a control bus, a status signal bus, etc. in addition to a data bus. However, for the sake of clarity, various buses are labeled as the bus system 750 in the figure.
[0309] The method disclosed in the embodiments of the present application can be applied to or implemented by the processor 730. The processor 730 may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor 730 or instructions in the form of software. The above-mentioned processor 730 may be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 740, and the processor 730 reads the information in the memory 740 and combines its hardware to complete the steps of the above method.
[0310] Figure 12 For the relevant description, please refer to Figure 3 the relevant description and effects in the method section for understanding, and no further elaboration will be made here.
[0311] The embodiments of the present application also provide another image display device, as Figure 13 shown. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown. For the specific technical details not disclosed, please refer to the method section of the embodiments of the present application. The terminal device may be any terminal device including a mobile phone, a tablet computer, a personal digital assistant (PDA), a point of sales (POS) device, an in-vehicle computer, etc. Taking the terminal device as a mobile phone as an example:
[0312] Figure 13 The block diagram shows a part of the structure of the mobile phone related to the terminal device provided by the embodiments of the present application. Refer to Figure 13, The mobile phone includes components such as a radio frequency (RF) circuit 810, a memory 820, an input unit 830, a display unit 840, sensors 850, an audio circuit 860, a wireless fidelity (WiFi) module 870, a processor 880, and a power supply 890. Those skilled in the art can understand that Figure 13 the mobile phone structure shown in
[0313] does not limit the mobile phone, and it may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. Figure 13 The following specifically introduces each component of the mobile phone:
[0314] The RF circuit 810 can be used for receiving and transmitting signals during information reception or call processes. Specifically, after receiving the downlink information from the base station, it is sent to the processor 880 for processing; in addition, the designed uplink data is sent to the base station. Generally, the RF circuit 810 includes but is not limited to antennas, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF circuit 810 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile Communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0315] The memory 820 can be used to store software programs and modules. The processor 880 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 820. The memory 820 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 820 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0316] The input unit 830 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the mobile phone. Specifically, the input unit 830 may include a touch panel 831 and other input devices 832. The touch panel 831, also known as a touch screen, can collect touch operations of the user on or near it (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel 831), and drive the corresponding connection device according to a pre-set program. Optionally, the touch panel 831 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 880, and can receive and execute commands sent by the processor 880. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 831. In addition to the touch panel 831, the input unit 830 may also include other input devices 832. Specifically, the other input devices 832 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, a joystick, etc.
[0317] The display unit 840 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone. The display unit 840 may include a display panel 841. Optionally, the display panel 841 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 831 can cover the display panel 841. When the touch panel 831 detects a touch operation on or near it, it is transmitted to the processor 880 to determine the type of touch event. Subsequently, the processor 880 provides a corresponding visual output on the display panel 841 according to the type of touch event. Although in Figure 13 the touch panel 831 and the display panel 841 are implemented as two independent components to realize the input and input functions of the mobile phone, in some embodiments, the touch panel 831 and the display panel 841 can be integrated to realize the input and output functions of the mobile phone.
[0318] The mobile phone may further include at least one sensor 850, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 841 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 841 and / or the backlight when the mobile phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity, and can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors that the mobile phone can also be configured with, they will not be elaborated here.
[0319] The audio circuit 860, the speaker 861, and the microphone 862 can provide an audio interface between the user and the mobile phone. The audio circuit 860 can transmit the electrical signal converted from the received audio data to the speaker 861, and the speaker 861 converts it into a sound signal for output; on the other hand, the microphone 862 converts the collected sound signal into an electrical signal, which is received by the audio circuit 860 and then converted into audio data. After the audio data is output to the processor 880 for processing, it is sent to another mobile phone through the RF circuit 810, for example, or the audio data is output to the memory 820 for further processing.
[0320] WiFi belongs to short-range wireless transmission technology. The mobile phone can help users send and receive emails, browse the web, and access streaming media through the WiFi module 870. It provides users with wireless broadband Internet access. Although Figure 13The WiFi module 870 is shown, but it can be understood that it does not belong to the essential components of the mobile phone and can be completely omitted within the scope of not changing the essence of the invention as needed.
[0321] The processor 880 is the control center of the mobile phone, connecting various parts of the entire mobile phone through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 820, and by calling data stored in the memory 820, it executes various functions of the mobile phone and processes data. Optionally, the processor 880 may include one or more processing units; optionally, the processor 880 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 880 either.
[0322] The mobile phone also includes a power supply 890 (such as a battery) for supplying power to each component. Optionally, the power supply can be logically connected to the processor 880 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system.
[0323] Although not shown, the mobile phone may also include a camera, a Bluetooth module, etc., which will not be elaborated here.
[0324] In the embodiment of the present application, the processor 880 included in the terminal device further has the following functions:
[0325] Obtain an image to be processed, where the image to be processed corresponds to a first magnification;
[0326] Receive an image adjustment instruction, where the image adjustment instruction carries an image magnification parameter, and the image magnification parameter is used to indicate the magnification multiple for magnifying the image to be processed;
[0327] In response to the image adjustment instruction, determine the number of cascades according to the image magnification parameter, where the number of cascades is an integer greater than or equal to 1, and the number of cascades represents the number of times of processing the image to be processed using the same image processing model;
[0328] According to the number of cascades, obtain the target image corresponding to the image to be processed through the image processing model, where the target image corresponds to a second magnification, the second magnification has an associated relationship with the number of cascades, and the second magnification is greater than the first magnification;
[0329] Display the target image.
[0330] Figure 14FIG. 0 is a schematic structural diagram of a server provided by an embodiment of the present application. The server 900 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 922 (for example, one or more processors) and a memory 932, and one or more storage media 930 (for example, one or more mass storage devices) for storing application programs 942 or data 944. Among them, the memory 932 and the storage media 930 may be transient storage or persistent storage. The programs stored in the storage media 930 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 922 may be configured to communicate with the storage media 930 and execute a series of instruction operations in the storage media 930 on the server 900.
[0331] The server 900 may further include one or more power supplies 926, one or more wired or wireless network interfaces 950, one or more input / output interfaces 958, and / or one or more operating systems 941, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0332] In the above embodiment, the steps executed by the server may be based on the Figure 14 server structure shown.
[0333] In the embodiment of the present application, the CPU 922 included in the server further has the following functions:
[0334] Obtain a set of to-be-trained image samples, where the set of to-be-trained image samples belongs to a set of to-be-trained image data, the set of to-be-trained image samples includes at least one to-be-trained image sample, each to-be-trained image sample includes a first image, a second image, and a third image, the first image and the second image have a preset sampling magnification, and the second image and the third image have the preset sampling magnification;
[0335] Generate a random value, where the random value is greater than or equal to 0 and less than or equal to 1;
[0336] Determine a set of to-be-trained objects from the set of to-be-trained image samples according to the random value and a ratio value, where the set of to-be-trained objects includes at least one to-be-trained object, each to-be-trained object includes the second image and the third image, or each to-be-trained object includes a first predicted image and the third image, and the first predicted image is obtained by passing the first image through a to-be-trained image processing model;
[0337] Train the image processing model to be trained using the set of objects to be trained, to obtain an image processing model.
[0338] Optionally, the CPU 922 is specifically configured to perform the following steps:
[0339] Determine whether the random value is greater than the ratio value;
[0340] If the random value is greater than the ratio value, obtain the first image and the third image corresponding to the image sample to be trained from the set of image samples to be trained;
[0341] Obtain the first predicted image corresponding to the first image through the image processing model to be trained;
[0342] Generate an object to be trained in the set of objects to be trained according to the first predicted image and the third image.
[0343] Optionally, the CPU 922 is specifically configured to perform the following steps:
[0344] Determine whether the random value is greater than the ratio value;
[0345] If the random value is less than or equal to the ratio value, obtain the second image and the third image corresponding to the image sample to be trained from the set of image samples to be trained;
[0346] Generate an object to be trained in the set of objects to be trained according to the second image and the third image.
[0347] Optionally, the CPU 922 is specifically configured to perform the following steps:
[0348] Obtain a first image sample to be trained, where the first image sample to be trained includes the first image, the second image, and the third image;
[0349] Obtain a second image sample to be trained, where the second image sample to be trained includes a fourth image, the first image, and the second image, and the fourth image and the first image have a preset sampling magnification;
[0350] Obtain a third image sample to be trained, where the third image sample to be trained includes a fifth image, the fourth image, and the first image, and the fifth image and the fourth image have a preset sampling magnification;
[0351] Generate an image sample to be trained according to the first image sample to be trained, the second image sample to be trained, and the third image sample to be trained.
[0352] Optionally, the CPU 922 is further configured to perform the following steps:
[0353] Obtain an offset value and a slope value;
[0354] Obtain the number of iterations corresponding to the set of training image samples to be trained;
[0355] Determine the ratio value corresponding to the set of training image samples to be trained according to the offset value, the slope value, and the number of iterations corresponding to the set of training image samples to be trained.
[0356] Optionally, the CPU 922 is specifically configured to perform the following steps:
[0357] Obtain, through the image processing model to be trained, a second predicted image corresponding to each object to be trained in the set of objects to be trained;
[0358] Determine network model parameters by using a target loss function according to the second predicted image corresponding to each object to be trained and the desired image corresponding to each object to be trained;
[0359] Train the image processing model to be trained by using the network model parameters to obtain the image processing model;
[0360] Determine the network model parameters in the following manner:
[0361]
[0362] Wherein, the L(θ) represents the target loss function, the θ represents the network model parameters, the n represents the total number of training image samples in the set of training image samples to be trained, and the x i represents the i-th object to be trained in the set of objects to be trained, the net(x i ,θ) represents the second predicted image corresponding to the i-th object to be trained, and the y i represents the desired image corresponding to the i-th object to be trained.
[0363] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0364] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0365] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0366] In addition, in each embodiment of this application, the functional units can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0367] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0368] As described above, the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of this application.
Claims
1. An image processing method, characterized in that, comprising: obtaining an image to be processed, wherein the image to be processed corresponds to a first magnification; obtaining magnification factor information corresponding to the image to be processed, wherein the magnification factor information is used to indicate a magnification multiple for magnifying the image to be processed or a reduction multiple for reducing the image to be processed; determining a cascade number according to the magnification factor information, wherein the cascade number is an integer greater than or equal to 1, and the cascade number represents the number of times of processing the image to be processed using the same image processing model, and the image processing model is used to achieve image super-resolution with any magnification; obtaining a target image corresponding to the image to be processed through the image processing model according to the cascade number, wherein the target image corresponds to a second magnification, the second magnification has an associated relationship with the cascade number, and the second magnification is different from the first magnification; The training process of the image processing model includes: obtaining a set of training image samples, wherein the set of training image samples belongs to a set of training image data, the set of training image samples includes at least one training image sample, each training image sample includes a first image, a second image and a third image, the first image and the second image have a preset sampling magnification, and the second image and the third image have the preset sampling magnification; generating a random value, wherein the random value is greater than or equal to 0 and less than or equal to 1; determining a set of training objects from the set of training image samples according to the random value and a ratio value, wherein the set of training objects includes at least one training object, each training object includes the second image and the third image, or each training object includes a first predicted image and the third image, the first predicted image is obtained by passing the first image through the training image processing model, and the ratio value is determined according to the iteration number corresponding to the set of training image samples, and the ratio value is a monotonically non-increasing function of the iteration number; training the training image processing model using the set of training objects to obtain an image processing model.
2. An image display method, characterized in that, comprising: obtaining an image to be processed, wherein the image to be processed corresponds to a first magnification; receiving an image adjustment instruction, wherein the image adjustment instruction carries an image magnification parameter, and the image magnification parameter is used to indicate a magnification multiple for magnifying the image to be processed; responding to the image adjustment instruction, determining a cascade number according to the image magnification parameter, wherein the cascade number is an integer greater than or equal to 1, and the cascade number represents the number of times of processing the image to be processed using the same image processing model, and the image processing model is used to achieve image super-resolution with any magnification; Obtain the target image corresponding to the image to be processed through the image processing model according to the number of cascades, where the target image corresponds to a second magnification factor, the second magnification factor has an associated relationship with the number of cascades, and the second magnification factor is greater than the first magnification factor; Display the target image; The training process of the image processing model includes: Obtain a set of training image samples, where the set of training image samples belongs to a set of training image data, the set of training image samples includes at least one training image sample, each training image sample includes a first image, a second image, and a third image, the first image and the second image have a preset sampling magnification factor, and the second image and the third image have the preset sampling magnification factor; Generate a random value, where the random value is greater than or equal to 0 and less than or equal to 1; Determine a set of training objects from the set of training image samples according to the random value and the ratio value, where the set of training objects includes at least one training object, each training object includes the second image and the third image, or each training object includes a first predicted image and the third image, the first predicted image is obtained after the first image passes through the image processing model to be trained, the ratio value is determined according to the number of iterations corresponding to the set of training image samples, and the ratio value is a non-increasing function of the number of iterations; Train the image processing model to be trained using the set of training objects to obtain an image processing model.
3. The method according to claim 2, wherein, the determining a set of training objects from the set of training image samples according to the random value and the ratio value includes: Determine whether the random value is greater than the ratio value; If the random value is greater than the ratio value, obtain the first image and the third image corresponding to the training image sample from the set of training image samples; Obtain the first predicted image corresponding to the first image through the image processing model to be trained; Generate a training object in the set of training objects according to the first predicted image and the third image.
4. The method according to claim 2, wherein, the determining a set of training objects from the set of training image samples according to the random value and the ratio value includes: Determine whether the random value is greater than the ratio value; If the random value is less than or equal to the ratio value, obtain the second image and the third image corresponding to the training image sample from the set of training image samples; Generate a training object in the set of training objects according to the second image and the third image.
5. The method according to claim 2, wherein, the obtaining a set of training image samples includes: Obtain a first training image sample, where the first training image sample includes the first image, the second image, and the third image; Obtain a second image sample to be trained, where the second image sample to be trained includes a fourth image, the first image, and the second image, and the fourth image and the first image have a preset sampling magnification; Obtain a third image sample to be trained, where the third image sample to be trained includes a fifth image, the fourth image, and the first image, and the fifth image and the fourth image have a preset sampling magnification; Generate an image sample to be trained according to the first image sample to be trained, the second image sample to be trained, and the third image sample to be trained.
6. The method according to claim 2, wherein, before determining the set of objects to be trained from the set of image samples to be trained according to the random value and the ratio value, the method further includes: Obtain an offset value and a slope value; Obtain the number of iterations corresponding to the set of image samples to be trained; Determine the ratio value corresponding to the set of image samples to be trained according to the offset value, the slope value, and the number of iterations corresponding to the set of image samples to be trained.
7. The method according to any one of claims 2 to 6, wherein, training the image processing model to be trained with the set of objects to be trained to obtain an image processing model includes: Obtain a second predicted image corresponding to each object to be trained in the set of objects to be trained through the image processing model to be trained; Determine network model parameters by using a target loss function according to the second predicted image corresponding to each object to be trained and the expected image corresponding to each object to be trained; Train the image processing model to be trained with the network model parameters to obtain the image processing model; Determining the network model parameters by using a target loss function according to the second predicted image corresponding to each object to be trained and the expected image corresponding to each object to be trained includes: Determine the network model parameters in the following manner: ; Among them, the represents the target loss function, the represents the network model parameters, the represents the total number of training image samples in the set of training image samples to be trained, the represents the th training object in the set of training objects to be trained, the represents the th second predicted image corresponding to the training object to be trained, the represents the th expected image corresponding to the training object to be trained.
8. An image processing device, wherein, comprises: an acquisition module, configured to acquire an image to be processed, where the image to be processed corresponds to a first magnification; the acquisition module is further configured to acquire scaling factor information corresponding to the image to be processed, where the scaling factor information is used to indicate a magnification factor for magnifying the image to be processed or a reduction factor for reducing the image to be processed; a determination module, configured to determine the number of cascades according to the scaling factor information acquired by the acquisition module, where the number of cascades is an integer greater than or equal to 1, and the number of cascades represents the number of times of processing the image to be processed by using the same image processing model; the acquisition module is further configured to obtain a target image corresponding to the image to be processed through the image processing model according to the number of cascades determined by the determination module, where the target image corresponds to a second magnification, the second magnification has an association relationship with the number of cascades, and the second magnification is different from the first magnification; The image processing apparatus further includes: an image processing model training apparatus, where the image processing model training apparatus includes: An acquisition module, configured to acquire a set of to-be-trained image samples. Among them, the set of to-be-trained image samples belongs to a set of to-be-trained image data. The set of to-be-trained image samples includes at least one to-be-trained image sample. Each to-be-trained image sample includes a first image, a second image, and a third image. The first image and the second image have a preset sampling magnification, and the second image and the third image have the preset sampling magnification; A generation module, configured to generate a random value, where the random value is greater than or equal to 0 and less than or equal to 1; A determination module, configured to determine a set of to-be-trained objects from the set of to-be-trained image samples according to the random value generated by the generation module and a ratio value. Among them, the set of to-be-trained objects includes at least one to-be-trained object. Each to-be-trained object includes the second image and the third image, or each to-be-trained object includes a first predicted image and the third image. The first predicted image is obtained by passing the first image through a to-be-trained image processing model; A training module, configured to train the to-be-trained image processing model by using the set of to-be-trained objects determined by the determination module to obtain an image processing model.
9. An image display apparatus, characterized in that it includes: An acquisition module, configured to acquire an image to be processed, where the image to be processed corresponds to a first magnification; A reception module, configured to receive an image adjustment instruction, where the image adjustment instruction carries an image magnification parameter, and the image magnification parameter is used to indicate the magnification multiple for magnifying the image to be processed; A determination module, configured to, in response to the image adjustment instruction received by the reception module, determine a cascade number according to the image magnification parameter. The cascade number is an integer greater than or equal to 1, and the cascade number represents the number of times of processing the image to be processed by using the same image processing model; The acquisition module is further configured to obtain a target image corresponding to the image to be processed through the image processing model according to the cascade number determined by the determination module. The target image corresponds to a second magnification. The second magnification has an associated relationship with the cascade number, and the second magnification is greater than the first magnification; A display module, configured to display the target image acquired by the acquisition module; The image display apparatus further includes an image processing model training apparatus, where the image processing model training apparatus includes: An acquisition module, configured to acquire a set of to-be-trained image samples. Among them, the set of to-be-trained image samples belongs to a set of to-be-trained image data. The set of to-be-trained image samples includes at least one to-be-trained image sample. Each to-be-trained image sample includes a first image, a second image, and a third image. The first image and the second image have a preset sampling magnification, and the second image and the third image have the preset sampling magnification; A generation module for generating a random value, where the random value is greater than or equal to 0 and less than or equal to 1; A determination module for determining a set of objects to be trained from the set of image samples to be trained according to the random value and the ratio value generated by the generation module, where the set of objects to be trained includes at least one object to be trained, and each object to be trained includes the second image and the third image, or each object to be trained includes a first predicted image and the third image, and the first predicted image is obtained after the first image passes through the image processing model to be trained; A training module for training the image processing model to be trained with the set of objects to be trained determined by the determination module to obtain an image processing model.
10. The apparatus according to claim 9, wherein, the determination module is specifically configured to determine whether the random value is greater than the ratio value; if the random value is greater than the ratio value, obtain the first image and the third image corresponding to the image sample to be trained from the set of image samples to be trained; obtain the first predicted image corresponding to the first image through the image processing model to be trained; generate an object to be trained in the set of objects to be trained according to the first predicted image and the third image.
11. The apparatus according to claim 9, wherein, the determination module is specifically configured to determine whether the random value is greater than the ratio value; if the random value is less than or equal to the ratio value, obtain the second image and the third image corresponding to the image sample to be trained from the set of image samples to be trained; generate an object to be trained in the set of objects to be trained according to the second image and the third image.
12. The apparatus according to claim 9, wherein, the acquisition module is specifically configured to acquire a first image sample to be trained, where the first image sample to be trained includes the first image, the second image, and the third image; acquire a second image sample to be trained, where the second image sample to be trained includes a fourth image, the first image, and the second image, and the fourth image has a preset sampling magnification with respect to the first image; acquire a third image sample to be trained, where the third image sample to be trained includes a fifth image, the fourth image, and the first image, and the fifth image has a preset sampling magnification with respect to the fourth image; generate an image sample to be trained according to the first image sample to be trained, the second image sample to be trained, and the third image sample to be trained.
13. The apparatus according to claim 9, wherein, the acquisition module is further configured to acquire an offset value and a slope value before the determination module determines a set of objects to be trained from the set of image samples to be trained according to the random value and the ratio value; the acquisition module is further configured to acquire the number of iterations corresponding to the set of image samples to be trained; The determining module is further configured to determine the ratio value corresponding to the set of to-be-trained image samples according to the offset value, the slope value, and the iteration number corresponding to the set of to-be-trained image samples obtained by the obtaining module.
14. The apparatus according to any one of claims 9 to 13, wherein: The training module is specifically configured to obtain, by using the to-be-trained image processing model, a second predicted image corresponding to each to-be-trained object in the set of to-be-trained objects; Determine network model parameters by using a target loss function according to the second predicted image corresponding to each to-be-trained object and the expected image corresponding to each to-be-trained object; Train the to-be-trained image processing model by using the network model parameters to obtain the image processing model; The training module specifically determines the network model parameters in the following manner: ; Among them, the represents the target loss function, the represents the network model parameters, the represents the total number of training image samples in the set of training image samples to be trained, the represents the th training object in the set of training objects to be trained, the represents the th second predicted image corresponding to the training object to be trained, the represents the th expected image corresponding to the training object to be trained.
15. A network device, wherein: It includes: A memory, a transceiver, a processor, and a bus system; Wherein, the memory is used for storing programs; The processor is configured to execute the programs in the memory, including the following steps: Obtain a to-be-processed image, wherein the to-be-processed image corresponds to a first magnification; Obtain the scaling factor information corresponding to the to-be-processed image, wherein the scaling factor information is used to indicate the magnification factor for magnifying the to-be-processed image or the reduction factor for reducing the to-be-processed image; Determine the number of cascading times according to the scaling factor information, wherein the number of cascading times is an integer greater than or equal to 1, and the number of cascading times represents the number of times of processing the to-be-processed image by using the same image processing model, and the image processing model is used to implement image super-resolution with any magnification factor; According to the number of cascading times, obtain the target image corresponding to the to-be-processed image by using the image processing model, wherein the target image corresponds to a second magnification, the second magnification has an associated relationship with the number of cascading times, and the second magnification is different from the first magnification; The training process of the image processing model includes: Obtain a set of to-be-trained image samples, wherein the set of to-be-trained image samples belongs to a set of to-be-trained image data, the set of to-be-trained image samples includes at least one to-be-trained image sample, each to-be-trained image sample includes a first image, a second image, and a third image, the first image and the second image have a preset sampling magnification, and the second image and the third image have the preset sampling magnification; Generate a random value, wherein the random value is greater than or equal to 0 and less than or equal to 1; Determine a set of objects to be trained from the set of image samples to be trained according to the random value and the ratio value, where the set of objects to be trained includes at least one object to be trained, and each object to be trained includes the second image and the third image, or each object to be trained includes the first predicted image and the third image, the first predicted image is obtained after the first image passes through the image processing model to be trained, the ratio value is determined according to the number of iterations corresponding to the set of image samples to be trained, and the ratio value is a monotonically non-increasing function of the number of iterations; Train the image processing model to be trained using the set of objects to be trained to obtain an image processing model; The bus system is used to connect the memory and the processor to enable communication between the memory and the processor.
16. A terminal device, characterized in that, comprising: a memory, a transceiver, a processor, and a bus system; wherein the memory is used to store programs; the processor is used to execute the programs in the memory, including the following steps: Obtain an image to be processed, where the image to be processed corresponds to a first magnification; Receive an image adjustment instruction, where the image adjustment instruction carries an image magnification parameter, and the image magnification parameter is used to indicate the magnification multiple for magnifying the image to be processed; In response to the image adjustment instruction, determine the number of cascades according to the image magnification parameter, where the number of cascades is an integer greater than or equal to 1, and the number of cascades represents the number of times of processing the image to be processed using the same image processing model, and the image processing model is used to implement image super-resolution with any magnification multiple; According to the number of cascades, obtain the target image corresponding to the image to be processed through the image processing model, where the target image corresponds to a second magnification, the second magnification has an association relationship with the number of cascades, and the second magnification is greater than the first magnification; Display the target image; The training process of the image processing model includes: Obtain a set of image samples to be trained, where the set of image samples to be trained belongs to a set of image data to be trained, the set of image samples to be trained includes at least one image sample to be trained, each image sample to be trained includes a first image, a second image, and a third image, the first image and the second image have a preset sampling magnification, and the second image and the third image have the preset sampling magnification; Generate a random value, where the random value is greater than or equal to 0 and less than or equal to 1; Determine a set of objects to be trained from the set of image samples to be trained according to the random numerical value and the ratio numerical value, wherein the set of objects to be trained includes at least one object to be trained, and each object to be trained includes the second image and the third image, or each object to be trained includes a first predicted image and the third image, the first predicted image is obtained by passing the first image through the image processing model to be trained, the ratio numerical value is determined according to the number of iterations corresponding to the set of image samples to be trained, and the ratio numerical value is a monotonically non-increasing function of the number of iterations; Train the image processing model to be trained by using the set of objects to be trained to obtain an image processing model; The bus system is used to connect the memory and the processor so that the memory and the processor can communicate with each other.
17. A computer-readable storage medium includes instructions which, when running on a computer, cause the computer to execute the method as claimed in claim 1 or execute the method as claimed in any one of claims 2 to 7.
Citation Information
Patent Citations
Method and device for image resizing
CN102246223A
Image processing method and device, electronic device and computer readable medium
CN109934773A