Method and device for model training, equipment and storage medium
By using feature extraction units and gating weight parameters in the machine learning model, the problems of data compression and packet loss in image transmission are solved, achieving efficient image super-resolution processing and improving image quality and processing efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2024-11-14
- Publication Date
- 2026-05-15
AI Technical Summary
During image transmission, due to limitations in network bandwidth and storage capacity, image data is difficult to meet users' demands for high-definition display. Existing technologies suffer from data compression and packet loss issues, resulting in low image processing efficiency.
By using the feature extraction unit in the machine learning model, multiple feature extraction operators are used to extract multiple types of features from the image. Combined with the gating weight parameters, a predictive super-resolution image is generated. The model parameters are then updated based on the difference between the predictive super-resolution image and the super-resolution version image, thereby improving the model's ability to understand complex features.
It improves the efficiency and quality of image super-resolution processing, reduces computing costs and storage requirements, and enhances the overall performance of image processing.
Smart Images

Figure CN122047351A_ABST
Abstract
Description
Technical Field
[0001] The exemplary embodiments disclosed herein relate generally to the field of computers, and more particularly to methods, apparatus, devices, and computer-readable storage media for model training. Background Technology
[0002] With the rapid development of computer technology, more and more high-definition display devices are becoming widespread, leading to increasingly higher demands from users for video clarity. However, in practical applications, image processing is affected by many factors. For example, during image transmission, limitations in network bandwidth and storage capacity can result in data compression and packet loss, making it difficult for image data to meet user needs. Therefore, handling the massive and rich amounts of data arriving at users' client devices has become a challenging problem. Summary of the Invention
[0003] In a first aspect of this disclosure, a method for model training is provided. The method includes: acquiring a first sample image and a second sample image for training a machine learning model, the second sample image being a super-resolution version of the first sample image; the machine learning model being configured to perform super-resolution processing on the input image to obtain an output image, the output image having a higher resolution than the input image; extracting feature information from the first sample image using a feature extraction unit in the machine learning model, wherein the feature extraction includes: extracting multiple types of features from the first sample image respectively through multiple feature extraction operators in multiple processing branches of the feature extraction unit, the multiple processing branches having corresponding gate weight parameters, each gate weight parameter indicating the importance of the corresponding type of feature in the feature information; generating a predicted super-resolution image for the first sample image based on the feature information; and updating the model parameter set of the machine learning model at least based on the difference between the predicted super-resolution image and the second sample image to obtain a trained machine learning model, the model parameter set including at least the gate weight parameters.
[0004] In a second aspect of this disclosure, an apparatus for model training is provided. The apparatus includes: a sample image acquisition module configured to acquire a first sample image and a second sample image for training a machine learning model, the second sample image being a super-resolution version of the first sample image, the machine learning model being configured to perform super-resolution processing on the input image to obtain an output image, the output image having a higher resolution than the input image; a feature information extraction module configured to extract feature information from the first sample image using a feature extraction unit in the machine learning model, wherein the feature information extraction includes: extracting multiple types of features from the first sample image respectively through multiple feature extraction operators in multiple processing branches of the feature extraction unit, the multiple processing branches having corresponding gate weight parameters, each gate weight parameter indicating the importance of the corresponding type of feature in the feature information; a prediction super-resolution image generation module configured to generate a prediction super-resolution image for the first sample image based on the feature information; and a parameter set update module configured to update the model parameter set of the machine learning model at least based on the difference between the prediction super-resolution image and the second sample image to obtain a trained machine learning model, the model parameter set including at least gate weight parameters.
[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.
[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.
[0007] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.
[0008] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0010] Figure 1 A schematic diagram of an example environment in which embodiments of the present disclosure can be implemented is shown;
[0011] Figure 2 A schematic diagram of an example model structure for training a machine learning model according to some embodiments of the present disclosure is shown;
[0012] Figure 3A A schematic diagram of an example architecture for a feature extraction unit to extract feature information from a sample image during the model training phase, according to some embodiments of the present disclosure, is shown.
[0013] Figure 3B A schematic structural diagram of a merged feature extraction unit in the model application stage according to some embodiments of the present disclosure is shown.
[0014] Figure 4 A schematic diagram of an example architecture for model training according to some embodiments of the present disclosure is shown;
[0015] Figure 5 A schematic diagram of an example architecture for acquiring sample images according to some embodiments of the present disclosure is shown;
[0016] Figure 6 A schematic diagram of an example architecture for acquiring sample images according to other embodiments of the present disclosure is shown;
[0017] Figure 7 A flowchart illustrating a model training process according to some embodiments of the present disclosure is shown;
[0018] Figure 8 A schematic structural block diagram of an apparatus for model training according to some embodiments of the present disclosure is shown; and
[0019] Figure 9 A block diagram of an electronic device in which one or more embodiments of the present disclosure may be implemented is shown. Detailed Implementation
[0020] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0021] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0022] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0023] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0024] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0025] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0026] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.
[0027] In this document, unless explicitly stated otherwise, performing a step in response to A does not mean that the step is performed immediately after A, but may include one or more intermediate steps.
[0028] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0029] As used in this paper, the term "model" refers to a system that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that uses multiple layers of processing units to process inputs and provide corresponding outputs. In this paper, "model" may also be referred to as a "machine learning model," a "machine learning network," or simply a "network," and these terms are used interchangeably. A model can also include different types of processing units or networks.
[0030] As used herein, a “unit,” “operation unit,” or “subunit” can consist of any suitable machine learning model or network. As used herein, a set of elements or similar expressions can include one or more such elements. For example, “a set of convolutional units” can include one or more convolutional units.
[0031] Image super-resolution technology has become extremely widespread in various applications. Efficient image super-resolution methods can significantly reduce the bandwidth costs required for image transmission and improve overall processing efficiency. Especially with the rapid development of generated content, diffusion models have emerged as a new paradigm in image generation, bringing high-quality image generation capabilities. However, the high quality of diffusion models is directly proportional to the enormous computing power required, increasing the overall computational cost.
[0032] Furthermore, due to the uneven distribution of computational pressure during inference in super-resolution models, a sharp increase in GPU memory (or system memory) occurs during the feature fusion upsampling stage. Conventionally, super-resolution can be implemented in stages using optimization techniques. That is, a complete image is divided into several small blocks, and each block is super-resolutioned sequentially before being stitched together to form the final image. However, this approach degrades latency. Correspondingly, methods based on diffusion models and utilizing control networks require significant computational resources.
[0033] In view of this, embodiments of the present disclosure propose a scheme for model training. According to various embodiments of the present disclosure, a first sample image and a second sample image are obtained for training a machine learning model. The second sample image is a super-resolution version of the first sample image. The machine learning model is configured to perform super-resolution processing on the input image to obtain an output image with a resolution higher than that of the input image. Further, feature information of the first sample image is extracted using a feature extraction unit in the machine learning model. Feature information is extracted in the following manner: in multiple processing branches of the feature extraction unit, multiple types of features of the first sample image are extracted respectively through multiple feature extraction operators. The multiple processing branches have corresponding gate weight parameters, and each gate weight parameter indicates the importance of the corresponding type of feature in the feature information. Subsequently, a predicted super-resolution image for the first sample image is generated based on the feature information. Accordingly, the model parameter set of the machine learning model is updated at least based on the difference between the predicted super-resolution image and the second sample image to obtain a trained machine learning model, the model parameter set including at least the gate weight parameters.
[0034] Therefore, by using the feature extraction unit in the machine learning model to train the machine learning model based on the acquired first and second sample images, the machine learning model's ability to understand complex features can be improved.
[0035] Example Environment
[0036] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. Figure 1 In environment 100, it is desirable to train and use a machine learning model 130 configured for various application environments. For example, in the case where machine learning model 130 is a super-resolution model, it can output a high-resolution image of the image based on the user-input image, and so on.
[0037] like Figure 1 As shown, environment 100 includes model training system 150. Figure 1The upper part illustrates the model training phase, and the lower part illustrates the model application phase. Before training, the parameter values of machine learning model 130 can have initial values or pre-trained parameter values obtained through a pre-training process. Machine learning model 130 can be trained via a reparameterized structural feature extraction unit, during which the parameter values of machine learning model 130 can be updated and adjusted. After training, machine learning model 130 can be distilled and pruned to obtain machine learning model 130′. At this point, the parameter values of machine learning model 130′ have been updated, and based on the updated parameter values, machine learning model 130 can be used in the model application phase to achieve the task of generating high-resolution images.
[0038] During the model training phase, the machine learning model 130 can be trained using a model training system 150 based on a sample pair set 110 comprising multiple sample pairs 112 used to train the machine learning model 130. Here, each sample pair 112 may include a low-resolution image (LR sample image) and a high-resolution image (HR sample image). For example, for a task referring to generating super-resolution images, the sample pair 112 may include an LR sample image 120 for model input and an HR sample image 122 for model output in the task referring to generating super-resolution images. The LR sample image 120 in the task referring to generating super-resolution images may, for example, include a low-resolution image, and the HR sample image 122 may, for example, include a high-resolution image corresponding to the low-resolution image. The sample pair 112, including the LR sample image 120 and the HR sample image 122, can be used to train the machine learning model 130. Specifically, the training process can be performed iteratively using a large number of training samples. After training is complete, the machine learning model 130 may include knowledge about the task to be processed. During the model application phase, the machine learning model 130′ (which at this point has the trained parameter values) can be used to perform the corresponding task. For example, it can receive the LR image 142, which refers to the task of generating a super-resolution image, and output the corresponding HR image 144.
[0039] exist Figure 1 In this context, the model training system 150 may include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, servers, etc. Terminal devices may involve any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. Servers include, but are not limited to, mainframes, edge computing nodes, computing devices in cloud environments, etc.
[0040] It should be understood that Figure 1 The components and arrangements shown in environment 100 are merely examples, and a computing system suitable for implementing the exemplary implementations described in this disclosure may include one or more different components, other components, and / or different arrangements. Implementations of this disclosure are not limited in this respect.
[0041] The following description will continue with reference to the accompanying drawings, which will be used to describe some exemplary embodiments of this disclosure. Hereinafter, the exemplary embodiments will be described primarily with respect to the model training system 150. It should be understood that the actions described with respect to the model training system 150 may be performed by the machine learning model 130 on the model training system 150, or may be performed by the machine learning model 130 in conjunction with its server (e.g., a server).
[0042] The following will refer to Figure 2 and Figure 3A This disclosure describes the scheme used for model training. Figure 2 A flowchart of an example model structure process 200 for a machine learning model 130 according to some embodiments of the present disclosure is shown. Figure 3A A schematic diagram of an example architecture for extracting feature information from sample images using a feature extraction unit during the model training phase, according to some embodiments of the present disclosure, is shown. In embodiments of the present disclosure, it is desirable to train a machine learning model 130 configured to perform super-resolution processing on an input image to obtain an output image with a higher resolution than the input image. That is, the input image may be a low-resolution image (i.e., LR image 202), while the output image may be a high-resolution image (i.e., HR image 204).
[0043] like Figure 2 The example model structure 200 shown has an LR image 202 input to a machine learning model 130. The machine learning model 130 generates an HR image (also known as a prediction super-resolution image) 204 for the LR image through operations such as feature extraction and super-resolution processing. In embodiments of this disclosure, the machine learning model 130 may be constructed based on one or more feature extraction units.
[0044] exist Figure 2In this example, for illustrative purposes only, the machine learning model 130 is constructed to include a feature extraction unit 220-1, a structured information extraction module 212, and a multi-scale edge information feature fusion module 214. In some embodiments, the structured information extraction module 212 includes a multi-angle information extraction residual unit 213 and a feature extraction unit 220-4. The multi-scale edge information feature fusion module 214 includes feature extraction units 220-5 and 220-6. These feature extraction units are collectively referred to as or individually as feature extraction unit 220. The processing units in the machine learning model 130 are interconnected to process the input LR image 202, where the output of one unit is input to the connected subsequent unit until the input is received. In some examples, edge information is a basic and intuitive structured feature in the image, corresponding to the contours of objects and important details. By correctly extracting edge information from the image, richer feature inputs can be provided to the machine learning model, thereby improving its reconstruction capabilities.
[0045] In some embodiments, the multi-angle information extraction residual unit 213 can be used to extract multi-angle feature information of the LR image more effectively and intuitively. The multi-angle information extraction residual unit 213 takes the intermediate feature information extracted by the previous feature extraction unit 220-1 as input. Subsequently, the multi-angle information extraction residual unit 213 obtains multi-angle feature information based on feature extraction units 220-2, 220-3, and the spatial information-oriented local feature extraction unit 221. The output of the multi-angle information extraction residual unit 213 includes the processing result of the spatial information-oriented local feature extraction unit 221 and the original input of the multi-angle information extraction residual unit 213 (multiplied by a single matrix and used as part of the output). The input of the multi-angle information extraction residual unit 213 can also be provided (after multiplication by a single matrix) as part of the input of the feature extraction unit 220-6, and can also be used to generate the HR image 204 after interpolation.
[0046] It should be understood that Figure 2 Only one example structure of the machine learning model 130 is given, in which the number of feature extraction units and the combination of the modules shown can be varied. In other embodiments, various model structures can be derived based on the feature extraction unit 220, and more other types of processing units can be introduced into the model.
[0047] To train a machine learning model 130 with super-resolution processing capabilities, a model training system 150 acquires a first sample image and a second sample image for training the machine learning model 130. The second sample image is a high-resolution or super-resolution version of the first sample image. In some embodiments, the first and second sample images may be constructed based on a degradation model and a diffusion model. In some examples, the first sample image may refer to a low-resolution image (LR sample image), and the second sample image may be referred to as a high-resolution image (HR sample image); the first and second sample images may be referred to as a sample pair. References will follow below. Figure 5 and Figure 6 The model training system 150 acquires first and second sample images for training the machine learning model 130. Understandably, during training, the model training system 150 can call the machine learning model 130 multiple times to generate predictive super-resolution images for the LR sample images. During supervised learning, the generated predictive super-resolution images are compared with the corresponding HR sample images, and the differences between them are used to guide the update of the model parameters of the machine learning model 130.
[0048] Since the feature extraction unit 220 is crucial to the super-resolution processing capability of the machine learning model 130, proper training of its parameters will contribute to higher performance in super-resolution processing. In some embodiments, the model training system 150 can utilize reparameterization training methods to train the machine learning model 130, particularly the feature extraction unit 220. In some embodiments, reparameterization is an efficient training technique that uses mergeable heterogeneous or homogeneous branches during training to achieve better results. Lossless merging of these structures during the inference phase can significantly improve the overall inference efficiency of the model. By creatively unifying the latent space and Gaussian distribution logically using the idea of reparameterization, the problem of fixed spatial representation and insufficient expressive power in traditional models is greatly addressed.
[0049] For ease of understanding, the following text will refer to Figure 3A This describes how the model training system 150 acquires feature information of the first sample image based on the feature extraction units. Understandably, the model training system 150 utilizes the feature extraction units 220-1, 220-2, 220-3, 220-4, ..., 220-N included in the machine learning model 130 to extract feature information of the first sample image. For ease of discussion, the feature extraction units 220-1, 220-2, 220-3, 220-4, ..., 220-N can be collectively referred to as or individually as feature extraction unit 220.
[0050] In some embodiments, when training the feature extraction model 220 based on reparameterization, the feature extraction model 220 can be constructed during the training phase to include multiple processing branches. Each processing branch can utilize a corresponding feature extraction operator to extract a class of image features, particularly structured feature information such as image edges. The multiple processing branches have corresponding gating weight parameters, each indicating the importance of the corresponding type of feature in the feature information. It is understood that the model training system 150 acquires input 310, and for each feature extraction unit 220, the acquired input 310 can be an LR sample image. Alternatively, the input 310 acquired by the feature extraction unit 220 can also be intermediate feature information extracted from the LR sample image by the previous feature extraction unit.
[0051] In some embodiments, edge detection and structured feature extraction play a crucial role in image super-resolution tasks. Super-resolution tasks aim to enhance the content details of an image by magnifying it, enabling machine learning models that output super-resolution data to reconstruct finer textures and details in high-resolution images. Therefore, effectively learning and extracting structured feature information from images during the training phase can improve the performance of machine learning models. Edge information is a fundamental and intuitive structured feature in images, indicating the contours of objects and important details. Therefore, correctly extracting edge information from images provides richer feature inputs for machine learning models.
[0052] In some embodiments, the model training system 150 extracts multiple types of features from the first sample image through multiple feature extraction operators in multiple processing branches of the feature extraction unit 220. In some embodiments, the feature extraction operators involved in the multiple processing branches may include the Soble operator. The Soble operator can efficiently capture the edge regions of an image by calculating the gradient of the image. Correspondingly, the multiple feature extraction operators involved in the multiple processing branches may also include operators such as Prewitt, Laplacian, Roberts, and Canny. The cross-differentiation algorithm (Roberts operator) detects edge lines by calculating local differences and can handle images with steep, low-noise edges. The Prewitt operator achieves edge detection by utilizing the differences generated by pixel grayscale values within a specific region. The Laplacian operator can be used for image enhancement and edge extraction. The Canny operator can determine the optimal edge detection solution or identify locations with strong grayscale intensity changes in an image. Introducing these types of operators during the training phase is to explore optimal combinations of these operators. Therefore, in the feature extraction unit 220, each processing branch (or each feature extraction operator) is also associated with a corresponding gating weight parameter, which is used to indicate the importance of the corresponding feature extraction operator.
[0053] Figure 3A An example structure of feature extraction unit 300 is shown. Feature extraction unit 220 may include 3*3 convolutional blocks 312, 1*1 convolutional blocks 311, and 3*3 convolutional blocks 312. Feature extraction unit 220 includes a processing branch consisting of 1*1 convolutional blocks 311, A operator_x313 (e.g., the x-axis direction of the Soble operator), and gate weight parameter a321. Feature extraction unit 220 includes a processing branch consisting of 1*1 convolutional blocks 311, A operator_y314 (e.g., the y-axis direction of the Soble operator), and gate weight parameter b322. Feature extraction unit 220 includes a processing branch consisting of 1*1 convolutional blocks 311, B operator_4315 (e.g., the small neighborhood of the Laplacian operator), and gate weight parameter c323. The feature extraction unit 220 includes a processing branch consisting of a 1*1 convolutional block 311, a B operator _8 316 (e.g., a large neighborhood of the Laplacian operator), and a gate weight parameter d 324.
[0054] Feature extraction unit 220 includes a processing branch consisting of a 1*1 convolutional block 311, a C operator _x 317 (e.g., the x-axis direction of the Prewitt operator), and a gate weight parameter e 325. Feature extraction unit 220 includes a processing branch consisting of a 1*1 convolutional block 311, a C operator _y 318 (e.g., the y-axis direction of the Prewitt operator), and a gate weight parameter f 326. Feature extraction unit 220 includes a processing branch consisting of a 1*1 convolutional block 311, a D operator _x 319 (e.g., the x-axis direction of the Roberts operator), and a gate weight parameter g 327. Feature extraction unit 220 includes a processing branch consisting of a 1*1 convolutional block 311, a D operator _320y (e.g., the x-axis direction of the Roberts operator), and a gate weight parameter h 328.
[0055] Understandable. Figure 3A Only one example structure of the feature extraction unit 220 is given. In practical applications, more, fewer, or different operators can be selected as needed to configure different example structures. The embodiments disclosed herein do not limit this aspect.
[0056] During training, the gating weight parameters of each processing branch of the feature extraction unit 220 are updated as the parameter set of the learning model is updated to determine the combination strategy between different operators. In some embodiments, the feature extraction unit 220 may also include an activation function 329, which can enable the feature extraction unit 220 to have stronger nonlinear learning capabilities. It is understood that the model training system 150, based on the multiple feature extraction operators and gating weight parameters in the multiple processing branches of the feature extraction unit 220, as well as the activation function 329, can obtain an output 330 for the LR sample image or intermediate feature information output by the feature extraction unit 220. For each feature extraction unit, the output 330 may be an HR image or intermediate feature information used as input to the next feature extraction unit.
[0057] In some embodiments, the model training system 150 generates a prediction super-resolution image for the first sample image based on feature information obtained via the feature extraction unit. Then, the model training system 150 updates the model parameter set of the machine learning model 130 based at least on the difference between the prediction super-resolution image and the second sample image to obtain a trained machine learning model 130. The model parameter set includes at least the gating weight parameters in the feature extraction unit 220. Of course, the model parameter set of the machine learning model 130 may also include parameters of other components in the feature extraction unit 220, such as parameters of convolutional blocks. Furthermore, if the machine learning model 130 includes other processing units besides the feature extraction unit 220, the model parameter set may also include parameters of those processing units.
[0058] During model training, supervised learning is used to compare the predicted super-resolution image obtained by machine learning model 130 based on the current model parameter set with the ground truth super-resolution sample image (i.e., the second sample image). This allows determination of whether the current model parameter set of machine learning model 130 has been updated to a satisfactory level. The training process of machine learning model 130 is performed iteratively based on multiple sample image pairs, and the model parameter set is continuously updated. The stopping condition for updating the model parameter set can be that the difference between the predicted super-resolution image output by the model and the ground truth super-resolution sample image is minimized or reduced to a predetermined target value, or that the number of iterations reaches a predetermined number. Various model training algorithms can be used to update the model parameter set. The following will refer to... Figure 4 This describes the model training process according to some embodiments of the present disclosure.
[0059] After training the feature extraction unit 220 through reparameterization, the feature extraction unit 220 contains a large number of parameters. In order to speed up model inference and save parameter storage space during the model inference stage, multiple processing branches in the trained feature extraction unit 220 can be merged to simplify the parameters. Figure 3B A schematic structural diagram of the merged feature extraction unit 220 in the model application stage according to some embodiments of the present disclosure is shown. The model application stage can be implemented in Figure 1 The model shown is applied to system 160.
[0060] In Figure 3A When multiple processing branches are merged, cascaded 1x1 and 3x3 convolutional blocks can be merged into a single convolutional block (e.g., a 3x3 convolutional block). Multiple feature extraction operators are treated as a special type of convolutional block, and 3x3 convolutional blocks with consistent shapes are merged. This transforms the feature extraction unit into a single convolutional block 352, thereby achieving efficient inference. Figure 3B As shown, the merged feature extraction unit 220 acquires the input LR image or intermediate feature information extracted from the LR image by the previous feature extraction unit. Input 351 (which may be the LR image or intermediate feature information) is provided to the merged convolutional block 352 to acquire feature information from the LR image or intermediate feature information. Activation function 353 is configured to enhance the non-linear learning capability of feature extraction unit 330, and by processing the output of convolutional block 352, the output 354 of the merged feature extraction unit 220 is obtained. Output 354 may be the final output of machine learning model 130 (i.e., the HR image) or intermediate feature information used as input to the next processing unit (e.g., the next feature extraction unit) in machine learning model 130.
[0061] The training process of machine learning model 130 will be described in detail below. As briefly described above, model training system 150 can generate a prediction super-resolution image for the LR sample image generated by the feature extraction unit 220 in machine learning model 130. Then, model training system 150 trains machine learning model 130 based on the difference between the prediction super-resolution image for the LR sample image and the HR sample image. (Continued below) Figure 4 To describe the set of model parameters for updating a machine learning model. Figure 4 A schematic diagram of an example architecture 400 for model training according to some embodiments of the present disclosure is shown.
[0062] In some embodiments, the model training system 150 utilizes a discriminator model to obtain predicted super-resolution partial image patches for a subset of image patches in a first sample image. Then, the model training system 150 updates the model parameter set of the machine learning model 130 based on the differences between the predicted super-resolution image and the second sample image, and the differences between the predicted super-resolution partial image patches and corresponding image patches in the second sample image. In some examples, the discriminator model (Patch Discriminator) is a discriminator used in generative adversarial networks (GANs). The discriminator model can divide the input image into multiple patches through convolutional operations and classify each patch as real or fake. In some examples, the discriminator model is updated during contrastive learning and can be discarded after training is complete.
[0063] like Figure 4 As shown, the model training system 150, based on the original image 510, performs a degradation operation on the original image using a degradation model 411 to obtain sample pairs 413 including HR sample images and LR sample images. Correspondingly, the model training system 150, based on the original image 410, obtains sample pairs 413 including HR sample images and LR sample images using a diffusion model 412. Subsequently, the model training system 150 uses a discriminator model 415 to obtain predicted super-resolution partial image patches for partial image patches in the LR sample images. Further reference will continue below. Figure 5 and Figure 6 The model training system 150 is described in detail to acquire the first sample image and the second sample image for training the machine learning model.
[0064] Then, the model training system 150 can determine a contrastive loss function (also known as GAN_loss) 417 based on the difference between the predicted super-resolution image patch and the predicted super-resolution image, and then update the model parameter set of the machine learning model 130 based at least on the contrastive loss function 417. In some examples, the model training system 150 can determine the contrastive loss function 417 based on the difference between the predicted super-resolution image and the second sample image, and the difference between the predicted super-resolution image patch and the second sample image patch, and then update the model parameter set of the machine learning model 130 based at least on the contrastive loss function 417.
[0065] In some examples, the machine learning model 130 can compute gradients based on the loss using a stochastic gradient descent training algorithm, and the model parameters of the machine learning model 130 to be trained can be updated via gradient backpropagation. In some examples, the model parameters of the discriminator model can also be updated as the machine learning model 130 to be trained is updated via gradient backpropagation.
[0066] In some embodiments, the model training system 150 may also perform a downsampling operation on the predicted super-resolution image to obtain a predicted downsampled image. In some embodiments, the resolution of the predicted downsampled image is lower than that of the predicted super-resolution image. Subsequently, the model training system 150 updates the model parameter set of the machine learning model 130 based on the differences between the predicted super-resolution image and the second sample image, and the differences between the predicted downsampled image and the first sample image.
[0067] like Figure 4 In the example architecture 400 shown, the model training system 150 performs a downsampling operation on the predicted super-resolution image obtained by the machine learning model 130 to obtain a predicted downsampled image with a lower resolution than the predicted super-resolution image. Then, the model training system 150 determines a multi-scale loss function 416 based on the differences between the predicted image and the HR sample image, and the differences between the predicted downsampled image and the LR sample image, thereby updating the model parameter set of the machine learning model 130.
[0068] Understandably, the model training system 150 can train the machine learning model 130 to be trained based on a multi-scale loss function and a contrastive loss function to obtain a trained machine learning model 130. In some examples, the model training system 150 can also collaboratively guide the training of the machine learning model 130 based on a Fourier frequency domain loss function. Thus, the reparameterized feature extraction module of this disclosure can convert the machine learning model 130 into a lightweight machine learning model 130 by merging redundant computational branches. In some embodiments, model compression can be further performed on the machine learning model 130 to obtain an even lighter machine learning model 130.
[0069] In some embodiments, the model training system 150 may utilize a trained benchmark machine learning model to remove a portion of the model parameters from the model parameter set of the machine learning model 130 through knowledge distillation. For example... Figure 4 As shown, the model training system 150 can utilize a trained baseline machine learning model (also known as a teacher model) 419 to remove some model parameters from the model parameter set of the machine learning model 130 using knowledge distillation. In this case, the machine learning model 130 can also be called a student model. Knowledge distillation (KD) is a model compression method that involves constructing a lightweight, smaller model and training it using the supervision information from a larger, more powerful model to achieve better performance and accuracy.
[0070] In some embodiments, the model training system 150 generates a first predictive super-resolution image based on a third sample image using a benchmark machine learning model 419. In some embodiments, the third sample image may be referred to as a low-resolution sample image (LR sample image). In some examples, the third sample image is different from the first sample image. Alternatively, the third sample image and the first sample image may be the same sample image. In some embodiments, the benchmark machine learning model 419 is configured to perform super-resolution processing on the input image to obtain the output image. In some examples, the model training system 150 generates a first predictive super-resolution image (also referred to as a high-resolution image) based on the LR sample image using the benchmark machine learning model 419.
[0071] Accordingly, the model training system 150 uses the machine learning model 130 to be compressed to generate a second predictive super-resolution image based on the third sample image. In some examples, the model training system 150 inputs the LR sample image into the machine learning model 130 to be compressed to obtain the second predictive super-resolution image (also referred to as a high-resolution image). Then, the model training system 150 removes some model parameters from the model parameter set of the machine learning model 130, at least based on the difference between the first and second predictive super-resolution images.
[0072] In some embodiments, the model training system 150 utilizes the machine learning model 130 to be compressed based on the noisy image I. Noise A third predicted super-resolution image is generated. Correspondingly, the model training system 150 uses a benchmark machine learning model to generate a fourth predicted super-resolution image based on the noisy image. For example... Figure 4 As shown, the noisy image may include random sample images 418. Then, the model training system 150 removes some model parameters from the model parameter set, at least based on the difference between the third and fourth prediction super-resolution images corresponding to the noisy image. Understandably, the model training system 150 may perform knowledge distillation on the machine learning model 130 based on a high-quality benchmark machine learning model as a teacher model, and introduce random noise input to guide the distillation direction of the machine learning model 130 to be compressed. In some embodiments, the model training system 150 may also use HR sample images to assist the machine learning model 130 to be compressed in performing knowledge distillation.
[0073] In some embodiments, the model training system 150 can determine the KD loss function 420 based on the differences between the first and second predicted super-resolution images, and the differences between the third and fourth predicted super-resolution images, thereby deleting some model parameters from the model parameter set. The KD loss function 420 can be expressed as follows:
[0074] L KD=MSE(T(I LR ), S(I LR ))+λ0MSE(T(I Noise ), S(I Noise )),
[0075] Where MSE(T(I) LR ), S(I LR The mean square error (MSE) between the first predicted super-resolution image generated by the baseline machine learning model 419 and the second predicted super-resolution image generated by the machine learning model 130 as a student model is indicated by the mean square error (MSE(T(I)). Noise ), s(I Noise The mean square error between the third predicted super-resolution image generated by the baseline machine learning model 419 and the fourth predicted super-resolution image generated by the machine learning model 130 is the error between the noisy super-resolution images.
[0076] In some embodiments, the model training system 150 may also remove some model parameters from the model parameters based on the image spectral difference between the second predicted super-resolution image and the fourth sample image. The fourth sample image is a super-resolution version of the third sample image, and can be referred to as a high-resolution sample image. In some examples, the model training system 150 bases its calculations on the spectrum (F(S(I)) of the second predicted super-resolution image obtained via the machine learning model 130 to be compressed. LR )), and the spectrum of HR sample images (F(I) HR The difference between the two is used to determine the Fourier loss function (L). Fourier This allows for the removal of some model parameters from the model parameters. The Fourier loss function can be expressed as follows:
[0077] L Fourier =L1(F(S(I LR ), F(I HR ))).
[0078] In some embodiments, the model training system 150 may also remove some model parameters from the model parameters based on the pixel differences between the second predicted super-resolution image and the fourth sample image. In some examples, the model training system 150 bases the pixel differences (S(I) of the second predicted super-resolution image obtained via the machine learning model 130 to be compressed on the pixel differences (S(I)). LR )), and the pixels of the HR sample image (I) HR The difference between the pixels determines the pixel loss function (L). Pixel This allows for the removal of some model parameters from the model parameters. The pixel loss function can be expressed as follows:
[0079] L Pixel =L1(S(ILR ), I HR ).
[0080] In some embodiments, the model training system 150 can remove some model parameters from the model parameter set based on a weighted sum of the KD loss function 420, the Fourier loss function, and the pixel loss function 421, to compress the machine learning model 130. The model training system 150 can be represented by the following formula based on the weighted sum of the KD loss function 420, the Fourier loss function, and the pixel loss function 421:
[0081] L total =λ1L KD +λ2L Fourier +λ3L Pixel ,
[0082] Where λ1 indicates the weight parameters corresponding to the KD loss function 420, λ2 indicates the weight parameters corresponding to the Fourier loss function, and λ3 indicates the weight parameters corresponding to the pixel loss function 421.
[0083] In some embodiments, the model training system 150 can also remove a portion of the model parameters from the model parameter set of the machine learning model through model pruning. In some embodiments, the machine learning model includes multiple processing kernels, each with corresponding model parameters. In some embodiments, the model training system 150 can modify the model parameter set of the machine learning model multiple times as follows: The model training system 150 sorts the model parameters of the multiple processing kernels based on the importance of their respective model parameters. Then, based on the sorting result, the model training system 150 selects at least one processing kernel from the multiple processing kernels for pruning. Accordingly, the model training system 150 configures the model parameters of the at least one selected processing kernel to zero.
[0084] The model training system 150 uses a machine learning model with modified model parameters to perform multiple training iterations based on sample images to obtain updated model parameters. Subsequently, if the target processing kernel among multiple processing kernels is selected for pruning more than a threshold number of times during multiple modifications and training iterations, the model training system 150 removes the model parameters corresponding to the target processing kernel from the model parameter set.
[0085] Understandably, the model training system 150 performs soft and hard pruning in stages during the compression of the machine learning model 130 to be compressed. Specifically, the model training system 150 performs repeated soft pruning at a high frequency and records the number of times each kernel is soft-pruned. After iterative distillation to a certain data interval, kernels that have been soft-pruned more than a threshold are selected for hard pruning. Understandably, the model training system 150 can directly delete the corresponding kernel. In some examples, the soft pruning instruction for a kernel resets its weight to 0 instead of directly deleting it, allowing it to continue being added to the training process.
[0086] The model training system 150 can perform soft and hard pruning in the following ways. The model training system 150 can use a Layer-Adaptive Magnitude Pruning (LAMP) pruning strategy, according to a certain data gap, to prune all weight matrices W in the l-th layer. i Calculate its L2 norm, w x,y These are the weight values at positions x and y in the i-th processing kernel: At this point, the model training system 150 can obtain the norm of all processing kernels in each layer, that is, the importance of all processing kernels in that layer.
[0087] Then, the model training system 150 can sort all the processing kernels in each layer according to their importance, let I be the set of importance vectors of all processing kernels in any layer. l Then it can be expressed as the following expression: I l =[(||W1||2,1),(||W2||2,2),…,(||W n [||2,n)], where (||W n ||2,n) represent the importance of the current processing kernel and the sequence number of the current processing kernel, respectively.
[0088] Subsequently, the model training system 150 can perform soft pruning by selecting at least one processing kernel of a corresponding proportion within each layer according to a preset pruning probability p. Wherein, top(1-p, I l This represents the index of the processing kernels that need to be retained, sorted by importance. At this point, the pruned weight matrix W can be defined. i 'for:
[0089]
[0090] For the weight matrix W iFor the processing kernels where '=0, the model training system 150 can set all weights of the corresponding processing kernels to zero using the method described above, performing one soft pruning. Accordingly, when performing soft pruning, the model training system 150 can maintain a pruning dictionary D to store the indices of the processing kernels undergoing soft pruning and their corresponding number of soft pruning operations. The structure of the pruning dictionary can be: D = {(i, c...} i | i∈Index of the processing kernel, c i This is the total number of times the nucleus was soft pruned.
[0091] In some embodiments, after each global soft pruning, all processing kernels (including those that have been soft pruned) re-enter the training process. After a certain number of training iterations, the model training system 150 can select processing kernels with a soft pruning hit rate greater than a threshold for hard pruning. The hit rate can be expressed as: Where N indicates the total number of processing kernels that have undergone soft pruning after training. For hard pruning, the final weight matrix W is removed. i "for:
[0092]
[0093] Among them, hard_prune_set is based on the hit rate r i The sorted selection is based on the set of weight indices that require hard pruning. For the weight matrix W... i For the processing kernels with "=0", the model training system 150 can directly delete the processing kernel and its coupled processing kernels from the network structure level to achieve hard pruning, following the method described above.
[0094] Therefore, this disclosure removes some model parameters from the parameter set of a machine learning model through knowledge distillation and intermittent pruning operations, enabling the machine learning model to be compressed to be trained into a more lightweight machine learning model. References will follow below. Figure 5 and Figure 6 The model training system 150 describes how it acquires first and second sample images for training a machine learning model. Figure 5 A schematic diagram of an example architecture 500 for acquiring sample images according to some embodiments of the present disclosure is shown. Figure 6 A schematic diagram of an example architecture 600 for acquiring sample images according to other embodiments of the present disclosure is shown.
[0095] In some embodiments, the model training system 150 may perform an image degradation operation on the original image to obtain a degraded image. In this scenario, the degraded image is determined as the first sample image, and the first original image is determined as the second sample image. Figure 5In the example architecture 500 shown, the model training system 150 performs a degradation operation 511 on the original image 510 using a degradation model. The model training system 150 then performs a fusion operation 516 on the original image after the degradation operation 511 to obtain a first sample image (LR sample image) 120. In this scenario, the original image 510 and the LR sample image 120 form a sample pair 522.
[0096] In some embodiments, the degradation operation includes at least one of the following: blurring, subsampling, image compression, and / or noise addition. It is understood that the model training system 150 can perform the degradation operation on the original image based on a gated degradation method. That is, the degradation operation d i This could be a blurring operation 512, a subsampling operation 513, an image compression operation 514, and / or a noise-adding operation 515. It can be represented by the following formula:
[0097] I LR =d i (I HR ), i = 1, 2, ..., N, I LR =D(I HR )=d1(d2(...d n (I HR ...)).
[0098] In some embodiments, the model training system 150 utilizes an image generative model to generate a super-resolution image based on the original image or the degraded image. In this scenario, the original image or the degraded image is determined as the first sample image, and the super-resolution image is determined as the second sample image. Figure 5 As shown, the model training system 150 can input LR sample images using the generative model 517 to obtain HR sample images 122. In some examples, the generative model 517 can call the diffusion model 518 to generate HR sample images 122 corresponding to the LR sample images 120. In this case, the LR sample images 120 and HR sample images 122 form a sample pair 523.
[0099] In some embodiments, the model training system 150 determines the original image as the first sample image and the super-resolution image as the second sample. For example... Figure 5 As shown, the original image 510 and the final HR sample image 122 form a sample pair 522.
[0100] In some examples, the model training system 150 can represent the process of super-resolution using a diffusion model as F SDSR In the process of using a diffusion model for super-resolution, the control network included in the model training system 150 will I LRAs additional control information input, it enables the diffusion model to output high-quality Ig. HR That is, the expression is: I HR =F SDSR (I LR Furthermore, the sample pair set 110 constructed by the model training system 150 based on the generative model can be represented as a set including multiple sample pairs: Where M represents the number of sample pairs in sample pair set 110.
[0101] like Figure 6 In the illustrated example architecture 600, the model training system 160 can also employ online data augmentation to acquire the first and second sample images. Online data augmentation enhances the randomness of the sample image readings by the machine learning model 130. Specifically, the model training system 160 performs an operation 612 to extract multiple patches from the original image 410. Subsequently, for a subset of image patches in the original image 410, degradation operations such as model operation 613, downsampling operation 614, upsampling operation 615, compression operation 616, and / or noise addition operation 617 are performed. Then, the model training system 150 performs a fusion operation 618 on the degraded portion of the original image 410 to obtain the LR sample image 120. Thus, this disclosure enables the construction of a high-quality sample pair set through a diffusion model and a degradation model.
[0102] In summary, this disclosure enhances the machine learning model's ability to understand complex features by utilizing reparameterized feature extraction units. Correspondingly, based on structural reparameterization, redundant computational branches can be merged to losslessly transform the machine learning model into a lightweight model. Furthermore, by setting a high-quality super-resolution model deployed off-client side as the teacher model, performing knowledge distillation on the machine learning model, and intermittently pruning using an interlayer weight importance strategy, the lightweight machine learning model becomes even more lightweight. This reduces computational complexity and cost.
[0103] Example process
[0104] Figure 7 A flowchart of a process 700 for model training according to some embodiments of the present disclosure is shown. Process 700 can be implemented at model training system 150. Reference is made below. Figure 1 Describe process 700.
[0105] In box 710, model training system 150 acquires a first sample image and a second sample image for training a machine learning model. The second sample image is a super-resolution version of the first sample image. The machine learning model is configured to perform super-resolution processing on the input image to obtain an output image with a higher resolution than the input image.
[0106] In box 720, the model training system 150 uses the feature extraction unit in the machine learning model to extract feature information of the first sample image. The feature extraction includes: in multiple processing branches of the feature extraction unit, multiple feature extraction operators are used to extract multiple types of features of the first sample image respectively. The multiple processing branches have corresponding gating weight parameters, and each gating weight parameter indicates the importance of the corresponding type of feature in the feature information.
[0107] In box 730, the model training system 150 generates a predicted super-resolution image for the first sample image based on feature information.
[0108] In box 740, the model training system 150 updates the model parameter set of the machine learning model based at least on the difference between the predicted super-resolution image and the second sample image to obtain a trained machine learning model, the model parameter set including at least gating weight parameters.
[0109] In some embodiments, updating the model parameter set of the machine learning model further includes: using a discriminator model to obtain a predicted super-resolution partial image patch for a partial image patch in a first sample image; and updating the model parameter set of the machine learning model based on the difference between the predicted super-resolution image and the second sample image, and the difference between the predicted super-resolution partial image patch and the corresponding image patch in the second sample image.
[0110] In some embodiments, updating the model parameters of the machine learning model includes: performing a downsampling operation on the predicted super-resolution image to obtain a predicted downsampled image, wherein the resolution of the predicted downsampled image is lower than that of the predicted super-resolution image; and updating the model parameter set of the machine learning model based on the difference between the predicted super-resolution image and the second sample image, and the difference between the predicted downsampled image and the first sample image.
[0111] In some embodiments, process 700 further includes: performing model compression on the trained machine learning model to obtain a compressed machine learning model by at least one of the following: using a trained baseline machine learning model to remove some model parameters from the model parameter set of the machine learning model through knowledge distillation; or removing some model parameters from the model parameter set of the machine learning model through model pruning.
[0112] In some embodiments, deleting a portion of the model parameters from the model parameter set of a machine learning model by knowledge distillation includes: generating a first predictive super-resolution image based on a third sample image using a benchmark machine learning model, the benchmark machine learning model being configured to perform super-resolution processing on the input image to obtain an output image; generating a second predictive super-resolution image based on the third sample image using the machine learning model to be compressed; and deleting a portion of the model parameters from the model parameter set based at least on the difference between the first predictive super-resolution image and the second predictive super-resolution image corresponding to the third sample image.
[0113] In some embodiments, deleting a portion of the model parameters from the model parameter set of a machine learning model by knowledge distillation further includes: generating a third predictive super-resolution image based on a noisy image using the machine learning model to be compressed; generating a fourth predictive super-resolution image based on a noisy image using a benchmark machine learning model; and deleting a portion of the model parameters from the model parameter set based at least on the difference between the third and fourth predictive super-resolution images corresponding to the noisy image.
[0114] In some embodiments, removing a portion of the model parameters from the model parameter set of a machine learning model by knowledge distillation further includes: removing a portion of the model parameters from the model parameter set based on the image spectral difference and / or pixel difference between the second predicted super-resolution image and the fourth sample image, wherein the fourth sample image is a super-resolution version of the third sample image.
[0115] In some embodiments, the machine learning model includes multiple processing kernels, each having corresponding model parameters. The process of deleting a portion of the model parameters from the model parameter set of the machine learning model through model pruning further includes: modifying the model parameter set of the machine learning model multiple times by: ranking the model parameters of the multiple processing kernels based on the importance of their respective model parameters; selecting at least one processing kernel from the multiple processing kernels for pruning based on the ranking result; configuring the model parameters of the selected at least one processing kernel to zero; performing multiple training operations on sample images using the machine learning model with the modified model parameter set; and deleting the model parameters corresponding to the target processing kernel from the model parameter set in response to the number of times the target processing kernel is selected for pruning in the multiple modifications and multiple training operations exceeding a threshold number.
[0116] In some embodiments, obtaining a first sample image and a second sample image for training a machine learning model includes one of the following: performing an image degradation operation on the original image to obtain a degraded image, wherein the degraded image is determined as the first sample image and the first original image is determined as the second sample image; generating a super-resolution image based on the original image or the degraded image using an image generative model, wherein the original image or the degraded image is determined as the first sample image and the super-resolution image is determined as the second sample image, or determining the original image as the first sample image and the super-resolution image as the second sample image.
[0117] In some embodiments, the image degradation operation includes at least one of the following: blurring operation, downsampling operation, image compression operation, and / or noise addition operation.
[0118] Example devices and equipment
[0119] Figure 8 A schematic structural block diagram of an apparatus 800 for model training according to certain embodiments of the present disclosure is shown. The apparatus 800 may be implemented as or included in a model training system 150. The various modules / components in the apparatus 800 may be implemented by hardware, software, firmware, or any combination thereof.
[0120] As shown in the figure, the device 800 includes a sample image acquisition module 810, configured to acquire a first sample image and a second sample image for training a machine learning model. The second sample image is a super-resolution version of the first sample image. The machine learning model is configured to perform super-resolution processing on the input image to obtain an output image, the resolution of which is higher than that of the input image. The device 800 also includes a feature information extraction module 820, configured to extract feature information from the first sample image using a feature extraction unit in the machine learning model. Feature information extraction includes: extracting multiple types of features from the first sample image through multiple feature extraction operators in multiple processing branches of the feature extraction unit. Each processing branch has a corresponding gate weight parameter, and each gate weight parameter indicates the importance of the corresponding type of feature in the feature information. The device 800 also includes a prediction super-resolution image generation module 830, configured to generate a prediction super-resolution image for the first sample image based on the feature information. The apparatus 800 also includes a parameter set update module 840, configured to update the model parameter set of the machine learning model based at least on the difference between the predicted super-resolution image and the second sample image, to obtain a trained machine learning model, the model parameter set including at least gating weight parameters.
[0121] In some embodiments, the parameter set update module 840 is further configured to use a discriminator model to obtain a predicted super-resolution partial image patch for a partial image patch in a first sample image; and to update the model parameter set of the machine learning model based on the difference between the predicted super-resolution image and the second sample image, and the difference between the predicted super-resolution partial image patch and the corresponding image patch in the second sample image.
[0122] In some embodiments, the parameter set update module 840 is further configured to perform a downsampling operation on the predicted super-resolution image to obtain a predicted downsampled image, the resolution of which is lower than that of the predicted super-resolution image; and to update the model parameter set of the machine learning model based on the difference between the predicted super-resolution image and the second sample image, and the difference between the predicted downsampled image and the first sample image.
[0123] In some embodiments, the apparatus 800 further includes a model parameter deletion module configured to perform model compression on the trained machine learning model by at least one of the following methods to obtain a compressed machine learning model: using a trained baseline machine learning model, deleting some model parameters from the model parameter set of the machine learning model through knowledge distillation; or deleting some model parameters from the model parameter set of the machine learning model through model pruning.
[0124] In some embodiments, the model parameter deletion module is further configured to generate a first predicted super-resolution image based on a third sample image using a benchmark machine learning model, the benchmark machine learning model being configured to perform super-resolution processing on the input image to obtain an output image; generate a second predicted super-resolution image based on the third sample image using a machine learning model to be compressed; and delete a portion of model parameters from the model parameter set based at least on the difference between the first predicted super-resolution image and the second predicted super-resolution image corresponding to the third sample image.
[0125] In some embodiments, the model parameter deletion module is further configured to generate a third predictive super-resolution image based on a noisy image using a machine learning model to be compressed; generate a fourth predictive super-resolution image based on a noisy image using a benchmark machine learning model; and delete a portion of the model parameters from the model parameter set based at least on the difference between the third and fourth predictive super-resolution images corresponding to the noisy image.
[0126] In some embodiments, the model parameter deletion module is further configured to delete a portion of the model parameters from the model parameter set based on the image spectral differences and / or pixel differences between the second predicted super-resolution image and the fourth sample image, where the fourth sample image is a super-resolution version of the third sample image.
[0127] In some embodiments, the machine learning model includes multiple processing kernels, each having corresponding model parameters. The model parameter deletion module is further configured to modify the model parameter set of the machine learning model multiple times by: ranking the model parameters of the multiple processing kernels based on the importance of their respective model parameters; selecting at least one processing kernel from the multiple processing kernels for pruning based on the ranking result; setting the model parameters of the selected at least one processing kernel to zero; performing multiple training operations on sample images using the machine learning model with the modified model parameter set to obtain an updated model parameter set; and deleting the model parameters corresponding to the target processing kernel from the model parameter set in response to the number of times the target processing kernel is selected for pruning in the multiple modifications and training operations exceeding a threshold number.
[0128] In some embodiments, the sample image acquisition module 810 is further configured to perform an image degradation operation on the original image to obtain a degraded image, wherein the degraded image is determined as a first sample image and the first original image is determined as a second sample image; and to generate a super-resolution image based on the original image or the degraded image using an image generative model, wherein the original image or the degraded image is determined as the first sample image and the super-resolution image is determined as the second sample image, or the original image is determined as the first sample image and the super-resolution image is determined as the second sample image.
[0129] In some embodiments, the image degradation operation includes at least one of the following: blurring operation, downsampling operation, image compression operation, and / or noise addition operation.
[0130] Figure 9 A block diagram is shown illustrating an electronic device 900 in which one or more embodiments of the present disclosure may be implemented. It should be understood that... Figure 9 The electronic device 900 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 9 The electronic device 900 shown can be used to achieve Figure 1 The model training system 150.
[0131] like Figure 9As shown, electronic device 900 is in the form of a general-purpose electronic device. Components of electronic device 900 may include, but are not limited to, one or more processors or processing units 910, memory 920, storage device 930, one or more communication units 940, one or more input devices 950, and one or more output devices 960. Processing unit 910 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 920. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 900.
[0132] Electronic device 900 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 900, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 920 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 930 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 900.
[0133] Electronic device 900 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 9 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 920 may include computer program product 925 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.
[0134] The communication unit 940 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 900 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 900 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0135] Input device 950 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 960 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 900 can also communicate with one or more external devices (not shown) via communication unit 940 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 900, or with any device that enables electronic device 900 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0136] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0137] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0138] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0139] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0140] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0141] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for training a model, comprising: A first sample image and a second sample image are obtained for training a machine learning model, the second sample image being a super-resolution version of the first sample image, the machine learning model being configured to perform super-resolution processing on the input image to obtain an output image, the output image having a higher resolution than the input image; The feature extraction unit in the machine learning model is used to extract feature information from the first sample image, wherein the extraction of feature information includes: In the multiple processing branches of the feature extraction unit, multiple types of features of the first sample image are extracted by multiple feature extraction operators respectively. The multiple processing branches have corresponding gate weight parameters, and each gate weight parameter indicates the importance of the corresponding type of feature in the feature information. Based on the feature information, a predicted super-resolution image for the first sample image is generated; as well as The model parameter set of the machine learning model is updated based at least on the difference between the predicted super-resolution image and the second sample image to obtain a trained machine learning model, wherein the model parameter set includes at least the gating weight parameters.
2. The method according to claim 1, wherein updating the model parameter set of the machine learning model further comprises: A discriminator model is used to obtain the predicted super-resolution partial image patches for a portion of the image patches in the first sample image; as well as Based on the difference between the predicted super-resolution image and the second sample image, and the difference between the predicted super-resolution image patch and the corresponding image patch in the second sample image, the model parameter set of the machine learning model is updated.
3. The method according to claim 1, wherein updating the model parameters of the machine learning model includes: A downsampling operation is performed on the predicted super-resolution image to obtain a predicted downsampled image, the resolution of which is lower than that of the predicted super-resolution image; as well as The model parameter set of the machine learning model is updated based on the difference between the predicted super-resolution image and the second sample image, and the difference between the predicted downsampled image and the first sample image.
4. The method according to claim 1, further comprising: The compressed machine learning model is obtained by performing model compression on the trained machine learning model using at least one of the following methods: Using a trained benchmark machine learning model, knowledge distillation is used to remove a portion of the model parameters from the model parameter set of the machine learning model; or Model pruning is used to remove a portion of the model parameters from the model parameter set of the machine learning model.
5. The method of claim 4, wherein deleting a portion of the model parameters from the model parameter set of the machine learning model by knowledge distillation comprises: The benchmark machine learning model is used to generate a first predicted super-resolution image based on a third sample image. The benchmark machine learning model is configured to perform super-resolution processing on the input image to obtain the output image. The machine learning model to be compressed is used to generate a second predictive super-resolution image based on the third sample image; as well as At least based on the difference between the first predicted super-resolution image and the second predicted super-resolution image corresponding to the third sample image, some model parameters are deleted from the model parameter set.
6. The method of claim 4, wherein deleting a portion of the model parameters from the model parameter set of the machine learning model by knowledge distillation further comprises: The machine learning model to be compressed is used to generate a third predictive super-resolution image based on the noisy image; The baseline machine learning model is used to generate a fourth predictive super-resolution image based on the noisy image; as well as At least based on the difference between the third predicted super-resolution image and the fourth predicted super-resolution image corresponding to the noisy image, some model parameters are deleted from the model parameter set.
7. The method of claim 5, wherein deleting a portion of the model parameters from the model parameter set of the machine learning model by knowledge distillation further comprises: Furthermore, based on the image spectral differences and / or pixel differences between the second predicted super-resolution image and the fourth sample image, a portion of the model parameters are removed from the model parameter set, wherein the fourth sample image is a super-resolution version of the third sample image.
8. The method according to claim 4, wherein the machine learning model comprises a plurality of processing kernels, each processing kernel having corresponding model parameters, wherein deleting a portion of the model parameters from the model parameter set of the machine learning model through model pruning further comprises: The model parameter set of the machine learning model is modified multiple times using the following methods. Based on the importance of the model parameters of each of the multiple processing cores, the model parameters of each of the multiple processing cores are ranked, and Based on the sorting results, at least one processing core is selected from the plurality of processing cores for pruning; Configure the model parameters of at least one of the selected processing cores to zero; Using the machine learning model with multiple modified model parameter sets, training is performed multiple times based on sample images to obtain an updated model parameter set; as well as In response to the fact that the number of times the target processing kernel among the plurality of processing kernels is selected for pruning in the multiple modifications and the multiple training sessions is greater than a threshold number, the model parameters corresponding to the target processing kernel are deleted from the model parameter set.
9. The method of claim 1, wherein acquiring the first sample image and the second sample image used for training the machine learning model comprises one of the following: An image degradation operation is performed on the original image to obtain a degraded image, wherein the degraded image is determined as the first sample image, and the first original image is determined as the second sample image; Using an image generative model, a super-resolution image is generated based on the original image or the degraded image, wherein the original image or the degraded image is determined as the first sample image, and the super-resolution image is determined as the second sample image, or... The original image is determined as the first sample image, and the super-resolution image is determined as the second sample image.
10. The method of claim 9, wherein the image degradation operation comprises at least one of the following: Fuzzy operation, Downsampling operation, Image compression operations, and / or Noise addition operation.
11. An apparatus for model training, comprising: The sample image acquisition module is configured to acquire a first sample image and a second sample image for training a machine learning model, wherein the second sample image is a super-resolution version of the first sample image, and the machine learning model is configured to perform super-resolution processing on the input image to obtain an output image, wherein the resolution of the output image is higher than the resolution of the input image. The feature information extraction module is configured to extract feature information of the first sample image using the feature extraction unit in the machine learning model, wherein the extraction of the feature information includes: In the multiple processing branches of the feature extraction unit, multiple types of features of the first sample image are extracted by multiple feature extraction operators respectively. The multiple processing branches have corresponding gate weight parameters, and each gate weight parameter indicates the importance of the corresponding type of feature in the feature information. A predictive super-resolution image generation module is configured to generate a predictive super-resolution image for the first sample image based on the feature information; and The parameter set update module is configured to update the model parameter set of the machine learning model based at least on the difference between the predicted super-resolution image and the second sample image to obtain a trained machine learning model, wherein the model parameter set includes at least the gating weight parameters.
12. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 10 when executed by the at least one processing unit.
13. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 10.
14. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1 to 10.