Data processing method for image enhancement model, electronic device, and storage medium
By setting the sample type acquisition ratio and adjusting the acquisition ratio during the training of the image enhancement model, the long-tail distribution problem in image enhancement model training is solved, and the overall processing effect of the model is improved.
Patent Information
- Application Number
- PCT/CN2024/130584
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-18
- Filing Date
- 2024-11-07
- Publication Date
- 2025-06-26
AI Technical Summary
Existing image enhancement models are prone to long-tail distribution problems during training, resulting in poor processing of certain types of image.
By setting the sample type collection ratio, training samples are collected on multiple sets of different types of training sample, forming training batch samples, and adjusting the collection ratio according to the training effect until the training termination condition is reached.
It effectively avoids the long-tail effect, improves the training effect of the image enhancement model, and makes the model more balanced in processing different types of images.
Smart Images

Figure CN2024130584_26062025_PF_FP_ABST
Abstract
Description
Data processing method, electronic device and storage medium for image enhancement model
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 18, 2023, with application number 202311750568.0, and application name “Data processing method, electronic device and storage medium for image enhancement model”, all contents of which are incorporated by reference into this application. Technical Field
[0002] The embodiments of the present application relate to the field of computer technology, and in particular to a data processing method, electronic device, and storage medium for an image enhancement model. Background Art
[0003] With the development of computer technology, people often need to electronically scan documents in their daily lives and work, such as electronic books, invoice reimbursement, work document scanning and printing, document material scanning, and other application scenarios.
[0004] Because professional scanners are expensive, scanning software has emerged. This software allows users to conveniently and cost-effectively scan documents anytime, anywhere. One current approach is to implement this functionality through image enhancement models. These models enhance the input image, producing a scanned image with an enhanced filter.
[0005] While the use of image enhancement models makes software scanning more convenient, they require training before use. Currently, a common training method involves randomly extracting batches of training samples from a total training sample set containing various types of scanned image samples to train the image enhancement model. However, this method is prone to long-tail distribution issues in model training, especially when the number of training samples for certain types is far less than that for other types. This can result in the trained image enhancement model performing well on some image types but poorly on others, resulting in poor overall image enhancement performance.
[0006] Summary of the Invention
[0007] In view of this, an embodiment of the present application provides a data processing solution for an image enhancement model to at least partially solve the above-mentioned problems.
[0008] According to a first aspect of an embodiment of the present application, a data processing method for an image enhancement model is provided, comprising: collecting training samples from a plurality of different types of training sample sets according to a set sample type collection ratio to obtain training batch samples; training the image enhancement model using the training batch samples, and obtaining losses corresponding to different types of samples after completing a preset round of training; adjusting the sample type collection ratio based on the loss, and re-collecting training samples from the plurality of different types of training sample sets according to the adjusted sample type collection ratio to form new training batch samples, and returning to the operation of training the image enhancement model using the training batch samples to continue execution until a training termination condition is reached.
[0009] According to a second aspect of an embodiment of the present application, a data processing method for an image enhancement model is provided, comprising: acquiring a scanned image to be processed; performing image enhancement processing on the scanned image through an image enhancement model to obtain an enhanced scanned image; wherein the image enhancement model is trained based on any one of the methods provided in the first aspect.
[0010] According to the third aspect of an embodiment of the present application, an electronic device is provided, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the method described in the first aspect or the second aspect.
[0011] According to a fourth aspect of the embodiments of the present application, a computer storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method described in the first aspect or the second aspect is implemented.
[0012] According to a fifth aspect of an embodiment of the present application, a computer program product is provided, comprising computer instructions, which instruct a computing device to perform operations corresponding to the method described in the first aspect or the second aspect.
[0013] According to the solution provided in the embodiment of the present application, when training the image enhancement model, unlike the traditional method of randomly collecting training samples to train the image enhancement model, the solution in the embodiment of the present application adopts a method of setting the sample type collection ratio, setting different collection ratios for different types of training samples. Thus, even if there are fewer training samples of certain types, according to the collection ratio of the sample type, a number of training samples that match its ratio will be extracted from its corresponding set, and combined with training samples of other types to form training batch samples for training the image enhancement model. In this way, even if the number of training samples of this type accounts for a small proportion of the total training samples, these fewer training samples can be collected according to the collection ratio to be combined with training samples of other types that account for a larger proportion, forming a large number of training batch samples that can meet the training needs. Thus, the long-tail effect of image enhancement model training caused by the small number of training samples of this type is avoided, and the training effect of the image enhancement model is improved. Furthermore, after the image enhancement model is trained for a preset number of rounds using the current training batch samples, the training quality of a certain type of training sample will be evaluated based on the model's training effect, that is, the losses corresponding to different types of samples obtained through model training. The sample type collection ratio will then be adjusted based on the loss to fully solve the long-tail effect problem generated in traditional training methods and further improve the model training effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0015] FIG1 is a schematic diagram of an exemplary system applicable to the embodiment of the present application.
[0016] FIG2 is a flowchart of an optional data processing method for an image enhancement model in the first aspect of the present application.
[0017] FIG3 is a flowchart of an optional sub-step of “obtaining losses corresponding to different types of samples” in step S104 .
[0018] FIG4 is a schematic diagram of an optional processing process of the data processing method in the first aspect of the present application.
[0019] FIG5 is a flowchart of an optional data processing method for an image enhancement model in the second aspect of the present application.
[0020] FIG6 is a schematic diagram of an optional scenario of the data processing method in the second aspect of the present application.
[0021] FIG7 is a schematic structural diagram of an optional electronic device of the present application. DETAILED DESCRIPTION
[0022] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field should fall within the scope of protection of the embodiments of the present application.
[0023] The specific implementation of the embodiment of the present application is further explained below in conjunction with the accompanying drawings of the embodiment of the present application.
[0024] Figure 1 illustrates an exemplary system applicable to the embodiments of the present application. As shown in Figure 1 , the system 100 may include a cloud service 102, a communication network 104, and / or one or more user devices 106, with Figure 1 illustrating multiple user devices. It should be noted that the embodiments of the present application can be independently implemented by the cloud service 102, or by a user device 106 with high-performance hardware and software.
[0025] The cloud server 102 can be any suitable device for storing information, data, programs, and / or any other suitable type of content, including but not limited to a distributed storage system device, a server cluster, a computing cloud server cluster, etc. In some embodiments, the cloud server 102 can perform any suitable function. For example, when the cloud server 102 independently implements the solution of the embodiments of the present application, in some embodiments, the cloud server 102 can be used to execute a data processing method for training an image enhancement model and / or execute a data processing method for performing image enhancement processing on a scanned image using the image enhancement model. As an optional example, in some embodiments, the cloud server 102 may first collect training samples from multiple different types of training sample sets according to a set sample type collection ratio to obtain training batch samples; then, the image enhancement model may be trained using the training batch samples, and after completing a preset round of training, the losses corresponding to the different types of samples are obtained; then, the sample type collection ratio is adjusted based on the losses, and training samples are collected again from multiple different types of training sample sets according to the adjusted sample type collection ratio to form a new training batch sample, and the operation of training the image enhancement model using the training batch samples is returned to continue until the training termination condition is met, thereby obtaining a trained image enhancement model. In some embodiments, after training the image enhancement model, the cloud server 102 may obtain a scanned image to be processed, perform image enhancement processing on the scanned image using the trained image enhancement model, and obtain an enhanced scanned image. In some embodiments, the cloud server 102 may receive a scanned image to be processed sent by the user device 106, perform image enhancement processing on the scanned image using the image enhancement model to generate an enhanced scanned image, and then return it to the user device 106.
[0026] In some embodiments, the communication network 104 can be any suitable combination of one or more wired and / or wireless networks. For example, the communication network 104 can include any one or more of the following: the Internet, an intranet, a wide area network (WAN), a local area network (LAN), a wireless network, a digital subscriber line (DSL) network, a frame relay network, an asynchronous transfer mode (ATM) network, a virtual private network (VPN), and / or any other suitable communication network. The user device 106 can be connected to the communication network 104 via one or more communication links (e.g., communication link 112), and the communication network 104 can be linked to the cloud service end 102 via one or more communication links (e.g., communication link 114). The communication link can be any communication link suitable for transmitting data between the user device 106 and the cloud service end 102, such as a network link, a dial-up link, a wireless link, a hard-wired link, any other suitable communication link, or any suitable combination of such links.
[0027] The user device 106 may include any one or more user devices suitable for presenting images, interacting with users, etc. When the solution of the embodiment of the present application is independently completed by the user device 106, the user device 106 can be used for a data processing method for training an image enhancement model, and / or a data processing method for performing image enhancement processing on a scanned image through an image enhancement model. As an optional example, in some embodiments, the user device 106 may first collect training samples from a plurality of different types of training sample sets according to a set sample type collection ratio to obtain training batch samples; then, the image enhancement model may be trained using the training batch samples, and after completing a preset round of training, the losses corresponding to the different types of samples are obtained; then, the sample type collection ratio is adjusted based on the loss, and training samples are collected again from a plurality of different types of training sample sets according to the adjusted sample type collection ratio to form a new training batch sample, and the operation of training the image enhancement model using the training batch samples is returned to continue until the training termination condition is reached, thereby obtaining a trained image enhancement model. In some optional embodiments, after training the image enhancement model, the user device 106 may obtain a scanned image to be processed and perform image enhancement processing on the scanned image using the trained image enhancement model to obtain an enhanced scanned image. In some embodiments, the user device 106 may include any suitable type of device. For example, in some embodiments, the user device 106 may include a mobile device, a tablet computer, a laptop computer, a desktop computer, and / or any other suitable type of user device.
[0028] Based on the above system, an embodiment of the present application provides a data processing solution for an image enhancement model, which is illustrated below through multiple embodiments.
[0029] FIG2 is a flowchart of a data processing method for an image enhancement model according to an embodiment of the present application. According to a first aspect of the present application, a data processing method for an image enhancement model is provided. As shown in FIG2 , the method includes steps S102, S104, and S106. Specifically:
[0030] S102: Collect training samples from a plurality of different types of training sample sets according to a set sample type collection ratio to obtain training batch samples.
[0031] It should be noted that the data processing method for the image enhancement model in this embodiment is a training method for training the image enhancement model. Using such a training method, the training effect of the image enhancement model can be improved. It should be understood that the image enhancement model in this application can be used to perform image enhancement processing on images, but the specific target is not limited in this application. To facilitate the description of the embodiments of this application, this application uses the application scenario of the image enhancement model used in scanning software to perform image enhancement processing on scanned images as an example.
[0032] For example, image enhancement processing can also be considered as filter processing, including but not limited to implementing enhancement filter functions, image document enhancement functions, background removal functions, image impurity removal functions, shadow removal filter functions, etc. Image enhancement processing can emphasize the overall or local characteristics of an image, such as making an originally unclear image clearer or emphasizing certain interesting features, expanding the differences between the features of different objects in the image, and suppressing uninteresting features, thereby improving image quality, enriching the amount of information, enhancing image interpretation and recognition, and meeting certain analytical needs, etc.
[0033] Unlike the traditional method where all training samples are uniformly located in the same large training data set, in the present application, the training samples are divided into multiple different sets according to the sample type, and each set contains multiple training samples of the same type. In a simple example, assume that there are three different types of training sample sets, where training set 1 is a set of sample images containing tables {11, 12, 13}, training set 2 is a set of sample images containing documents {21, 22, 23, ... 2M}, and training set 3 is a set of sample images containing business licenses {31, 32, 33, 34 ... 3N}. In the traditional method, these samples are collected in a large set, but it is obvious that the number of samples in training set 1 is small. Therefore, when randomly sampling the large set, the three training samples {11, 12, 13} may not be collected, or even if they are collected, when the collected training samples contain a large number of samples from the other two training sets, because the sample images containing tables account for too small a proportion, the model's training effect on this type of image is also poor, resulting in a long tail effect. However, through this application, assuming that the sample type collection ratio is 1:1:1, although the training set 1 only includes three sample images, the sample images in the training set 1 can be collected every time to form training batch samples with other types of sample images in other training sets, ensuring that their proportion in each training batch sample is relatively balanced to overcome the long-tail effect.
[0034] In this application, the type of training samples can be determined by the content included in the training samples. For example, taking the image enhancement model used to perform image enhancement processing on scanned images as an example, the training sample is a scanned image, and the type of training samples can be determined by the content mainly contained in the scanned image. For example, the content can include but is not limited to books, invoices, contracts, business licenses, work documents, certificate materials, identity cards, driver's licenses, etc., and the type of training samples can include but is not limited to book types, invoice types, contract types, business license types, work document types, certificate material types, identity card types, driver's license types, etc. Correspondingly, the types of training sample sets can include but are not limited to book types, invoice types, contract types, business license types, work document types, certificate material types, identity card types, driver's license types, etc. It should be noted that the above types are only exemplary. In actual applications, those skilled in the art can set the required types according to actual needs, and this application does not impose specific restrictions on this.
[0035] This means that, in some optional embodiments, the multiple different types of training sample sets in this application may include: dividing the multiple different types of training sample sets according to the different types of content contained in the sample images. Based on this, the multiple different types of training sample sets divided in this manner in this application can effectively meet the requirements for training image enhancement models used in scanning software, which is conducive to improving model training results.
[0036] Optionally, in the present application, an optional way of obtaining the source of training samples is that various users use scanning software to scan documents. For example, taking the document as an invoice as an example, the training sample can be obtained in the following optional way: the user places the invoice to be scanned on a support (such as a desktop, etc.) or holds it in his hand, and then uses the camera on the mobile terminal to shoot the invoice from the direction of the invoice to obtain the corresponding document image. The document image can be used as a training sample and divided into a training sample set of the invoice type (which can be manually marked or divided by a machine learning model). Of course, this is just an example and does not serve as any limitation to the present application. It should be understood that in this case, the training sample can be obtained by shooting when needed and then directly obtained, or it can be shot in advance and then the training sample image is stored in a predetermined storage space (including but not limited to storage media such as disks, hard disks, memories, or databases, etc.), and directly obtained from the storage space when needed. There is no limitation on this in the present application. Alternatively, another optional way to obtain the training samples is to randomly or specifically generate the required training samples through any suitable algorithm for generating training samples (such as a machine learning model) to divide them into corresponding training sample sets. Similarly, the generated training samples can be first stored in a predetermined storage space and directly retrieved from the storage space when needed. Alternatively, training samples can be obtained by other optional methods to form a training sample set, and this application does not limit this.
[0037] In step S102, the sample type collection ratio can be set as needed. For example, in a feasible embodiment, for a plurality of different types of training sample sets, initially, the sample type collection ratio can be an evenly divided ratio. When collecting training samples for a plurality of different types of training sample sets, the collection can be performed randomly or according to predetermined rules. For example, when training a model, two types of training sample sets are used (for example, training sample set A and training sample set B), then a sample type collection ratio of 1:1 (that is, 50% each) can be set. For example, if a total of N training samples are required for training, 0.5N training samples can be extracted and collected from training sample set A, and then 0.5N training samples can be extracted and collected from training sample set B to generate training batch samples for training the image enhancement model. For another example, when training a model, four types of training sample sets are used (for example, training sample sets A, B, C, and D). Then, a sample type collection ratio of 1:1:1:1 (i.e., 25% each) can be set. For example, if a total of N training samples are required for training, 0.25N training samples can be extracted from each of the training sample sets A, B, C, and D to generate training batch samples for training the image enhancement model. It should be understood that this is only for some examples and is not a limitation of this application. Other sample type collection ratios can also be used. For example, taking two types of training sample sets as an example, 9:11 (i.e., 45% and 55% respectively), 49:51 (i.e., 49% and 51% respectively), etc. can be used to meet the needs.
[0038] After collecting different types of training samples according to the sample type collection ratio, a training batch sample, also called a batch sample, can be obtained, which includes multiple different types of training sample images.
[0039] Optionally, when training the image enhancement model, supervised training can be performed. The collected training samples are images that have not been processed by image enhancement. For each collected training sample, the generated training batch samples include multiple different types of training sample image pairs. Each training sample image pair includes the original training sample image and the enhanced processed image corresponding to the original training sample image (which can be pre-stored).
[0040] It is understandable that, when training the image enhancement model in the present application, unlike the traditional method of randomly collecting training samples to train the image enhancement model, the solution of the embodiment of the present application adopts a method of setting the sample type collection ratio, setting different collection ratios for different types of training samples. Thus, even if there are fewer training samples of certain types, according to the sample type collection ratio, a number of training samples that match its ratio will be extracted from its corresponding set, and combined with other types of training samples to form training batch samples for training the image enhancement model. In this way, even if the number of training samples of this type is relatively small in the total training samples, these fewer training samples can be collected according to the collection ratio to be combined with other types of training samples that account for a larger number, forming a large number of training batch samples that can meet the training needs. Thus, the long-tail effect of image enhancement model training caused by the small number of training samples of this type is avoided, and the training effect of the image enhancement model is improved.
[0041] S104: The image enhancement model is trained using the training batch samples, and after completing a preset round of training, the losses corresponding to different types of samples are obtained.
[0042] In the present application, after obtaining the training batch samples, the image enhancement model can be trained for a predetermined round. Before the end of the predetermined round of training, the training samples are not replaced to obtain the losses corresponding to different types of samples in the training batch samples, so as to facilitate further processing based on the losses in the subsequent step S106.
[0043] The specific method for training the image enhancement model using training batch samples in this application can be selected based on the model structure of the image enhancement model. The image enhancement model in this application can adopt any suitable model structure, as long as it can achieve the image enhancement function. In some optional embodiments, the image enhancement model can be a UNet model, for example, a machine learning model using an encoder-decoder structure. In traditional training methods, each iterative training for each sample in each training batch will output a corresponding loss value. However, in order to save model training costs and more specifically address the long-tail effect problem of traditional training methods, in this application, the corresponding loss will not be obtained until the image enhancement model has been trained for a preset number of rounds. Before this, there is no need to calculate the model loss. The specific number of preset rounds can be appropriately set by those skilled in the art according to actual needs and is not limited by this application. For example, it can be set to 10 times. Each round means that all samples in a training batch have undergone one round of training.
[0044] Since adjustments need to be made for different types of training samples in the future, the losses corresponding to the different types of samples obtained in step S104 can be used to evaluate the quality of training of the image enhancement model by different types of training samples. Based on this, the loss function can adopt, for example, L1 loss function, L2 loss function, loss function for evaluating image quality, etc., wherein the loss function for evaluating image quality may include but is not limited to Peak Signal-to-Noise Ratio (PSNR) loss function, Structural Similarity Index Measure (SSIM) loss function, etc. These loss functions may include multiple loss calculation parts, and different loss calculation parts may correspond to different types of training samples to calculate the loss.
[0045] In some optional embodiments, referring to the flowchart shown in FIG3 , the step S104 of “obtaining losses corresponding to different types of samples” includes sub-steps S1041 and S1042 :
[0046] S1041: Obtain a test sample.
[0047] In order to more accurately obtain the effect of the image enhancement model after a preset round of training, in this application, test samples are used to determine the training effect of the image enhancement model and obtain the corresponding loss. Similar to training samples, test samples also include many different types.
[0048] Generally, the sample data set can be divided into a training sample set and a test sample set. The training sample set is mainly used to train the model, while the test sample set is used to test the trained model. In this application, the test sample can come from the test sample set.
[0049] The test samples in the test sample set are not used for model training, but only for determining the loss. Test samples can be collected from different types of test sample sets, such as 1%, 3%, 5%, etc. The specific selection can be based on needs.
[0050] S1042: Perform image enhancement processing on the test sample using the image enhancement model that has completed the preset rounds of training, and obtain the losses corresponding to different types of test samples according to the processing results.
[0051] Taking the preset number of rounds as 10 as an example, the image enhancement model after 10 rounds of training can be used to perform image enhancement processing on the test samples. Then, the loss values corresponding to different types of test samples are obtained by calculating the loss values according to the aforementioned loss function.
[0052] In one feasible approach, a batch of test samples can be collected from the test sample set, including test samples of various types, and then processed through the image enhancement model to obtain the losses corresponding to each type.
[0053] Based on this, in this application, through the optional implementation of the above sub-steps S1041 to S1042, the losses corresponding to different types of test samples can be accurately and effectively obtained, so as to facilitate the subsequent model training process based on the losses.
[0054] In some optional embodiments, "obtaining losses corresponding to different types of test samples according to the processing results" in step S1042 includes: obtaining losses corresponding to different types of test samples according to the processing results and a preset loss function; wherein the loss function includes: multiple loss calculation parts that reflect the losses corresponding to each type of test sample.
[0055] As mentioned above, the loss function in this application can be: L1 loss function, L2 loss function, loss function for evaluating image quality, etc.
[0056] In this application, based on the above-mentioned processing results of the image enhancement processing of the test samples, the losses corresponding to each type of training samples are calculated respectively through multiple loss calculation parts of the loss function to obtain the losses corresponding to different types of test samples, thereby accurately and effectively obtaining the losses corresponding to different types of test samples, so as to facilitate the subsequent model training process based on the loss.
[0057] In some optional embodiments, the image enhancement model is a machine learning model based on an encoder and decoder structure. Unlike traditional encoder and decoder structures, the encoder in the image enhancement model of the present application includes an encoding layer and an attention layer; then, "training the image enhancement model using training batch samples" in step S104 may include: for each sample in the training batch samples, encoding processing and attention processing of the sample based on the encoding layer and attention layer in the encoder, as well as decoding processing by the decoder, to train the image enhancement model.
[0058] Based on this, on the one hand, in this application, the image enhancement model can be effectively trained by encoding each sample in the training batch samples through the encoding processing of the encoding layer in the encoder and the attention processing of the attention layer, as well as the decoding processing of the decoder; on the other hand, since the image resolution of the samples in the training batch samples is generally high, the traditional encoder-decoder structure model will reduce resource consumption by reducing the image resolution in the encoding stage, but it will also bring about the problem of insufficient receptive field (receptive field refers to the size of the area where the pixels on the feature map output by each layer of the convolutional neural network are mapped back to the input image), thereby affecting the restoration effect of large areas. The encoder of the image enhancement model structure in this application includes an attention layer, and then the attention layer can perform attention processing on each sample in the training batch samples during model training, thereby effectively improving the receptive field of the trained image enhancement model and improving the model training effect. Moreover, when the image enhancement model is used for image enhancement processing of scanned images, it can effectively improve the restoration effect of the image enhancement model for large areas of scanned images.
[0059] The attention layer is implemented based on the attention mechanism. Due to the existence of the attention layer, during the feature extraction stage when performing attention processing on samples in the training batch, the features of the current area of the sample and the features of the surrounding areas of the current area can be comprehensively considered before extracting features, which is beneficial to improving the receptive field of the trained image enhancement model and improving the model training effect.
[0060] In this application, the specific structure of the encoder can be implemented through various optional structures. For example, in some optional embodiments, the encoder includes multiple encoding layers and a self-attention layer connected to each encoding layer; the "encoding processing and attention processing performed on the sample by the encoding layer and attention layer in the encoder" includes: sequentially performing encoding processing and self-attention processing on the sample through each encoding layer in the encoder and the self-attention layer connected to each encoding layer.
[0061] In this way, the attention processing is implemented as self-attention processing. Based on this, through such an optional encoder structure, on the one hand, in this application, each sample in the training batch samples can be encoded and self-attended in sequence through each encoding layer in the encoder and the self-attention layer connected to each encoding layer, and decoded by the decoder, so that the image enhancement model can be effectively trained; on the other hand, since each encoding layer in the multiple encoding layers included in the encoder is connected to a self-attention layer, multiple self-attention layers can perform self-attention processing on each sample in the training batch samples during model training, which can more effectively improve the receptive field of the trained image enhancement model and better improve the model training effect. Moreover, when the image enhancement model is used for image enhancement processing of scanned images, the restoration effect of the image enhancement model for large areas of the scanned images can be more effectively improved.
[0062] It should be noted that in such an optional encoder structure, the number of encoding layers and self-attention layers can be reasonably selected according to the needs of actual model training, and is not specifically limited in this application.
[0063] For example, an optional encoder includes 10 encoding layers (recorded as encoding layers 1 to 10) and 10 self-attention layers (recorded as self-attention layers 1 to 10). The encoder structure can be connected in order (from beginning to end): encoding layer 1, self-attention layer 1, encoding layer 2, self-attention layer 2, encoding layer 3, self-attention layer 3, ..., encoding layer 10, self-attention layer 10; when encoding and attention processing are performed on the training samples, they can be performed in sequence according to the above connection relationship. It should be understood that this is only an example for ease of understanding and is not intended to limit the present application.
[0064] In other optional embodiments, the encoder includes multiple encoding layers, and some of the multiple encoding layers are connected to self-attention layers; the encoding processing and attention processing performed by the encoding layer and the attention layer in the encoder on the sample include: inputting the sample into the encoder, and performing encoding processing and self-attention processing on the sample in sequence according to the connection relationship between the encoding layer and the self-attention layer in the encoder.
[0065] Based on this, through such an optional encoder structure, on the one hand, in this application, each sample in the training batch samples can be encoded and self-attention processed in sequence through the encoding layer in the encoder and the self-attention layer connected to part of the encoding layer, and the decoder can be decoded, so that the image enhancement model can be effectively trained; on the other hand, since some of the multiple encoding layers included in the encoder are connected to the self-attention layer, multiple self-attention layers can perform self-attention processing on each sample in the training batch samples during model training, which can more effectively improve the receptive field of the trained image enhancement model and better improve the model training effect. Moreover, when the image enhancement model is used for image enhancement processing of scanned images, the image enhancement model can more effectively improve the restoration effect of large areas of the scanned images; on the other hand, for the case where there are many encoding layers in the encoder structure, if a self-attention layer is connected after each encoding layer, it is easy to cause the video memory to be occupied and the prediction time to be too long during the model training process. Therefore, connecting the self-attention layer after part of the encoding layer can also improve the calculation speed and reduce the calculation time.
[0066] It should be noted that in such an optional encoder structure, the number of coding layers and self-attention layers can be reasonably selected according to the needs of actual model training, and is not specifically limited in this application. In addition, the setting position of the self-attention layer can be determined as needed and is not limited here. For example, a self-attention layer may be connected to each coding layer in the latter part of the encoder (such as the last three coding layers), while no self-attention layer may be set after each coding layer in the front part; or a self-attention layer may be connected after every N (N≥2 and a positive integer) coding layers; or the setting position of the self-attention layer may be determined according to other rules.
[0067] For example, taking an optional encoder including 10 coding layers (recorded as coding layers 1 to 10 in sequence) and 4 self-attention layers (recorded as self-attention layers 1 to 4 in sequence) as an example, the encoder structure can be in the order of connection (from beginning to end): coding layer 1, coding layer 2, coding layer 3, coding layer 4, self-attention layer 1, coding layer 5, coding layer 6, coding layer 7, coding layer 8, self-attention layer 2, coding layer 9, self-attention layer 3, coding layer 10, self-attention layer 4. Then, when encoding and attention processing are performed on the training samples, they can be performed in sequence according to the above connection relationship. It should be understood that this is only an example for ease of understanding and is not intended to limit the present application.
[0068] S106: Adjust the sample type collection ratio based on the loss, and re-collect training samples from multiple different types of training sample sets according to the adjusted sample type collection ratio to form new training batch samples, and return to the operation of training the image enhancement model using the training batch samples to continue executing until the training termination condition is reached.
[0069] During the training process, the sample type collection ratio can be adjusted based on the losses corresponding to each type of test samples, and training samples can continue to be collected from different types of training sample sets according to the adjusted ratio to form new training batch samples, and then the image enhancement model can continue to be trained until it reaches the training termination condition.
[0070] Based on this, the technical solution of the above steps S102 to S106 in this application, when training the image enhancement model, is different from the traditional method of randomly collecting training samples to train the image enhancement model. The solution of the embodiment of the present application adopts a method of setting the sample type collection ratio, setting different collection ratios for different types of training samples. Thus, even if there are fewer training samples of certain types, according to the sample type collection ratio, a number of training samples that match its ratio will be extracted from its corresponding set, and combined with other types of training samples to form training batch samples for training the image enhancement model. In this way, even if the number of training samples of this type accounts for a small proportion of the total training samples, these fewer training samples can be collected according to the collection ratio to be combined with other types of training samples that account for a larger proportion, forming a large number of training batch samples that can meet the training needs. Thus, the long-tail effect of image enhancement model training caused by the small number of training samples of this type is avoided, and the training effect of the image enhancement model is improved. Furthermore, after the image enhancement model is trained for a preset number of rounds using the current training batch samples, the training quality of a certain type of training sample will be evaluated based on the model's training effect, that is, the losses corresponding to different types of samples obtained through model training. The sample type collection ratio will then be adjusted based on the loss to fully solve the long-tail effect problem generated in traditional training methods and further improve the model training effect.
[0071] In this application, when adjusting the sample type collection ratio based on loss, any appropriate strategy may be adopted, and is not limited herein. In some optional embodiments, the "adjusting the sample type collection ratio based on loss" in step S106 includes: for each type of training sample, if the corresponding loss is greater than a preset loss threshold, increasing the collection ratio of training samples of that type, and reducing the collection ratio of training samples of other types.
[0072] Based on this, in this application, the sample type collection ratio can be adjusted more reasonably and effectively based on the loss in this way, so that training samples can be re-collected from multiple different types of training sample sets according to the adjusted sample type collection ratio to form new training batch samples to continue model training, thereby facilitating a more comprehensive solution to the long-tail effect problem generated in traditional training methods and further improving the model training effect.
[0073] Optionally, "adjusting the sample type collection ratio based on the loss" in step S106 may further include: for each type of training sample, if the corresponding loss is less than a preset loss threshold, then either no adjustment is made or the collection ratio of training samples of that type is reduced. When the loss is less than the preset loss threshold, it indicates that model training for that type of training sample has met the required requirements and there is no need to increase the sample collection ratio. Alternatively, the collection ratio may be reduced to ensure that the training batch includes more samples that require a high collection ratio.
[0074] The preset loss threshold corresponding to each type of training sample can be set appropriately based on the circumstances, and this application does not limit this. For example, for types of content related to pixel changes, such as types containing images or color blocks, for example, such as document material types and identity card types, the preset loss threshold can be set slightly larger.
[0075] For example, taking the previous article as an example, "When training the model, two types of training sample sets are used (for example, training sample set A and training sample set B), a sample type collection ratio of 1:1 (that is, 50% each) can be set", and the scheduled number of rounds is 10 rounds, the preset loss threshold a can be set for the training samples of type A in advance, and the preset loss threshold b can be set for the training samples of type B in advance. After collecting 0.5N training samples from the training sample set A and the training sample set B at a sample type collection ratio of 1:1 (that is, 50% each) to form training batch samples, and performing 10 rounds of image enhancement model through these training batch samples: if the loss corresponding to the training samples of type A is calculated to be 0.5a and the loss corresponding to the training samples of type B is 1.5b, since 0.5a is less than the preset loss threshold a, it can be considered that the training of type A has achieved the training effect, and since 1.5b is greater than the preset loss threshold b, it is considered that the training of type B has not achieved the training effect. Therefore, before the next 10 rounds of model training, the collection ratio of type B training samples can be increased, while the collection ratio of type A training samples can be reduced, so as to collect training samples from training sample set A and training sample set B respectively to form training batch samples for the next 10 rounds of model training.
[0076] For another example, if the loss corresponding to the training samples of type A is calculated to be 1.2a and the loss corresponding to the training samples of type B is 1.5b, then before the next 10 rounds of model training, the collection ratio of the corresponding samples of one of type A and type B can be increased, and the collection ratio of the corresponding samples of the other type can be reduced. For example, taking the example of first increasing the collection ratio of type A and reducing the collection ratio of type B, the training batch samples for the next 10 rounds of model training are formed for training until the loss of the training samples of type A is less than a, and then the collection ratio of type B is increased, and training continues in the next 10 rounds of training.
[0077] Other situations can be deduced by analogy, and will not be described in detail. It should be understood that these are only some examples and are not intended to be limiting in this application.
[0078] For example, the collection ratio can be lowered and increased in steps of 3% to 10%. Taking a step size of 5% as an example, before the next 10 rounds of model training, the collection ratio of type B training samples is increased by 5%, and the collection ratio of type A training samples is reduced by 5%. Then the sample type collection ratio changes from 1:1 (that is, 50% each) to 9:11 (that is, 45% for type A and 55% for type B). For example, if a total of N training samples are required for training, 0.45N training samples can be extracted and collected from training sample set A, and then 0.55N training samples can be extracted and collected from training sample set B to generate training batch samples for the next 10 rounds of model training to train the image enhancement model. It should be understood that this is only for some examples and is not a limitation in this application.
[0079] The above-adjusted training process is iterated until a training termination condition is reached. For example, in some optional embodiments, the "until a training termination condition is reached" in step S106 includes: until the loss corresponding to each type of training sample meets a corresponding convergence threshold, where the convergence threshold is determined according to the type of training sample.
[0080] In this application, when the losses corresponding to each type of training sample meet their respective corresponding convergence thresholds, it can be said that the image enhancement model has good training effects on different types of training samples and there is no long-tail effect, so the training can be terminated. Among them, the convergence threshold can be set as needed, for example, it can be set to 0.01-0.1, etc., and this application does not make specific limitations. For example, for the type of document content, such as book type, work document type, etc., the convergence threshold can be slightly smaller, such as 0.01; and for types containing images or color blocks, such as identity card type, certificate material type, business license type, etc., the convergence threshold can be slightly larger, such as 0.1. However, the above is only an exemplary explanation. In actual applications, those skilled in the art can flexibly set it according to actual needs and sample types.
[0081] Based on this, the present application uses such training termination conditions to effectively determine whether the image enhancement model has been trained, so as to end the training and ensure that the trained image enhancement model can obtain better image enhancement effects for different types of scanned images.
[0082] FIG4 is a schematic diagram of an optional processing process of the above-mentioned data processing method. As shown in FIG4 , the exemplary image enhancement model is a machine learning model based on an encoder and decoder structure, wherein the encoder includes multiple encoding layers and multiple self-attention layers, wherein the encoding layers can be used for encoding processing, the self-attention layers can be used for self-attention processing, and the decoder includes multiple decoding layers, which can be used for decoding processing. First, in step S102, training samples are collected from multiple different types of training sample sets according to a set sample type collection ratio to obtain training batch samples. Exemplarily, a dataloader (a data loading module in the training framework) can be used to collect training samples according to the sample type collection ratio to obtain training batch data, which is then input into the image enhancement model to be trained. Then, in step S104, the image enhancement model is trained for a predetermined number of rounds using the training batch samples to obtain the losses corresponding to the different types of samples. For example, after the image enhancement model has undergone a predetermined round of training, the image enhancement model that has completed the preset round of training can be used to perform image enhancement processing on the test samples, thereby calculating the loss on different types of test samples. The test samples can be collected from different types of test sample sets, for example, they can be collected according to 1%, 3%, 5%, etc., and can be selected specifically as needed. It should be understood that the functions of test samples and training samples are different. They are not used for the predetermined round of model training, but only for determining the loss. Then, after step S106, the sample type collection ratio is adjusted based on the obtained loss, and training samples are collected from multiple different types of training sample sets according to the adjusted sample type collection ratio to form a new training batch sample, and the operation of using the training batch sample to train the image enhancement model is returned to continue until the training termination condition is reached, the model training is terminated, and the trained image enhancement model is obtained. It should be understood that the example shown in Figure 4 is only used to facilitate the understanding of the embodiment of the present application and does not constitute any limitation to the present application.
[0083] In summary, the solution provided in the embodiment of the present application is different from the traditional method of randomly collecting training samples to train the image enhancement model. The solution of the embodiment of the present application adopts a method of setting the sample type collection ratio, setting different collection ratios for different types of training samples. Thus, even if there are fewer training samples of certain types, according to the collection ratio of the sample type, a number of training samples that are consistent with its ratio will be extracted from its corresponding set, and combined with training samples of other types to form training batch samples for training the image enhancement model. In this way, even if the number of training samples of this type accounts for a small proportion of the total training samples, these fewer training samples can be collected according to the collection ratio to be combined with training samples of other types that account for a larger proportion, to form a large number of training batch samples that can meet the training needs. Thereby, the long-tail effect of image enhancement model training caused by the small number of training samples of this type is avoided, and the training effect of the image enhancement model is improved. Furthermore, after the image enhancement model is trained for a preset number of rounds using the current training batch samples, the training quality of a certain type of training sample will be evaluated based on the model's training effect, that is, the losses corresponding to different types of samples obtained through model training. The sample type collection ratio will then be adjusted based on the loss to fully solve the long-tail effect problem generated in traditional training methods and further improve the model training effect.
[0084] It can be understood that the above description of the data processing method for the image enhancement model of the first aspect is only an exemplary description of the present application and does not constitute any limitation to the present application.
[0085] According to a second aspect of the present application, a data processing method for an image enhancement model is provided. The data processing method is implemented based on the trained image enhancement model. Referring to the flowchart shown in FIG5 , the method includes steps S202 and S204, specifically:
[0086] S202: Acquire a scanned image to be processed.
[0087] S204: Performing image enhancement processing on the scanned image through an image enhancement model to obtain an enhanced scanned image.
[0088] The image enhancement model in step S204 is obtained by training based on any one of the methods provided in the first aspect.
[0089] Based on the data processing method for the image enhancement model provided in the second aspect of this application, since the scanned image can be enhanced by the image enhancement model to obtain an enhanced scanned image, and the image enhancement model is trained based on any one of the methods provided in the first aspect, the enhanced scanned image obtained by this solution has a better effect.
[0090] It should be noted that the data processing method for the image enhancement model provided in the second aspect of this application is a method for using the trained image enhancement model. By using such a method, the effect of image enhancement can be improved. It should still be understood that the image enhancement model in this application can be used to perform image enhancement processing on images, but its specific function is not limited in this application. To facilitate the description of the embodiments of this application, this application uses the application scenario of the image enhancement model used in scanning software to perform image enhancement processing on scanned images as an example.
[0091] In this application, the scanned image to be processed may include but is not limited to books, invoices, contracts, business licenses, work documents, certificate materials, identity cards, driver's licenses, etc. obtained by software scanning. The image enhancement model in this application can then be used to perform image enhancement processing on the scanned image to be processed to obtain an enhanced scanned image. For example, taking an invoice as an example, the scanned image can be obtained in the following optional manner: the user places the invoice to be scanned on a support (such as a desktop, etc.) or holds it in his hand, and then uses the camera on the mobile terminal to shoot the invoice from the direction of the invoice to obtain the corresponding scanned image. Of course, this is just an example and does not serve as any limitation to this application. It should be understood that in this case, the training sample can be obtained by shooting when needed and then directly obtained, or it can be shot in advance and then the training sample image is stored in a predetermined storage space (including but not limited to storage media such as disks, hard drives, memories, or databases, etc.), and directly obtained from the storage space when needed. This application does not impose any restrictions on this.
[0092] In another type, the scanned image to be processed may also be a document image that has undergone preliminary model processing, and can be input into the image enhancement model in this application for further image enhancement processing to obtain an enhanced scanned image.
[0093] In another type, the scanned image to be processed can also be obtained through a pre-generated electronic file, for example, it can be obtained by printing or scanning with software. The image enhancement model in this application can then be used to perform image enhancement processing on the scanned image to be processed to obtain an enhanced scanned image. The electronic file can be in PDF format, or it can be in another file format. The electronic file can include one or more pages, or a scanned image can be generated based on these pages. Similarly, the scanned image generated in this way can also be first stored in a predetermined storage space (including but not limited to storage media such as disks, hard disks, memories, or databases, etc.), and directly retrieved from the storage space when needed. This application does not impose any restrictions on this.
[0094] Figure 6 is an optional scenario diagram of the data processing method in the second aspect of the present application. With reference to Figure 6, it can be understood that the content of the scanned image to be processed in this example contains text, color blocks, charts and seals (or at least one of them), and there are also some impurity defects (noise data). These impurity defects make the scanned image effect poor. The scanned image to be processed is input into the image enhancement model trained by the method of the first aspect. After image enhancement processing by the image enhancement model, the enhanced scanned image is output. It can be seen from the figure that the enhanced scanned image is compared with the original scanned image to be processed. The impurity defects have been removed and the effect of the scanned image has been significantly enhanced. However, it is not limited to this and can also include other enhancement effects, such as document enhancement, background removal, etc., which are not illustrated one by one here. It should be understood that the scenario example shown in Figure 6 is only used to facilitate the understanding of the embodiment of the present application and does not serve as any limitation to the present application.
[0095] It can be understood that the above description of the data processing method for the image enhancement model of the second aspect is merely an exemplary description of the present application and does not constitute any limitation to the present application.
[0096] According to the third aspect of an embodiment of the present application, an electronic device is provided, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the method described in the first aspect or the second aspect.
[0097] 7 , a schematic structural diagram of an electronic device according to an embodiment of the present application is shown. The specific embodiment of the present application does not limit the specific implementation of the electronic device.
[0098] As shown in FIG. 7 , the electronic device 700 may include a processor 702 , a communications interface 704 , a memory 706 , and a communication bus 708 .
[0099] in:
[0100] The processor 702 , the communication interface 704 , and the memory 706 communicate with each other via a communication bus 708 .
[0101] The communication interface 704 is used to communicate with other electronic devices or servers.
[0102] The processor 702 is used to execute the program 710, and specifically can execute the relevant steps in the embodiment of the data processing method for the image enhancement model.
[0103] Specifically, the program 710 may include program codes, which include computer operation instructions.
[0104] Processor 702 may be a CPU, a Graphics Processing Unit (GPU), an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs, or may be processors of different types, such as one or more CPUs and one or more ASICs.
[0105] The memory 706 is used to store the program 710. The memory 706 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0106] The program 710 may include multiple computer instructions. Specifically, the program 710 may enable the processor 702 to execute operations corresponding to the data processing method for the image enhancement model described in any of the aforementioned method embodiments through the multiple computer instructions.
[0107] The specific implementation of each step in program 710 can refer to the corresponding description of the corresponding steps and units in the above-mentioned method embodiment, and has corresponding beneficial effects, which will not be repeated here. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working process of the above-mentioned devices and modules can refer to the corresponding process description in the above-mentioned method embodiment, and will not be repeated here.
[0108] According to a third aspect of the embodiments of the present application, the embodiments of the present application further provide a computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any of the aforementioned method embodiments. The computer storage medium includes, but is not limited to, a compact disc read-only memory (CD-ROM), random access memory (RAM), a floppy disk, a hard disk, or a magneto-optical disk.
[0109] An embodiment of the present application also provides a computer program product, including computer instructions, which instruct a computing device to perform operations corresponding to any data processing method for an image enhancement model in the multiple method embodiments in the first aspect or the second aspect above.
[0110] The electronic device 700 / computer storage medium / computer program product embodiment in the embodiment of the present application has been described in detail in the aforementioned data processing method embodiment for the image enhancement model. Therefore, its relevant content and beneficial effects can be understood with reference to the above-mentioned method embodiment and will not be repeated here.
[0111] In addition, it should be noted that the user-related information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to sample data used to train the model, data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0112] It should be pointed out that, according to the needs of implementation, the various components / steps described in the embodiments of the present application can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present application.
[0113] The above-mentioned method according to the embodiment of the present application can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk or magneto-optical disk), or as computer code that is originally stored in a remote recording medium or a non-temporary machine-readable medium downloaded via a network and will be stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor or programmable or dedicated hardware (such as an application-specific integrated circuit (ASIC) or a field programmable gate array (FPGA)). It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component (e.g., random access memory (RAM), read-only memory (ROM), flash memory, etc.) that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown here, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown here.
[0114] Those skilled in the art will appreciate that the units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for specific applications, but such implementation should not be considered to be beyond the scope of the embodiments of this application.
[0115] The above implementation methods are only used to illustrate the embodiments of the present application, and are not intended to limit the embodiments of the present application. Ordinary technicians in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of the present application, and the scope of patent protection of the embodiments of the present application should be defined by the claims.
Claims
1. A data processing method for an image enhancement model, comprising: According to the set sample type collection ratio, training sample collection is performed on multiple different types of training sample sets to obtain training batch samples; The image enhancement model is trained using the training batch samples, and after completing a preset round of training, the losses corresponding to different types of samples are obtained; The sample type collection ratio is adjusted based on the loss, and training samples are re-collected from the multiple different types of training sample sets according to the adjusted sample type collection ratio to form new training batch samples, and the operation of training the image enhancement model using the training batch samples is returned to continue execution until the training termination condition is reached.
2. The method according to claim 1, wherein: The losses corresponding to different types of samples are obtained, including: Obtain test samples; The image enhancement model after completing the preset rounds of training is used to perform image enhancement processing on the test sample, and the losses corresponding to different types of test samples are obtained according to the processing results.
3. The method according to claim 2, wherein: The losses corresponding to different types of test samples are obtained according to the processing results, including: According to the processing results and a preset loss function, the losses corresponding to different types of test samples are obtained; wherein the loss function includes: a plurality of loss calculation parts reflecting the losses corresponding to each type of test sample.
4. The method according to any one of claims 2 to 3, wherein: The step of reaching the training termination condition includes: until the loss corresponding to each type of test sample satisfies the corresponding convergence threshold, wherein the convergence threshold is determined according to the type of training sample.
5. The method according to any one of claims 1 to 4, wherein: The plurality of different types of training sample sets include: a plurality of different types of training sample sets divided into corresponding sets according to different types of content contained in the image.
6. The method according to any one of claims 1 to 5, wherein: The image enhancement model is a machine learning model based on an encoder and decoder structure, wherein the encoder includes an encoding layer and an attention layer; The training of the image enhancement model using the training batch samples includes: for each sample in the training batch samples, encoding and attention processing of the sample based on the encoding layer and attention layer in the encoder The image enhancement model is trained by the intention processing and the decoding processing of the decoder.
7. The method according to claim 6, wherein: The encoder comprises a plurality of encoding layers and a self-attention layer connected after each encoding layer; The encoding and attention processing performed on the sample by the encoding layer and the attention layer in the encoder includes: sequentially performing encoding and self-attention processing on the sample through each encoding layer in the encoder and a self-attention layer connected to each encoding layer.
8. The method according to claim 6, wherein: The encoder comprises a plurality of encoding layers, and some of the encoding layers in the plurality of encoding layers are connected to a self-attention layer; The encoding processing and attention processing performed on the sample by the encoding layer and the self-attention layer in the encoder include: inputting the sample into the encoder, and performing encoding processing and self-attention processing on the sample in sequence according to the connection relationship between the encoding layer and the self-attention layer in the encoder.
9. The method according to any one of claims 1 to 8, wherein: The adjusting the sample type collection ratio based on the loss includes: For each type of training sample, if the corresponding loss is greater than the preset loss threshold, the collection ratio of training samples of this type is increased, and the collection ratios of training samples of other types are reduced.
10. A data processing method for an image enhancement model, comprising: Acquire a scanned image to be processed; Performing image enhancement processing on the scanned image through an image enhancement model to obtain an enhanced scanned image; The image enhancement model is obtained by training based on the method described in any one of claims 1 to 9.
11. An electronic device, comprising: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the method according to any one of claims 1 to 10.
12. A computer storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.
13. A computer program product, comprising computer instructions, wherein the computer instructions instruct a computing device to execute operations corresponding to the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Model training method and device and electronic equipment
CN110717515A
Image model training method, electronic equipment, roadside equipment and cloud control platform
CN113420792A
Data processing method for image enhancement model, electronic equipment and storage medium
CN117746121A
Method and apparatus for training image model, and method and apparatus for category prediction
US20190286940A1
Neural network training method and apparatus, image processing method and apparatus, and device and storage medium
WO2023040629A1