Image enhancement method and device, storage medium and electronic equipment
By selecting large models and image enhancement large models for training areas, the problem that global processing in the prior art cannot meet local needs is solved, and efficient and targeted image enhancement effect is achieved.
Patent Information
- Application Number
- CN202411814199.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-06-06
AI Technical Summary
The image enhancement methods in the prior art are usually based on global processing, and cannot meet the targeted needs of users for local images, while increasing unnecessary computing volume.
The local image indicated by the user is found by the training area selection model, and the local image is enhanced by the image enhancement model, which improves targetedness and reduces the amount of calculation.
Efficient enhancement processing of local images is realized, targeted and efficient image enhancement is improved, and the calculation amount is reduced.
Smart Images

Figure CN120107111A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and more specifically, to an image enhancement method, device, storage medium and electronic device in the field of computer technology. Background Art
[0002] Image enhancement technology is a technology that improves image quality through various algorithms and processing methods. Its main purpose is to make the image clearer or more in line with specific visual requirements. Image enhancement technology can help users in daily life enhance images with blur, fog, etc. to obtain clearer images. However, image enhancement in existing technologies is often based on global processing, directly enhancing the entire image, which cannot meet users' targeted needs for local areas and also increases the unnecessary amount of computation for image enhancement processing. It is necessary to provide an image enhancement method with stronger targeting and less computation. Summary of the invention
[0003] The embodiments of the present application provide an image enhancement method, device, storage medium and electronic device. The method can find the local image indicated by the user through a trained region selection model, and then use the image enhancement model to enhance the local image. The processing of the local image improves the specificity of the image enhancement and reduces the amount of image processing calculations.
[0004] In a first aspect, an embodiment of the present application provides an image enhancement method, the method comprising:
[0005] Acquire a first sample data set, where the first sample data set includes a first sample image, a first sample text, and sample area coordinates obtained by performing a frame selection process on the first sample image based on the first sample text;
[0006] Performing model training on the initial region selection large model based on the first sample data set to obtain the region selection large model;
[0007] Acquire a second sample data set, where the second sample data set includes a second sample image, a second sample text, and a sample clear image obtained by enhancing the second sample image based on the second sample text;
[0008] Performing model training on the initial image enhancement large model based on the second sample data set to obtain the image enhancement large model;
[0009] Image enhancement processing is performed based on the region selection large model and the image enhancement large model.
[0010] In a second aspect, an embodiment of the present application provides an image enhancement device, the device comprising:
[0011] A first sample acquisition unit is used to acquire a first sample data set, where the first sample data set includes a first sample image, a first sample text, and sample area coordinates obtained by performing a frame selection process on the first sample image based on the first sample text;
[0012] A region selection model training unit, used for performing model training on an initial region selection large model based on the first sample data set to obtain a region selection large model;
[0013] A second sample acquisition unit is used to acquire a second sample data set, where the second sample data set includes a second sample image, a second sample text, and a sample clear image obtained by enhancing the second sample image based on the second sample text;
[0014] An image enhancement model training unit, used for performing model training on an initial image enhancement large model based on the second sample data set to obtain an image enhancement large model;
[0015] An image enhancement processing unit is used to perform image enhancement processing based on the region selection large model and the image enhancement large model.
[0016] In a third aspect, an embodiment of the present application provides a computer storage medium, wherein the computer storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the above-mentioned method steps.
[0017] In a fourth aspect, an embodiment of the present application provides an electronic device, which may include: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the above-mentioned method steps.
[0018] In one or more embodiments of the present application, a first sample data set is obtained, wherein the first sample data set includes a first sample image, a first sample text, and sample region coordinates obtained by selecting the first sample image based on the first sample text, a large initial region selection model is trained based on the first sample data set to obtain a large region selection model, a second sample data set is obtained, wherein the second sample data set includes a second sample image, a second sample text, and a sample clear image obtained by enhancing the second sample image based on the second sample text, a large initial image enhancement model is trained based on the second sample data set to obtain an image enhancement model, and image enhancement processing is performed based on the large region selection model and the image enhancement model. The local image indicated by the user is found through the trained large region selection model, and then the large image enhancement model is used to enhance the local image, and processing the local image improves the pertinence of image enhancement and reduces the amount of image processing calculations. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0020] Figure 1 is an example schematic diagram of an image enhancement process provided in an embodiment of the present application;
[0021] Figure 2 is a flowchart of an image enhancement method provided in an embodiment of the present application;
[0022] Figure 3 is a flowchart of an image enhancement method provided in an embodiment of the present application;
[0023] Figure 4 It is a flowchart of a large model training for region selection provided in an embodiment of the present application;
[0024] Figure 5 It is a flowchart of a large image enhancement model training provided in an embodiment of the present application;
[0025] Figure 6 is a flowchart of an image enhancement method provided in an embodiment of the present application;
[0026] Figure 7 is a structural schematic diagram of an image enhancement device provided in an embodiment of the present application;
[0027] Figure 8 is a structural schematic diagram of an image enhancement device provided in an embodiment of the present application;
[0028] Fig. 9 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0030] Due to the influence of environmental factors such as surrounding lighting and weather, or the shaking of the shooting machine, the acquired image quality may be poor and unclear, which cannot meet the user's expectations or usage standards. For example, the image taken in rainy weather may be dim and blurred due to low light conditions, or the image taken in a fast moving situation may cause the object in the picture to be blurred, or the car exhaust may cause the local image to be blurred and the details to be difficult to recognize when shooting a car. In addition, due to the loss or damage of some data during the image transmission process, the state of partial or complete unclearness may also occur. In order to make the picture clearer and the details in the picture more complete, the user can use image enhancement technology to enhance the image. The image enhancement technology in the prior art is to enhance the overall image. For example, the image obtained by the user has fog and causes the image to be blurred. The user expects to enhance the image by removing fog. In the prior art, the entire image will be enhanced by removing fog, but not all parts of the image may have fog, or the user does not need to remove all fog in the image. The prior art cannot meet the user's demand for enhancing the local image. The enhancement of the entire image will also reduce the efficiency of image enhancement and increase the amount of calculation.
[0031] The image enhancement method provided in the embodiment of the present application can be implemented by a computer program and can be run on an image enhancement device based on the von Neumann system. The computer program can be integrated in an application or run as an independent tool application. The image enhancement device can obtain a large model for region selection and a large model for image enhancement through model training. It can be understood that the large model for region selection and the large model for image enhancement can be large models, which are deep learning models with a large number of parameters and complex structures, and have the ability to process tasks such as natural language processing and computer vision. When the user needs to enhance the target image, the target image and the target instruction text can be input to the image enhancement device. The target instruction text is a text containing the user's enhanced processing requirements for the target image. The image enhancement device can use the large model for region selection and the large model for image enhancement to enhance the target image based on the target instruction text, thereby obtaining the enhanced target image, wherein the large model for region selection is used to determine the target local image that the user needs to enhance from the target image, and the large model for image enhancement is used to enhance the target local image.
[0032] Please also see Figure 1, an example schematic diagram of image enhancement processing is provided for an embodiment of the present application. The image enhancement device may include a large region selection model and an image enhancement model, or the image enhancement device may control and call the large region selection model and the image enhancement model. When the image enhancement device obtains the target image and the target instruction text input by the user, the target image and the target instruction text may be first input into the large region selection model. The large region selection model may extract the target local image from the target instruction text. For example, the target image may be a photo of a large tree, but due to shooting or data transmission errors, it is blurred. The target instruction text input by the user may be "deblur and enhance the tree in the image", and the target local image may be a local image containing the tree, for example, it may be a bounding box containing the tree in the target image. Then the image enhancement device may input the target local image and the target instruction text into the large image enhancement model. Since the user expects to perform deblur and enhance processing in the target instruction text, the large image enhancement model may perform deblur and enhance processing on the target local image based on the target instruction text to obtain a target enhanced image. The target enhanced image is a clear target local image that has completed deblur and enhance processing. Then the image enhancement device can replace the target local image in the target image with the target enhanced image to obtain the enhanced target image. The image enhancement device does not perform deblurring enhancement processing on the entire target image, but deblurring enhancement processing on the local image containing the "tree" desired by the user, which improves the pertinence of image enhancement, saves the amount of image enhancement calculation, and improves the efficiency of image enhancement.
[0033] The image enhancement method provided by the present application is described in detail below with reference to specific embodiments.
[0034] See also Figure 2 , which is a flow chart of an image enhancement method provided in an embodiment of the present application. Figure 2 As shown, the method of the embodiment of the present application may include the following steps S101-S103.
[0035] S101, obtaining a first sample data set, wherein the first sample data set includes a first sample image, a first sample text, and sample area coordinates obtained by selecting the first sample image based on the first sample text.
[0036] Specifically, the image enhancement device can obtain the first sample image and the first sample text. The first sample text can be used to select a local image in the first sample image, so relevant professionals can use manual annotation to select the first sample image based on the first sample text to obtain the sample area coordinates. The sample area coordinates are used to indicate the position of the local image. For example, if the local image is a rectangular image, the sample area coordinates may include the coordinates of the four vertices of the local image in the sample image. The image enhancement device can generate a first sample data set based on the first sample image, the first sample text, and the corresponding sample area coordinates.
[0037] S102: Perform model training on the initial region selection large model based on the first sample data set to obtain the region selection large model.
[0038] Specifically, the image enhancement device can create an initial region selection large model for model training. The initial region selection large model can be a target detection model pre-trained on a large-scale data set. The initial region selection large model can have the function of performing frame selection processing on the image according to the text content. In order to enhance the frame selection processing capability of the initial region selection large model, the image enhancement device can perform model training on the initial region selection large model based on the first sample data set, perform back propagation and iterative optimization during the model training to complete the fine-tuning of the initial region selection large model, so that the initial region selection large model is more focused on frame selection processing and improves the accuracy of frame selection processing, and finally obtains the region selection large model.
[0039] S103, obtaining a second sample data set, where the second sample data set includes a second sample image, a second sample text, and a sample clear image obtained by enhancing the second sample image based on the second sample text.
[0040] Specifically, the image enhancement device can obtain a second sample image and a second sample text. The second sample text can be used to indicate that the second sample image is enhanced. It can be understood that the enhancement process can include different image enhancement methods, such as deblurring, defogging, removing water droplets, improving brightness or improving sharpness, etc. The second sample text can also be used to indicate the image enhancement method for the second sample image. The second sample image can be an image obtained from the Internet, etc., and can include a local image obtained by selecting the first sample image through a large area selection model.
[0041] Relevant professionals can enhance the second sample image based on the second sample text to generate a sample clear image, and the image enhancement device can generate a second sample data set based on the second sample image, the second sample text and the corresponding sample clear image.
[0042] S104: Perform model training on the initial image enhancement large model based on the second sample data set to obtain the image enhancement large model.
[0043] Specifically, the image enhancement device can create an initial image enhancement large model, which has the function of enhancing the image according to the text content. In order to strengthen the enhancement processing capability of the initial image enhancement large model, the image enhancement device can perform model training on the initial image enhancement large model based on the second sample data set, and perform back propagation and iterative optimization during the model training, so that the initial image enhancement large model is more focused on enhancement processing and the enhanced image is more in line with user needs, and finally the image enhancement large model is obtained.
[0044] S105, performing image enhancement processing based on the region selection large model and the image enhancement large model.
[0045] Specifically, the image enhancement device can perform image enhancement processing based on the region selection large model and the image enhancement large model, wherein the region selection large model is used to determine the local image from the image, and the image enhancement large model is used to perform enhancement processing on the local image, so as to improve the targetedness of image enhancement, save the amount of image enhancement calculation, and improve the efficiency of image enhancement.
[0046] In an embodiment of the present application, a first sample data set is obtained, the first sample data set includes a first sample image, a first sample text, and sample area coordinates obtained by selecting the first sample image based on the first sample text, the initial area selection model is trained based on the first sample data set to obtain the area selection model, a second sample data set is obtained, the second sample data set includes a second sample image, a second sample text, and a sample clear image obtained by enhancing the second sample image based on the second sample text, the initial image enhancement model is trained based on the second sample data set to obtain the image enhancement model, and image enhancement is performed based on the area selection model and the image enhancement model. The local image indicated by the user is found through the trained area selection model, and the image enhancement model is used to enhance the local image. The processing of the local image improves the pertinence of the image enhancement and reduces the amount of image enhancement calculations.
[0047] See also Figure 3 , which is a flow chart of an image enhancement method provided in an embodiment of the present application, such as Figure 3 As shown, in one or more embodiments of the present application, the following steps S201-S203 may be included before step S101.
[0048] S201, obtaining sample images and sample texts.
[0049] Specifically, the image enhancement device can obtain sample images from the Internet, etc., and obtain sample text for the sample images. There is a one-to-one correspondence between the sample text and the sample image. The sample text can be obtained from the Internet, etc., or generated by relevant professionals or the image enhancement device.
[0050] Among them, the position indicator words are used to indicate the local image in the sample image, so that the relevant professionals and the regional selection model can perform frame selection processing on the sample image later. The position indicator words can be directional words or objects in the sample image. For example, the position indicator words can be directional words such as "upper left", "middle", "bottom" etc. used to indicate the position of the local image in the sample image, or they can be objects existing in the sample image such as "tree", "car", "person". It can be understood that in addition to indicating the local image, the position indicator words can also indicate the entire sample image. When the position indicator words are words such as "full image", it can indicate that the user expects to enhance the entire sample image. At this time, the local image is equal to the sample image.
[0051] The enhancement indicating words are used to indicate the image enhancement method for enhancing the sample image, for example, they may be "deblur", "defogging", "improving sharpness", etc. The sample text contains at least one of the position indicating words and the enhancement indicating words. For example, the sample text may contain only the position indicating words such as "select the upper left corner of the image", or only the enhancement indicating words such as "perform deblurring and enhancing processing", or may contain both the position indicating words and the enhancement indicating words, for example, "perform deblurring and enhancing processing on the upper left corner of the image".
[0052] S202: If there are position indicating words in the sample text, a frame selection process is performed on the sample image based on the position indicating words to obtain sample area coordinates.
[0053] Specifically, if there are position indicating words in the sample text, relevant professionals can perform frame selection processing on the sample image based on the position indicating words to obtain corresponding sample area coordinates. The sample area coordinates are the area coordinates corresponding to the sample image. The area coordinates are used to indicate the position of the local image indicated by the position indicating words in the image. If the local image is a rectangular image, the sample area coordinates can be the coordinates of the four vertices of the local image indicated by the position indicating words in the sample image.
[0054] Optionally, the area coordinates may be the coordinates of four vertices of a local image arranged in a preset order. The preset order may be an initial setting of the image enhancement device, or may be configured by a user or relevant staff. For example, the preset order may be the upper left corner coordinate, the upper right corner coordinate, the lower right corner coordinate, and the lower left corner coordinate.
[0055] S203: If there are enhancement indicating words in the sample text, the sample image is enhanced based on the enhancement indicating words to obtain a sample clear image.
[0056] Specifically, if there are enhancement indication words in the sample text, then the relevant professionals or image enhancement devices can enhance the sample image based on the enhancement indication words to obtain a sample clear image, and the sample clear image is the clear image after the enhancement processing. The enhancement indication words can indicate the image enhancement method of the enhancement processing, so the relevant professionals or image enhancement devices can enhance the sample image based on the image enhancement method corresponding to the enhancement indication words.
[0057] Optionally, if both position indicating words and enhancement indicating words exist in the sample text, the sample image can be first framed based on the position indicating words to obtain the sample area coordinates, and then a sample local image can be obtained from the sample image based on the sample area coordinates, and then the sample local image can be enhanced based on the enhancement indicating words to obtain a clear sample image.
[0058] Optionally, the image enhancement device can generate a total sample data set based on the sample image, sample text, sample area coordinates and sample clear image. When the initial area selection large model and the initial image enhancement large model need to be trained subsequently, the first sample data set and the second sample data set can be obtained from the total sample data set.
[0059] In the embodiment of the present application, a sample image and a sample text are obtained. If there are position indicating words in the sample text, the sample image is framed based on the position indicating words to obtain the sample area coordinates. If there are enhancement indicating words in the sample text, the sample image is enhanced based on the enhancement indicating words to obtain a sample clear image. By processing the sample image according to the indicating words in the sample text, the subsequent model can learn the ability to process images based on text, thereby improving the accuracy of image enhancement.
[0060] In one or more embodiments of the present application, step S101 may include the following steps:
[0061] The total sample data set is screened to obtain a first sample data set.
[0062] Specifically, the image enhancement device can obtain sample text containing position indicating words in the total sample data set and determine it as the first sample text, then obtain a sample image corresponding to the first sample text in the total sample data set and determine it as the first sample image, and can also obtain sample area coordinates corresponding to the first sample image, and then generate a first sample data set based on the first sample text, the first sample image and the sample area coordinates.
[0063] See also Figure 4 , a schematic diagram of a process of training a large model for selecting a region is provided for the embodiment of the present application, such as Figure 4 As shown, in one or more embodiments of the present application, step S101 may further include the following steps:
[0064] S301, creating an initial region selection model.
[0065] Specifically, the image enhancement device can create an initial region selection large model, which has the function of frame selection processing on the image according to the text. The initial region selection large model can be a target detection model pre-trained on a large-scale data set.
[0066] S302, inputting the first sample image and the first sample text into the initial region selection large model, and obtaining the training region coordinates obtained by the initial region selection large model by performing frame selection processing on the first sample image based on the first sample text.
[0067] Specifically, the image enhancement device can input the first sample image and the first sample text into the initial area selection large model. The initial area selection large model can perform a frame selection process on the first sample image based on the first sample text and obtain training area coordinates. The training area coordinates represent the position of the local image indicated by the first sample text obtained by the initial area selection large model in the first sample image.
[0068] Optionally, the image enhancement device may input the first sample image and the first sample text into the initial area selection large model. Since the first sample text may contain position indicating words, the initial area selection large model may obtain the position indicating words from the first sample text, and perform frame selection processing on the first sample image based on the position indicating words and generate training area coordinates.
[0069] S303, based on the training area coordinates and the sample area coordinates, the initial area selection large model is adjusted in parameters until the model training is completed to obtain the area selection large model.
[0070] Specifically, the image enhancement device can calculate the first loss function of the initial region selection large model based on the training area coordinates and the sample area coordinates, and perform parameter adjustment processing on the initial region selection large model during the back propagation training process based on the first loss function until the initial region selection large model completes model training to obtain the region selection large model.
[0071] Optionally, the image enhancement device may use supervised fine-tuning (SFT) technology and direct preference optimization (DPO) technology during the model training process. The SFT technology fine-tunes the initial region selection model on the first sample data set to reduce the model's bias and thus improve the accuracy of the frame selection processing function. The DPO technology helps reduce the bias of the region selection model when generating content, making the output more reasonable.
[0072] Optionally, the first loss function may be an intersection over union (IoU) loss function, which may calculate the ratio of the intersection area between two regions, namely, the region corresponding to the training region coordinates and the region corresponding to the sample region coordinates, to their union area. The calculation formula is as follows:
[0073]
[0074] Among them, Loss IoU is the first loss function, Box low_crop is the coordinate of the training area, Box gt_crop are the sample area coordinates.
[0075] In an embodiment of the present application, model training of the region selection large model is completed based on the first sample data set and the intersection-over-union loss function, which improves the accuracy of the region selection large model in object positioning, thereby enabling the region selection large model to further improve the accuracy of frame selection processing.
[0076] In one or more embodiments of the present application, step S103 may include the following steps:
[0077] The total sample data set is screened to obtain a second sample data set.
[0078] Specifically, the image enhancement device can obtain sample text containing enhancement indication words in the total sample data and determine it as the second sample text, and then obtain a sample image corresponding to the second sample text in the total sample data set and determine it as the second sample image, and can also obtain a sample clear image corresponding to the second sample image, and then generate a second sample data set based on the second sample text, the second sample image and the sample clear image.
[0079] Optionally, the second sample text may be the first sample text. If the first sample text contains both position indicating words and enhancement indicating words, the first sample text may also be used as the second sample text. In this case, the image enhancement device may input the second sample text and the first sample image corresponding to the second sample text into the region selection large model, and obtain the retrained local image obtained by the region selection large model performing frame selection processing on the first sample image, and then use the retrained local image as the second sample image corresponding to the second sample text. Moreover, the region selection large model may continue to be parameter adjusted based on the retrained local image and the sample local image corresponding to the first sample image, so that the region selection large model may continue to be optimized during the process of model training of the initial image enhancement large model.
[0080] See also Figure 5 , provides a flow chart of a large image enhancement model training process for an embodiment of the present application, such as Figure 5 As shown, in one or more embodiments of the present application, step S104 may further include the following steps:
[0081] S401, creating an initial image enhancement model.
[0082] Specifically, the image enhancement device can create an initial image enhancement large model, and the initial image enhancement large model has the function of enhancing the image according to the text.
[0083] Optionally, the image enhancement device can create an initial image enhancement model based on a mixture of experts (MoE) architecture, and the initial image enhancement model can include at least one expert layer and each expert layer corresponds to a different image enhancement method. In actual use, the final image enhancement model only needs to activate the expert layer corresponding to the image enhancement method required by the user to enhance the image, thereby reducing the amount of model calculation and improving the efficiency of image enhancement. Since different expert layers can focus more on the corresponding image enhancement method during the training process, the quality of the enhanced image is also improved.
[0084] S402, inputting the second sample image and the second sample text into the initial image enhancement large model, and obtaining a training clear image obtained by the initial image enhancement large model by enhancing the second sample image based on the second sample text.
[0085] Specifically, the image enhancement device can input the second sample image and the second sample text into the initial image enhancement large model. The initial image enhancement large model can enhance the second sample image based on the second sample text and obtain a training clear image. The training clear image is the enhanced sample image obtained by the initial image enhancement large model.
[0086] Optionally, the image enhancement device can input the second sample image and the second sample text into the initial image enhancement large model. Since the second sample text may contain enhancement indicator words, the initial image enhancement large model can obtain the enhancement indicator words from the second sample text and determine the training image enhancement method corresponding to the enhancement indicator words. The training image enhancement method is the image enhancement method determined by the initial image enhancement large model based on the enhancement indicator words. Then the image enhancement device can control the initial image enhancement large model to adopt the target expert layer corresponding to the training image enhancement method to enhance the second sample image to obtain a clear training image, thereby improving the focus and accuracy of the target expert layer on the training image enhancement method.
[0087] S403, based on the training clear image and the sample clear image, the parameters of the initial image enhancement large model are adjusted until the model training is completed to obtain the image enhancement large model.
[0088] Specifically, the image enhancement device can calculate the second loss function of the initial image enhancement large model based on the training clear image and the sample clear image, and adjust the parameters of the initial image enhancement large model based on the second loss function during the back propagation training process until the initial image enhancement large model completes the model training to obtain the image enhancement large model.
[0089] Optionally, the second loss function may be a mean squared error (MSE) loss function, and the calculation formula is as follows:
[0090]
[0091] Among them, Loss MSE is the second loss function, W is the width of the training clear image and the sample clear image, H is the height of the training clear image and the sample clear image, I high_crop To train clear images, I high_gt A clear image of the sample.
[0092] In an embodiment of the present application, a large image enhancement model is created based on the MoE architecture, and each expert layer focuses on a different image enhancement method, so that the large image enhancement model only needs to activate the corresponding expert layer in actual application, thereby improving the accuracy of the enhancement processing of the corresponding image enhancement method and reducing the amount of image enhancement calculations.
[0093] See also Figure 6 , which is a flow chart of an image enhancement method provided in an embodiment of the present application, such as Figure 6 As shown, in one or more embodiments of the present application, step S105 may further include the following steps:
[0094] S501, obtaining a target image and a target instruction text input by a user.
[0095] Specifically, when the user wants to perform image enhancement processing on the target image, the target image and the target instruction text for instructing the image enhancement processing can be input into the image enhancement device, and the image enhancement device can obtain the target image and target instruction text input by the user, wherein the target instruction text can include position indication words and enhancement indication words.
[0096] S502, input the target image and the target instruction text into the region selection large model, and obtain the target region coordinates output by the region selection large model.
[0097] Specifically, the image enhancement device can input the target image and the target instruction text into the area selection large model. The area selection large model will obtain the position indication words in the target instruction text and perform a frame selection process on the target image based on the position indication words. The image enhancement device can then obtain the target area coordinates output by the area selection large model.
[0098] S503: Acquire a target partial image in the target image based on the target region coordinates.
[0099] Specifically, the image enhancement device may acquire a target partial image in the target image based on the target region coordinates.
[0100] Optionally, after the large region selection model generates the target region coordinates based on the target instruction text and the target image, the target local image can also be obtained based on the target region coordinates, and then the image enhancement device can directly obtain the target local image output by the large region selection model.
[0101] S504, inputting the target local image and the target instruction text into the image enhancement large model, and obtaining the target enhanced image output by the image enhancement large model.
[0102] Specifically, the image enhancement device can input the obtained target local image and target instruction text into the image enhancement large model. The image enhancement large model can obtain the enhancement instruction words in the target instruction text and determine the target image enhancement method corresponding to the target instruction text based on the enhancement instruction words. Then, the target expert layer corresponding to the target enhancement method is used to enhance the target local image to obtain the target enhanced image. The image enhancement device can then obtain the target enhanced image output by the image enhancement large model.
[0103] S505, replacing the target partial image in the target image with the target enhanced image to obtain the target image after image enhancement processing.
[0104] Specifically, the image enhancement device may replace the target local image in the target image with the target enhanced image, thereby obtaining the target image after image enhancement processing, thereby realizing local enhancement processing for the target image.
[0105] In an embodiment of the present application, the target enhanced image replaces the target local image in the target image to obtain the target image after image enhancement processing, which improves the accuracy of the local enhancement processing of the target image, further avoids unnecessary processing of the remaining parts of the target image, reduces the amount of calculation, and improves the image enhancement efficiency.
[0106] The following will be combined with the attached Figure 7 -Attached Figure 8 , the image enhancement device provided in the embodiment of the present application is introduced in detail. It should be noted that the attached Figure 7 -Attached Figure 8 The image enhancement device is used to execute the present application Figure 1-Figure 6 For the convenience of explanation, only the part related to the embodiment of the present application is shown. For the specific technical details not disclosed, please refer to the present application. Figure 1-Figure 6 The embodiment shown.
[0107] See also Figure 7 , which shows a schematic diagram of the structure of an image enhancement device provided by an exemplary embodiment of the present application. The image enhancement device can be implemented as all or part of the device through software, hardware or a combination of both. The device 1 includes a first sample acquisition unit 11, a region selection model training unit 12, a second sample acquisition unit 13, an image enhancement model training unit 14 and a picture enhancement processing unit 15.
[0108] A first sample acquisition unit 11 is used to acquire a first sample data set, where the first sample data set includes a first sample image, a first sample text, and sample region coordinates obtained by selecting the first sample image based on the first sample text;
[0109] A region selection model training unit 12 is used to perform model training on an initial region selection large model based on the first sample data set to obtain a region selection large model;
[0110] A second sample acquisition unit 13 is used to acquire a second sample data set, where the second sample data set includes a second sample image, a second sample text, and a sample clear image obtained by enhancing the second sample image based on the second sample text;
[0111] An image enhancement model training unit 14 is used to perform model training on the initial image enhancement large model based on the second sample data set to obtain the image enhancement large model;
[0112] The image enhancement processing unit 15 is used to perform image enhancement processing based on the region selection large model and the image enhancement large model.
[0113] In this embodiment, a first sample data set is obtained, wherein the first sample data set includes a first sample image, a first sample text, and sample region coordinates obtained by selecting the first sample image based on the first sample text, and the initial region selection model is trained based on the first sample data set to obtain the region selection model, and a second sample data set is obtained, wherein the second sample data set includes a second sample image, a second sample text, and a sample clear image obtained by enhancing the second sample image based on the second sample text, and the initial image enhancement model is trained based on the second sample data set to obtain the image enhancement model, and image enhancement processing is performed based on the region selection model and the image enhancement model. The local image indicated by the user is found through the trained region selection model, and then the image enhancement model is used to enhance the local image, and processing the local image improves the pertinence of image enhancement and reduces the amount of image processing calculations.
[0114] See also Figure 8 , which shows a schematic diagram of the structure of an image enhancement device provided by an exemplary embodiment of the present application. The image enhancement device can be implemented as all or part of the device through software, hardware or a combination of both. The device 1 includes a sample set acquisition unit 16, a first sample acquisition unit 11, a region selection model training unit 12, a second sample acquisition unit 13, an image enhancement model training unit 14 and a picture enhancement processing unit 15.
[0115] A sample collection acquisition unit 16, configured to acquire sample images and sample texts, wherein the sample texts contain at least one of a position indicating word and an enhancement indicating word;
[0116] If there are position indicating words in the sample text, the sample image is framed based on the position indicating words to obtain sample area coordinates;
[0117] If there are enhancement indicating words in the sample text, the sample image is enhanced based on the enhancement indicating words to obtain a sample clear image.
[0118] A first sample acquisition unit 11 is used to acquire a first sample data set, where the first sample data set includes a first sample image, a first sample text, and sample region coordinates obtained by selecting the first sample image based on the first sample text;
[0119] A region selection model training unit 12 is used to perform model training on an initial region selection large model based on the first sample data set to obtain a region selection large model;
[0120] Optionally, the region selection model training unit 12 is specifically used to create an initial region selection large model;
[0121] Inputting the first sample image and the first sample text into the initial region selection large model, and obtaining the training region coordinates obtained by the initial region selection large model performing frame selection processing on the first sample image based on the first sample text;
[0122] Based on the training area coordinates and the sample area coordinates, the initial area selection large model is subjected to parameter adjustment processing until the model training is completed to obtain the area selection large model.
[0123] Optionally, the region selection model training unit 12 is specifically used to input the first sample image and the first sample text into the initial region selection model;
[0124] The initial region selection model is controlled to obtain position indicating words in the first sample text, and the first sample image is frame-selected based on the position indicating words to generate training region coordinates.
[0125] A second sample acquisition unit 13 is used to acquire a second sample data set, where the second sample data set includes a second sample image, a second sample text, and a sample clear image obtained by enhancing the second sample image based on the second sample text;
[0126] An image enhancement model training unit 14 is used to perform model training on the initial image enhancement large model based on the second sample data set to obtain the image enhancement large model;
[0127] Optionally, the image enhancement model training unit 14 is specifically used to create an initial image enhancement large model;
[0128] Inputting the second sample image and the second sample text into the initial image enhancement large model, and obtaining a training clear image obtained by the initial image enhancement large model performing enhancement processing on the second sample image based on the second sample text;
[0129] Based on the training clear image and the sample clear image, the parameters of the initial image enhancement large model are adjusted until the model training is completed to obtain the image enhancement large model.
[0130] Optionally, the image enhancement model training unit 14 is specifically used to create an initial image enhancement large model based on a hybrid expert architecture, and the initial image enhancement large model includes at least one expert layer, and each expert layer corresponds to a different image enhancement method.
[0131] Optionally, the image enhancement model training unit 14 is specifically used to input the second sample image and the second sample text into the initial image enhancement large model;
[0132] Controlling the initial image enhancement large model to obtain the enhancement indication words in the second sample text, and determining the training image enhancement mode corresponding to the enhancement indication words;
[0133] The initial image enhancement model is controlled to adopt the target expert layer corresponding to the training image enhancement method to perform enhancement processing on the second sample image to obtain a training clear image.
[0134] The image enhancement processing unit 15 is used to perform image enhancement processing based on the region selection large model and the image enhancement large model.
[0135] Optionally, the image enhancement processing unit 15 is specifically used to obtain the target image and target instruction text input by the user;
[0136] Input the target image and the target instruction text into the region selection large model, and obtain the target region coordinates output by the region selection large model;
[0137] Acquire a target partial image in the target image based on the target area coordinates;
[0138] Inputting the target local image and the target instruction text into the image enhancement large model, and obtaining the target enhanced image output by the image enhancement large model;
[0139] The target enhanced image replaces the target partial image in the target image to obtain the target image after image enhancement processing.
[0140] In this embodiment, a sample image and a sample text are obtained. If there are position indicating words in the sample text, the sample image is framed based on the position indicating words to obtain the sample area coordinates. If there are enhancement indicating words in the sample text, the sample image is enhanced based on the enhancement indicating words to obtain a sample clear image. The sample image is processed according to the indicating words in the sample text so that the subsequent model can learn the ability to process images based on text, thereby improving the accuracy of image enhancement. A first sample data set is obtained, and the first sample data set includes a first sample image, a first sample text, and sample area coordinates. Model training of a large model for region selection is completed based on the first sample data set and the intersection-over-union loss function, thereby improving the accuracy of the large model for region selection in object positioning, so that the large model for region selection further improves the accuracy of frame selection. A second sample data set is obtained, which includes a second sample image, a second sample text, and a sample clear image. The initial image enhancement model is trained based on the second sample data set to obtain the image enhancement model. The image enhancement model is created based on the MoE architecture. Each expert layer focuses on different image enhancement methods, so that the image enhancement model only needs to activate the corresponding expert layer in actual application, thereby improving the accuracy of the enhancement processing of the corresponding image enhancement method and reducing the amount of image enhancement calculation. Image enhancement processing is performed based on the region selection model and the image enhancement model. The local image indicated by the user is found through the trained region selection model, and then the image enhancement model is used to enhance the local image. The processing of the local image improves the pertinence of the image enhancement and reduces the amount of image enhancement calculation. And the target enhanced image replaces the target local image in the target image, further avoiding unnecessary processing of the rest of the target image, reducing the amount of calculation and improving the efficiency of image enhancement.
[0141] It should be noted that, when the image enhancement device provided in the above embodiment executes the image enhancement method, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the image enhancement device provided in the above embodiment and the image enhancement method embodiment belong to the same concept, and the implementation process thereof is detailed in the method embodiment, which will not be repeated here.
[0142] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0143] The present application also provides a computer storage medium, which can store multiple instructions, and the instructions are suitable for being loaded and executed by a processor as described above. Figure 1-Figure 6The image enhancement method of the embodiment shown in the figure can be specifically executed by referring to Figure 1-Figure 6 The specific description of the illustrated embodiment will not be repeated here.
[0144] The present application also provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded and executed by the processor as described above. Figure 1-Figure 6 The image enhancement method of the embodiment shown in the figure can be specifically executed by referring to Figure 1-Figure 6 The specific description of the illustrated embodiment will not be repeated here.
[0145] Please refer to Fig. 9 , which shows a block diagram of the structure of an electronic device provided by an exemplary embodiment of the present application. The electronic device in the present application may include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, the memory 120, the input device 130, and the output device 140 may be connected via the bus 150.
[0146] The processor 110 may include one or more processing cores. The processor 110 uses various interfaces and lines to connect various parts of the entire electronic device, and executes various functions and processes data of the terminal 100 by running or executing instructions, programs, code sets or instruction sets stored in the memory 120, and calling data stored in the memory 120. Optionally, the processor 110 can be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 110 can integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user pages, and applications; the GPU is responsible for rendering and drawing display content; and the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 110, but may be implemented separately through a communication chip.
[0147] The memory 120 may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory 120 includes a non-transitory computer-readable medium (Non-Transitory Computer-Readable Storage Medium). The memory 120 may be used to store instructions, programs, codes, code sets or instruction sets. The memory 120 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc. The operating system may be an Android system, including a system deeply developed based on the Android system, an IOS system developed by Apple, including a system deeply developed based on the IOS system or other systems.
[0148] The memory 120 can be divided into an operating system space and a user space. The operating system runs in the operating system space, and native and third-party applications run in the user space. In order to ensure that different third-party applications can achieve good operating results, the operating system allocates corresponding system resources to different third-party applications. However, the requirements for system resources in different application scenarios in the same third-party application are also different. For example, in the local resource loading scenario, the third-party application has higher requirements for disk reading speed; in the animation rendering scenario, the third-party application has higher requirements for GPU performance. The operating system and third-party applications are independent of each other, and the operating system often cannot perceive the current application scenario of the third-party application in a timely manner, resulting in the operating system being unable to perform targeted system resource adaptation according to the specific application scenario of the third-party application.
[0149] In order for the operating system to distinguish the specific application scenarios of third-party applications, it is necessary to open up data communication between third-party applications and the operating system so that the operating system can obtain the current scenario information of third-party applications at any time, and then perform targeted system resource adaptation based on the current scenario.
[0150] The input device 130 is used to receive input commands or data, and includes but is not limited to a keyboard, a mouse, a camera, a microphone, or a touch device. The output device 140 is used to output commands or data, and includes but is not limited to a display device and a speaker. In one example, the input device 130 and the output device 140 can be combined, and the input device 130 and the output device 140 are touch screen displays.
[0151] The touch display screen can be designed as a full screen, a curved screen or a special-shaped screen. The touch display screen can also be designed as a combination of a full screen and a curved screen, or a combination of a special-shaped screen and a curved screen, which is not limited in the embodiments of the present application.
[0152] In addition, those skilled in the art will appreciate that the structure of the electronic device shown in the above drawings does not constitute a limitation on the electronic device, and the electronic device may include more or fewer components than shown, or combine certain components, or arrange the components differently. For example, the electronic device also includes a radio frequency circuit, an input unit, a sensor, an audio circuit, a wireless fidelity (WiFi) module, a power supply, a Bluetooth module and other components, which will not be described in detail here.
[0153] exist Fig. 9 In the electronic device shown, the processor 110 may be used to call the image enhancement application stored in the memory 120 and specifically perform the following operations:
[0154] Acquire a first sample data set, where the first sample data set includes a first sample image, a first sample text, and sample area coordinates obtained by performing a frame selection process on the first sample image based on the first sample text;
[0155] Performing model training on the initial region selection large model based on the first sample data set to obtain the region selection large model;
[0156] Acquire a second sample data set, where the second sample data set includes a second sample image, a second sample text, and a sample clear image obtained by enhancing the second sample image based on the second sample text;
[0157] Performing model training on the initial image enhancement large model based on the second sample data set to obtain the image enhancement large model;
[0158] Image enhancement processing is performed based on the region selection large model and the image enhancement large model.
[0159] In one embodiment, before acquiring the first sample data set, the processor 110 further performs the following operations:
[0160] Acquire a sample image and a sample text, wherein the sample text contains at least one of a position indicating word and an enhancement indicating word;
[0161] If there are position indicating words in the sample text, the sample image is framed based on the position indicating words to obtain sample area coordinates;
[0162] If there are enhancement indicating words in the sample text, the sample image is enhanced based on the enhancement indicating words to obtain a sample clear image.
[0163] In one embodiment, when the processor 110 performs model training on the initial region selection large model based on the first sample data set to obtain the region selection large model, the processor 110 specifically performs the following operations:
[0164] Create a large model for initial area selection;
[0165] Inputting the first sample image and the first sample text into the initial region selection large model, and obtaining the training region coordinates obtained by the initial region selection large model performing frame selection processing on the first sample image based on the first sample text;
[0166] Based on the training area coordinates and the sample area coordinates, the initial area selection large model is subjected to parameter adjustment processing until the model training is completed to obtain the area selection large model.
[0167] In one embodiment, when the processor 110 inputs the first sample image and the first sample text into the initial region selection model and obtains the training region coordinates obtained by the initial region selection model performing a frame selection process on the first sample image based on the first sample text, the processor 110 specifically performs the following operations:
[0168] Inputting the first sample image and the first sample text into the initial area selection model;
[0169] The initial region selection model is controlled to obtain position indicating words in the first sample text, and the first sample image is frame-selected based on the position indicating words to generate training region coordinates.
[0170] In one embodiment, when the processor 110 performs model training on the initial image enhancement large model based on the second sample data set to obtain the image enhancement large model, the processor 110 specifically performs the following operations:
[0171] Create an initial image enhancement model;
[0172] Inputting the second sample image and the second sample text into the initial image enhancement large model, and obtaining a training clear image obtained by the initial image enhancement large model performing enhancement processing on the second sample image based on the second sample text;
[0173] Based on the training clear image and the sample clear image, the parameters of the initial image enhancement large model are adjusted until the model training is completed to obtain the image enhancement large model.
[0174] In one embodiment, when the processor 110 creates the initial image enhancement model, the processor 110 specifically performs the following operations:
[0175] An initial image enhancement large model is created based on a hybrid expert architecture, wherein the initial image enhancement large model comprises at least one expert layer, and each expert layer corresponds to a different image enhancement method.
[0176] In one embodiment, when the processor 110 inputs the second sample image and the second sample text into the initial image enhancement large model and obtains the training clear image obtained by the initial image enhancement large model enhancing the second sample image based on the second sample text, the processor 110 specifically performs the following operations:
[0177] Inputting the second sample image and the second sample text into the initial image enhancement model;
[0178] Controlling the initial image enhancement large model to obtain the enhancement indication words in the second sample text, and determining the training image enhancement mode corresponding to the enhancement indication words;
[0179] The initial image enhancement model is controlled to adopt the target expert layer corresponding to the training image enhancement method to perform enhancement processing on the second sample image to obtain a training clear image.
[0180] In one embodiment, when the processor 110 performs image enhancement processing based on the region selection large model and the image enhancement large model, the processor 110 specifically performs the following operations:
[0181] Obtain the target image and target instruction text input by the user;
[0182] Input the target image and the target instruction text into the region selection large model, and obtain the target region coordinates output by the region selection large model;
[0183] Acquire a target partial image in the target image based on the target area coordinates;
[0184] Inputting the target local image and the target instruction text into the image enhancement large model, and obtaining the target enhanced image output by the image enhancement large model;
[0185] The target enhanced image replaces the target partial image in the target image to obtain the target image after image enhancement processing.
[0186] In this embodiment, a sample image and a sample text are obtained. If there are position indicating words in the sample text, the sample image is framed based on the position indicating words to obtain the sample area coordinates. If there are enhancement indicating words in the sample text, the sample image is enhanced based on the enhancement indicating words to obtain a sample clear image. The sample image is processed according to the indicating words in the sample text so that the subsequent model can learn the ability to process images based on text, thereby improving the accuracy of image enhancement. A first sample data set is obtained, and the first sample data set includes a first sample image, a first sample text, and sample area coordinates. Model training of a large model for region selection is completed based on the first sample data set and the intersection-over-union loss function, thereby improving the accuracy of the large model for region selection in object positioning, so that the large model for region selection further improves the accuracy of frame selection. A second sample data set is obtained, which includes a second sample image, a second sample text, and a sample clear image. The initial image enhancement model is trained based on the second sample data set to obtain the image enhancement model. The image enhancement model is created based on the MoE architecture. Each expert layer focuses on different image enhancement methods, so that the image enhancement model only needs to activate the corresponding expert layer in actual application, thereby improving the accuracy of the enhancement processing of the corresponding image enhancement method and reducing the amount of image enhancement calculation. Image enhancement processing is performed based on the region selection model and the image enhancement model. The local image indicated by the user is found through the trained region selection model, and then the image enhancement model is used to enhance the local image. The processing of the local image improves the pertinence of the image enhancement and reduces the amount of image enhancement calculation. And the target enhanced image replaces the target local image in the target image, further avoiding unnecessary processing of the rest of the target image, reducing the amount of calculation and improving the efficiency of image enhancement.
[0187] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only storage memory, or a random access memory, etc.
[0188] The above disclosure is only the preferred embodiment of the present application, which certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.
[0189] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in the embodiments of this specification are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions. For example, the sample images, sample texts, etc. involved in this specification are all obtained with full authorization.
Claims
1. An image enhancement method, characterized in that: The method comprises: Acquire a first sample data set, where the first sample data set includes a first sample image, a first sample text, and sample area coordinates obtained by performing a frame selection process on the first sample image based on the first sample text; Performing model training on the initial region selection large model based on the first sample data set to obtain the region selection large model; Acquire a second sample data set, where the second sample data set includes a second sample image, a second sample text, and a sample clear image obtained by enhancing the second sample image based on the second sample text; Performing model training on the initial image enhancement large model based on the second sample data set to obtain the image enhancement large model; Image enhancement processing is performed based on the region selection large model and the image enhancement large model.
2. The method according to claim 1, characterized in that Before obtaining the first sample data set, the method further includes: Acquire a sample image and a sample text, wherein the sample text contains at least one of a position indicating word and an enhancement indicating word; If there are position indicating words in the sample text, the sample image is framed based on the position indicating words to obtain sample area coordinates; If there are enhancement indicating words in the sample text, the sample image is enhanced based on the enhancement indicating words to obtain a sample clear image.
3. The method according to claim 1, characterized in that The performing model training on the initial region selection large model based on the first sample data set to obtain the region selection large model includes: Create a large model for initial area selection; Inputting the first sample image and the first sample text into the initial region selection large model, and obtaining the training region coordinates obtained by the initial region selection large model performing frame selection processing on the first sample image based on the first sample text; Based on the training area coordinates and the sample area coordinates, the initial area selection large model is subjected to parameter adjustment processing until the model training is completed to obtain the area selection large model.
4. The method according to claim 3, characterized in that The step of inputting the first sample image and the first sample text into the initial region selection large model, and obtaining the training region coordinates obtained by the initial region selection large model performing frame selection processing on the first sample image based on the first sample text, comprises: Inputting the first sample image and the first sample text into the initial area selection model; The initial region selection model is controlled to obtain position indicating words in the first sample text, and the first sample image is frame-selected based on the position indicating words to generate training region coordinates.
5. The method according to claim 1, characterized in that The performing model training on the initial image enhancement large model based on the second sample data set to obtain the image enhancement large model includes: Create an initial image enhancement model; Inputting the second sample image and the second sample text into the initial image enhancement large model, and obtaining a training clear image obtained by the initial image enhancement large model performing enhancement processing on the second sample image based on the second sample text; Based on the training clear image and the sample clear image, the parameters of the initial image enhancement large model are adjusted until the model training is completed to obtain the image enhancement large model.
6. The method according to claim 5, characterized in that The step of creating an initial image enhancement model comprises: An initial image enhancement large model is created based on a hybrid expert architecture, wherein the initial image enhancement large model comprises at least one expert layer, and each expert layer corresponds to a different image enhancement method.
7. The method according to claim 6, characterized in that The step of inputting the second sample image and the second sample text into the initial image enhancement large model, and obtaining the training clear image obtained by performing enhancement processing on the second sample image by the initial image enhancement large model based on the second sample text, comprises: Inputting the second sample image and the second sample text into the initial image enhancement model; Controlling the initial image enhancement large model to obtain the enhancement indication words in the second sample text, and determining the training image enhancement mode corresponding to the enhancement indication words; The initial image enhancement model is controlled to adopt the target expert layer corresponding to the training image enhancement method to perform enhancement processing on the second sample image to obtain a training clear image.
8. An image enhancement device, characterized in that: The device comprises: A first sample acquisition unit is used to acquire a first sample data set, where the first sample data set includes a first sample image, a first sample text, and sample area coordinates obtained by performing a frame selection process on the first sample image based on the first sample text; A region selection model training unit, used for performing model training on an initial region selection large model based on the first sample data set to obtain a region selection large model; A second sample acquisition unit is used to acquire a second sample data set, where the second sample data set includes a second sample image, a second sample text, and a sample clear image obtained by enhancing the second sample image based on the second sample text; An image enhancement model training unit, used for performing model training on an initial image enhancement large model based on the second sample data set to obtain an image enhancement large model; An image enhancement processing unit is used to perform image enhancement processing based on the region selection large model and the image enhancement large model.
9. A computer storage medium, characterized in that: The computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the method steps according to any one of claims 1 to 7.
10. An electronic device, characterized in that: include: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the method steps as claimed in any one of claims 1 to 7.
Citation Information
Patent Citations
Image local reinforcement method and device
CN103310411A
Image recognition method and device, electronic equipment and storage medium
CN115331150A
Image data enhancement method and device and image recognition method and device
CN115496965A
Image enhancement method and device, equipment and storage medium
CN118429191A
Monitoring Method of The Status of Garbage Discharge Using the Renewal Volume-Based Garbage Bag and CCTV
KR102856298B1
Cited By
Guardrail detection method, device and equipment and storage medium
CN120976709A