Image processing method and device, electronic equipment, storage medium and program product
By performing global or local image quality processing on images, generating and fusing mask images, the problem of obtaining high-definition-low-definition image pairs is solved, a rich dataset is constructed, and the actual effect of image processing models is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2025-01-21
- Publication Date
- 2026-07-21
AI Technical Summary
Existing technologies struggle to effectively obtain abundant high-definition-low-definition image pairs, resulting in poor performance of image processing models in practical applications.
By acquiring the first image, performing global or local image quality processing, determining the target area, generating a mask image, and fusing it with the original image, image pairs with different image qualities are formed.
The generated image pairs can build datasets, reduce the difference between the datasets and real-world scenarios, help train models, and improve the practical application of image processing models.
Smart Images

Figure CN122434745A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, specifically to image processing methods, apparatus, electronic devices, storage media, and program products. Background Technology
[0002] In image processing, high-resolution-low-resolution image pairs are often used for supervised training. That is, a low-resolution image is used as input, and the corresponding high-resolution image is used to supervise the image after enhancement or restoration by the algorithm. However, real high-resolution-low-resolution image pairs are very difficult to obtain, resulting in image models trained using datasets of image pairs performing poorly in practical applications. Therefore, there is a need to provide an image processing method to obtain abundant high-resolution-low-resolution image pairs. Summary of the Invention
[0003] In view of this, the present disclosure provides an image processing method, apparatus, electronic device, storage medium, and program product to solve the problem of obtaining image pairs.
[0004] In a first aspect, this disclosure provides an image processing method, including:
[0005] Get the first image of the highest quality;
[0006] The first image is processed to obtain a second image with a second quality, and the first quality is different from the second quality.
[0007] Based on the image information of the first image, a target area in the first image that needs to be processed is determined, wherein the size of the target area is smaller than the size of the first image.
[0008] Based on the location information of the target region in the first image and the size of the first image, a mask image of the first image is obtained. The pixel value of the target region in the mask image is different from the pixel value of other regions in the mask image. The other regions are the regions in the mask image other than the target region.
[0009] A third image is obtained by fusing the first image, the second image, and the mask image;
[0010] The first image is associated with the third image to obtain image pairs with different image qualities.
[0011] Secondly, this disclosure provides an image processing apparatus, comprising:
[0012] The image acquisition module is used to acquire a first image of first quality.
[0013] The quality processing module is used to perform image quality processing on the entire first image to obtain a second image with a second image quality, wherein the first image quality is different from the second image quality.
[0014] The region determination module is used to determine, based on the image information of the first image, a target region in the first image that needs to undergo image quality processing, wherein the size of the target region is smaller than the size of the first image.
[0015] A mask processing module is used to obtain a mask image of the first image based on the position information of the target region in the first image and the size of the first image. The pixel value of the target region in the mask image is different from the pixel value of other regions in the mask image. The other regions are regions in the mask image other than the target region.
[0016] An image fusion module is used to obtain a third image by fusing the first image, the second image, and the mask image;
[0017] The image pairing module is used to associate the first image with the third image to obtain image pairs with different image qualities.
[0018] Thirdly, this disclosure provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the image processing method described in the first aspect or any corresponding embodiment thereof.
[0019] Fourthly, this disclosure provides a computer-readable storage medium storing computer instructions for causing a computer to perform the image processing method described in the first aspect or any corresponding embodiment thereof.
[0020] Fifthly, this disclosure provides a computer program product, including computer instructions for causing a computer to perform the image processing method described in the first aspect or any corresponding embodiment thereof.
[0021] The image processing method provided in this disclosure involves acquiring a first image of a first quality, performing global quality processing on it to obtain a second image of a second quality; determining the target region in the first image that needs quality processing based on the image information of the first image, and obtaining a mask image of the first image. The pixel values of the target region in this mask image are different from the pixel values of the other regions, thus distinguishing the target region from the rest. A third image is then obtained by fusing the first image, the second image, and the mask image. Image pairs of different quality are obtained based on the first image and the third image. In this method, based on the first image, the target region in the first image that needs quality processing is determined according to the image information of the first image, and a corresponding mask image is obtained. This allows for quality processing of a portion of the first image as needed, and then, combined with the original first image, image pairs of different quality are obtained. This method can utilize the first image to obtain corresponding image pairs. The first image can be a low-resolution image, which can be enhanced to obtain a corresponding high-resolution image; or the first image can be a high-resolution image, which can be degraded to obtain a corresponding low-resolution image. Therefore, this method can obtain a rich variety of image pairs. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the specific embodiments or related technologies of this disclosure, the accompanying drawings used in the description of the specific embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a schematic flowchart of an image processing method according to an embodiment of the present disclosure;
[0024] Figure 2 This is a schematic flowchart of another image processing method according to an embodiment of the present disclosure;
[0025] Figure 3 This is a schematic diagram of image processing according to an embodiment of the present disclosure;
[0026] Figure 4 This is a schematic diagram of another image processing according to an embodiment of the present disclosure;
[0027] Figure 5 This is a schematic diagram of another image processing according to an embodiment of the present disclosure;
[0028] Figure 6 This is a schematic flowchart of another image processing method according to an embodiment of the present disclosure;
[0029] Figure 7 This is a schematic diagram of another image processing according to an embodiment of the present disclosure;
[0030] Figure 8 This is a structural block diagram of an image processing apparatus according to an embodiment of the present disclosure;
[0031] Figure 9 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this disclosure. Detailed Implementation
[0032] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0033] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0034] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0035] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0036] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0037] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0038] In related technologies, obtaining image pairs typically involves performing lossy quality processing on high-resolution images to generate corresponding low-resolution images. However, the distribution of low-resolution images generated by lossy quality processing differs from the distribution of real low-resolution images, and some types of low-resolution images are difficult to reproduce by algorithms, such as old photos. This can lead to poor performance of trained image processing models in practical applications.
[0039] Based on this, embodiments of this disclosure provide an image processing method. First, the target region requiring image quality processing is determined based on the image information of a first image. Then, image quality processing is performed on the target region to obtain a third image. Accordingly, the first image and the third image form an image pair with different image qualities. This method can perform image quality processing on all or part of a region in a first image to obtain corresponding globally or locally processed images forming image pairs, which are used to construct a dataset for training an image processing model.
[0040] Image pairs obtained based on the image processing method provided in this disclosure can be used to construct a dataset. This dataset can obtain real low-resolution to high-resolution image pairs that match the application scenario, reducing the data discrepancy between the dataset and the actual scenario, and facilitating the implementation of image processing models. Furthermore, in this disclosure, the image processing method is for enhancing low-resolution images, or it can be for processing high-resolution images with lossy quality; therefore, a richer dataset can be obtained.
[0041] Furthermore, the constructed dataset can be used to train image enhancement models for low-resolution images, and multiple enhancement models can be employed for image enhancement or repair for different types of image quality defects. Then, a student model with relatively low complexity learns the enhancement results of these complex teacher models, enabling the student model to efficiently acquire the capabilities of multiple teacher models.
[0042] The image processing method provided in this disclosure, in addition to improving image quality, can also obtain corresponding image processing information for constructing a multimodal dataset. This data serves as labels for corresponding image pairs, facilitating the training of multimodal models and the implementation of personalized applications. For example, the trained model can enhance the image quality of a specified region or target by inputting prompts. In other words, the dataset constructed based on the image pairs obtained by the image processing method provided in this disclosure can be used for knowledge distillation of complex models, simulating real-world data distribution, and assisting in model deployment, among other things.
[0043] According to an embodiment of this disclosure, an image processing method embodiment is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0044] This embodiment provides an image processing method that can be used in electronic devices, such as servers. Figure 1 This is a flowchart of an image processing method according to an embodiment of the present disclosure, such as... Figure 1 As shown, the process includes the following steps:
[0045] Step S101: Obtain the first image of the first image quality.
[0046] It should be noted that the first image quality can be the low-resolution quality described above, and correspondingly, the third image quality is the high-resolution quality described above. That is, the resulting image pair is a low-resolution-high-resolution image pair. Therefore, the subsequent image quality processing of the entire image described is image quality enhancement processing.
[0047] Furthermore, the first image quality can also be the high-definition quality described above, and correspondingly, the third image quality is the low-definition quality described above. Therefore, the subsequent image quality processing for the entire image is a lossy processing.
[0048] Accordingly, the image processing method provided in this disclosure can either enhance the image quality of a low-resolution image to obtain a corresponding high-resolution image, thereby forming a low-resolution-high-resolution image pair; or it can perform lossy image quality processing on a high-resolution image to obtain a corresponding low-resolution image, thereby forming a high-resolution-low-resolution image pair. Therefore, the type of the first image quality is determined according to actual needs, and no limitation is made here.
[0049] The first image in first-quality mode can be captured by video, extracted from a pre-existing dataset, and so on. Of course, the first image in first-quality mode can also be obtained in other ways; there are no restrictions on the acquisition method here, and it can be set according to actual needs.
[0050] Step S102: Perform image quality processing on the entire first image to obtain a second image with second image quality.
[0051] Image quality processing methods include overall image quality processing and local image quality processing. Overall image quality processing is also known as global image quality processing, while local image quality processing is known as local image quality processing. Image quality processing involves combining different image processing techniques with appropriate algorithm parameters to generate images of varying quality. Global image quality processing is suitable for repairing global image quality defects generated during image capture or encoding, or for simulating global image quality defects in specific scenarios. For example, global image quality enhancement can be applied to image noise generated during shooting under certain lighting conditions, or to address blockage artifacts that may occur during image encoding. Local image quality processing, on the other hand, applies image quality enhancement or processing to a specific area of the image. In image processing applications, local image quality processing often involves spot repair or simulating spots.
[0052] In this embodiment, a second image of a second quality is obtained by performing global image quality processing on the first image. For example, if the first image quality is low-resolution, then the global image quality processing here is global image quality enhancement, and correspondingly, the second image quality is high-resolution. If the first image quality is high-resolution, then the global image quality processing here is global image quality loss, and correspondingly, the second image quality is low-resolution.
[0053] For ease of description below, this embodiment of the disclosure uses low-resolution image quality as an example and high-resolution image quality as an example. That is, the image processing method provided in this embodiment of the disclosure can obtain a high-resolution third image corresponding to the low-resolution first image, forming a low-resolution-high-resolution image pair.
[0054] Step S103: Based on the image information of the first image, determine the target area in the first image that needs to be processed.
[0055] The size of the target region is smaller than the size of the first image.
[0056] The image information of the first image includes, but is not limited to, edge information, target information, graphic element information, etc., and no limitations are imposed here. The target region in the first image can be determined through the image information of the first image, so that subsequent image processing can be performed on the target region.
[0057] Different first images require different areas for image quality processing because some contain more text, some more lines, and some may have more edge information, etc. Therefore, it is necessary to use the image information of the first image to determine the target area within it. It should be noted that the target area is the region in the first image that requires image quality processing.
[0058] The target region can be determined by detecting image elements based on the image information of the first image, detecting a specific target, or performing edge detection, etc. Therefore, the target region in the first image that needs image quality processing is obtained based on the detection results of the corresponding detection method.
[0059] Step S104: Based on the position information of the target region in the first image and the size of the first image, obtain the mask image of the first image.
[0060] In this mask image, the pixel values of the target region are different from the pixel values of other regions in the mask image.
[0061] After determining the target region, a mask image of the first image is obtained by combining the target region with a mask of the same size as the first image. This mask image is then used for subsequent image processing targeting the target region. The mask image is the same size as the first image, and its position in the mask image is obtained by mapping the target region's location from its position in the first image. Accordingly, corresponding pixel values can be set for the target region and other regions in the mask image. For example, the pixel value corresponding to the target region in the mask image is 1, and the pixel value for the remaining regions in the mask image is 0.
[0062] Step S105: Based on the fusion of the first image, the second image, and the mask image, a third image is obtained.
[0063] The first image can be referred to as the original image for image processing, the second image is the image after global image quality processing, and the mask image is used to perform image processing on the target area in the first image. Based on this, fusing the second image with the mask image can obtain an image after image processing only on the target area; then fusing this image with the first image yields the third image.
[0064] Step S106: Associate the first image with the third image to obtain image pairs with different image qualities.
[0065] The third image is the first image after image enhancement or lossy image processing. As described above: if the first image is a low-resolution image, the third image is obtained after image enhancement; if the first image is a high-resolution image, the third image is obtained after lossy image processing. Therefore, the first image and the third image can form image pairs with different image qualities.
[0066] The image processing method provided in this embodiment, based on a first image, determines the target area in the first image that needs image quality processing according to the image information of the first image, and obtains the corresponding mask image. This allows for image quality processing of specific areas of the first image as needed. Combined with the original first image, image pairs of different image qualities are obtained. This method can utilize the first image to obtain corresponding image pairs. The first image can be a low-resolution image, which can be enhanced to obtain a corresponding high-resolution image; alternatively, the first image can be a high-resolution image, which can be degraded to obtain a corresponding low-resolution image. Therefore, this method can obtain a rich variety of image pairs.
[0067] This embodiment provides an image processing method that can be used in electronic devices, such as servers. Figure 2 This is a flowchart of an image processing method according to an embodiment of the present disclosure, such as... Figure 2 As shown, the process includes the following steps:
[0068] Step S201: Acquire the first image of the first quality. See details... Figure 1 Step S101 of the illustrated embodiment will not be described again here.
[0069] Step S202: Perform image quality processing on the entire first image to obtain a second image with second-highest image quality. See details... Figure 1 Step S102 of the illustrated embodiment will not be described again here.
[0070] Step S203: Based on the image information of the first image, determine the target region in the first image to obtain the mask image of the first image.
[0071] In this case, the pixel values of the target region in the mask image are different from the pixel values of other regions in the mask image.
[0072] The image information of the first image can be used to perform target detection, image element segmentation, or edge detection to determine the target area in the first image that needs image quality processing.
[0073] In some optional implementations, step S203 above includes:
[0074] Step a1: Obtain the target corpus, which includes image elements that need to be processed.
[0075] Step a2: Based on the language model and the first prompt information, determine the selectable image elements in the first image. The first prompt information is used to instruct the language model to recognize the image elements in the first image.
[0076] Step a3: Use the target corpus to filter the selectable image elements to obtain the target image elements in the first image.
[0077] Step a4: Determine the target region where the target image element is located in the first image.
[0078] The target corpus includes image elements that require image quality processing. These image elements include, but are not limited to, targets in the images, such as cars, trees, and buildings. Image elements in the target corpus can be described using natural language or represented by unique identifiers. No restrictions are placed on the representation format of image elements in the target corpus; the specific format can be set according to actual needs.
[0079] It should be noted that the target corpus can be related to the current image processing task; that is, different image processing tasks correspond to different target corpora.
[0080] The first prompt information is used to instruct the language model to recognize image elements in the first image. That is, the first prompt information and the first image are input into the language model, and the language model outputs optional image elements from the first image. These optional image elements represent the image elements in the first image recognized by the language model. The language model can be a Multimodal Large Language Model (MLLM), or other types of language models, etc., and its specific form is not limited here.
[0081] As described above, the target corpus includes image elements that require image quality processing. This means that not all selectable image elements identified by the language model require image quality processing. Therefore, by filtering the selectable image elements from the target corpus, the target image elements in the first image that require image quality processing are obtained.
[0082] For example, the language model outputs optional image elements: car, building, road surface. However, the target corpus does not include the image element of building. Therefore, by comparing and removing the image element of building from the optional image elements, the target image elements in the first image that need to be processed are finally determined to be: car and road surface.
[0083] It should be noted that if multiple target image elements are identified, one can be randomly selected for subsequent target region determination; alternatively, target regions can be determined sequentially for each target image element. That is, after image quality processing is performed on the first target image element, image quality processing is then performed on the second target image element based on the processing, and so on.
[0084] After identifying the target image element in the first image, the position of the target image element in the first image can be obtained, and the target area where the target image element is located can be determined accordingly.
[0085] By using the first prompt information and language model to identify the optional image elements in the first image, the comprehensiveness of image element identification in the first image can be guaranteed. By combining the target corpus to filter the optional image elements, the amount of image element processing can be reduced and the efficiency of image processing can be improved.
[0086] In some alternative implementations, step a4 above includes:
[0087] Step a41: Generate second prompt information based on the target image elements. The second prompt information is used to instruct the image segmentation model to segment the target image elements.
[0088] Step a42: Based on the image segmentation model and the second prompt information, the first image is segmented into target image elements to obtain the target region.
[0089] After obtaining the target image elements in the first image, the information of the target image elements can be filled into the prompt template to obtain the second prompt information. That is, a second prompt is obtained to instruct the image segmentation model to perform target image element segmentation on the first image.
[0090] The second prompt information and the first image are input into the image segmentation model to segment the target image elements in the first image to obtain the corresponding target regions. The recording of the target regions can be either simply retaining the target regions in an image of the same size as the first image, or recording the position information of the target regions within the first image, etc., and no specific limitations are imposed here.
[0091] Figure 3 The diagram illustrates image processing. A first image and a first prompt are input into a language model. The output of the language model is filtered through a target corpus to obtain target image elements. Based on this, a second prompt is generated. The second prompt and the first image are then input into an image segmentation model to segment the target image elements from the first image, thus obtaining the target region.
[0092] By combining the second prompt information with the image segmentation model, the image elements of the first image are segmented, which improves the efficiency and accuracy of image element segmentation.
[0093] In some optional implementations, step S203 above includes:
[0094] Step b1: Based on the image information of the first image, detect the specified target in the first image and determine the detection result of the specified target.
[0095] Step b2: Based on the location information in the detection results, determine the target region in the first image.
[0096] If the target to be enhanced contains many thin lines (such as text), image segmentation-based enhancement algorithms are not suitable. In this case, object detection-based algorithms can be used for local enhancement. The detection of a specific target can be implemented using an object detection algorithm, which is determined by analyzing the image information of the first image. The object detection algorithm obtains the detection result of the specified target, which can be the location of the detection box or other forms representing the detection result.
[0097] The detection results include the location information of the specified target, and the target region in the first image can be determined using this location information.
[0098] Figure 4 The diagram illustrates another image processing step, where target detection is performed on the first image to obtain the detection result, i.e., the target region corresponding to the specified target.
[0099] The detection result of the specified target in the first image is determined by the target detection method. Based on this, the corresponding target area can be accurately determined, so that images with many fine lines of the specified target can be processed.
[0100] In some optional implementations, step S203 includes: performing edge detection on the first image based on the image information of the first image to obtain the edge information of the first image, so as to determine the target region in the first image.
[0101] Compared to flat areas, edge regions in an image are often more prone to image quality defects, such as jagged edges. Therefore, edge detection-based algorithms can be used for local enhancement of image quality defects present at edges. Figure 5 The diagram illustrates another image processing method, in which edge information in the first image can be obtained through an edge detection algorithm, thereby determining the target region in the first image.
[0102] By determining the edge information of the first image through edge detection, the image quality of the first image's edges can be improved.
[0103] Step S204: Based on the position information of the target region in the first image and the size of the first image, obtain the mask image of the first image.
[0104] In this mask image, the pixel values of the target region are different from the pixel values of other regions in the mask image.
[0105] Specifically, step S204 includes:
[0106] Step S2041: Obtain an initial mask image with the same size as the first image. Step S2042: Based on the position information of the target region in the first image, perform mapping in the initial mask image to determine the mapped region in the initial mask image.
[0107] Step S2043: Set the pixel values in the mapped area to the first pixel value, and set the pixel values of the remaining areas in the initial mask image to the second pixel value to obtain the processed mask image.
[0108] Step S2044: Determine the optimization processing method based on the method for determining the target region. The method for determining the target region includes at least one of image element detection based on image information from the first image, detection of a specified target, and edge detection.
[0109] Step S2045: Optimize the edges of the processed mask image according to the optimization processing method to obtain the mask image of the first image.
[0110] The size of the initial mask image is the same as the size of the first image. That is, by obtaining the size of the first image, another image with all pixel values of 0 can be generated as the initial mask image; or an image with all pixel values of 1 can be generated as the initial mask image, and so on.
[0111] The mask image of the first image is used to distinguish the target area from the rest of the first image. That is, by setting the pixel values, the pixel values of the target area are different from the pixel values of the rest of the image.
[0112] After determining the target region, its position information within the first image can be obtained. Since the initial mask image is the same size as the first image, the target region can be mapped onto the initial mask image, resulting in the mapped region within the initial mask image.
[0113] Based on this, the pixel values within the mapped area are set to the first pixel value, and the pixel values of the remaining areas in the initial mask image are set to the second pixel value, thus obtaining the processed mask image.
[0114] Depending on the method used to determine the target region, the optimization methods for the processed mask image will also differ. For example, for target regions determined by object detection or image segmentation, the corresponding optimization process could be edge feathering; for target regions determined by edge detection, the corresponding optimization processes could be dilation, erosion, and so on.
[0115] For example, such as Figure 3 As shown, edge feathering of the segmentation results output by the image segmentation model can yield a mask image; as... Figure 4As shown, the target detection results are feathered at the edges to obtain a mask image; as... Figure 5 As shown, the output of edge detection may contain problems such as thin lines and excessive textures. Therefore, post-processing such as dilation and erosion is performed on the output of edge detection to obtain a mask image.
[0116] After determining the mapping region of the first image in the initial mask image using its positional information, the pixel values of the initial mask image are set accordingly to obtain the processed mask image, thus improving the accuracy of the mask image. Simultaneously, the processed mask image is optimized to obtain a mask image of the first image, ensuring a smooth transition at the edges of the mask image, which facilitates subsequent image fusion.
[0117] It should be noted that the above steps describe three methods for determining the target region. In practical applications, these three methods can be applied to the first image. For example, the first image can be first segmented to obtain a mask image, which is then fused with the first and second images to obtain an image based on image segmentation. Alternatively, the image obtained through image segmentation can be subjected to object detection to obtain a mask image, which is then fused with the first and second images to obtain an image based on object detection. Finally, the image obtained through object detection can be subjected to edge detection to obtain a mask image, which is then fused with the first and second images to obtain an image based on edge detection. Of course, the processing order of the first image using image segmentation, object detection, and edge detection is not limited to the order described above; other orders can also be used.
[0118] Furthermore, if the first image does not contain any image elements from the target corpus, the pixel values of the corresponding mask image can all be 0. The fusion of the mask image with the first and second images will still result in the first image. In other words, after processing the first image using image segmentation, target detection, and edge detection methods, the resulting mask image can be empty, indicating the absence of any corresponding image elements, specified targets, or edges. Here, "absent" can mean a genuine absence, or it can mean that the effective pixel values in the corresponding detection results are below a preset threshold, i.e., there are no effective image elements, effective targets, or effective edges.
[0119] Step S205: Based on the fusion of the first image, the second image, and the mask image, a third image is obtained. See details... Figure 1 Step S105 of the illustrated embodiment will not be described again here.
[0120] Step S206: Associate the first image with the third image to obtain image pairs with different image qualities. See details... Figure 1 Step S106 of the illustrated embodiment will not be described again here.
[0121] The image processing method provided in this embodiment uses image information to determine the target area in the first image that needs to be processed for image quality, and then combines it with an initial mask image of the same size as the first image to process the pixel values to obtain a mask image of the first image, thereby ensuring the accuracy of the obtained mask image.
[0122] This embodiment provides an image processing method that can be used in electronic devices, such as servers. Figure 6 This is a flowchart of an image processing method according to an embodiment of the present disclosure, such as... Figure 6 As shown, the process includes the following steps:
[0123] Step S601: Obtain the first image of the first image quality. See details... Figure 1 Step S101 of the illustrated embodiment will not be described again here.
[0124] Step S602: Perform image quality processing on the entire first image to obtain a second image with second quality. See details... Figure 1 Step S102 of the illustrated embodiment will not be described again here.
[0125] Step S603: Based on the image information of the first image, determine the target area in the first image that needs to be processed.
[0126] The target region is smaller than the first image. See details... Figure 2 Step S203 of the illustrated embodiment will not be described again here.
[0127] Step S604: Based on the position information of the target region in the first image and the size of the first image, obtain the mask image of the first image.
[0128] In this mask image, the pixel values of the target region are different from the pixel values of other regions in the mask image.
[0129] Step S605: Based on the fusion of the first image, the second image, and the mask image, a third image is obtained.
[0130] Specifically, step S605 includes:
[0131] Step S6051: Fuse the second image with the mask image to obtain the fourth image.
[0132] The fusion of the second image and the mask image involves obtaining the target region from the second image, thereby generating the fourth image. In other words, through the fusion process, only the image corresponding to the template region in the second image is retained.
[0133] Step S6052: The fourth image is fused with the remaining regions of the first image except for the target region to obtain the third image.
[0134] For example, the method of obtaining the third image can be represented by the following formula:
[0135] I l =I ori *(1-M)+I g *M
[0136] Among them, I l For the third image, I ori M is the first image, and I is the mask image. g This is the second image.
[0137] Step S606: Associate the first image with the third image to obtain image pairs with different image qualities.
[0138] The first image and the third image form an image pair, used to represent two images with different image qualities.
[0139] The image processing method provided in this embodiment fuses the second image with the mask image to obtain a fourth image that performs image quality processing on the target area; then, it combines the remaining areas of the first image (excluding the target area) to fill the remaining areas of the fourth image to obtain a third image. The third image is obtained through the mask image processing method, which ensures the image processing effect of the obtained third image.
[0140] In some optional embodiments, the image processing method described above further includes:
[0141] Step d1: Obtain the target region and a description of the global image quality processing of the first image.
[0142] Step d2 involves performing semantic processing based on the target region and processing description to obtain image processing information.
[0143] Step d3: Based on the image processing information, the image pairs are associated to obtain the association results of the image pairs.
[0144] Image processing information is used to characterize the processing described in obtaining the third image from the first image, such as the processing method, target region, etc. The image processing information is associated with the image pairs to obtain the association results of the image pairs. The image processing information can be used as labels for image pair samples during subsequent training. The image pair samples used for training can be image pairs consisting of the first image and the third image.
[0145] The target region can be described in natural language, that is, it represents which part of the first image the target region is, such as vehicles and road surfaces. The description of global image quality processing includes the methods used for global image quality processing, such as the image quality processing algorithms employed.
[0146] Furthermore, during data construction, image processing information can be saved as text labels for use in training multimodal models. Text Labels S prompt Including the target area S for image quality processing object Enhancement level S degree and the image processing methods used S method ,Right now:
[0147] S proept =S object +S degree +S method
[0148] If global image quality processing is used for the first image, then S object You can use "the picture"; if you are applying local image quality processing to the first image, you can use S. object "the[object] in the picture", where [object] can be replaced with the target area in the first image that needs local image quality processing.
[0149] The image processing information includes descriptions of the target region and the overall image quality, making the image processing information rich in image processing content, which is convenient for subsequent model training.
[0150] Since the processing from the first image to the third image is obtained by combining the mask image and the target region, and this processing can obtain the corresponding image processing information, associating the image processing information with the image pairs can be used as labels for model training, which facilitates the training of subsequent image-related models.
[0151] As a specific application embodiment of this disclosure, taking image quality enhancement as an example, it is necessary to enhance the image quality of a low-resolution image to obtain a corresponding high-resolution image. For example, such as Figure 7As shown, image quality enhancement can be achieved through either global or local enhancement. Global enhancement involves applying an enhancement algorithm to the original image to obtain a globally enhanced image. Local enhancement can be performed sequentially based on image segmentation, object detection, and edge detection. In other words, the result of the previous local enhancement serves as the source image for the next. For example, after performing local enhancement on the original image based on image segmentation, the resulting enhanced image is used as the source image for object detection, and so on.
[0152] This embodiment also provides an image processing apparatus for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0153] This embodiment provides an image processing device, such as... Figure 8 As shown, it includes:
[0154] Image acquisition module 801 is used to acquire a first image of first quality.
[0155] The quality processing module 802 is used to perform image quality processing on the entire first image to obtain a second image with a second quality, wherein the first quality and the second quality are different.
[0156] The region determination module 803 is used to determine the target region in the first image that needs to be processed based on the image information of the first image. The size of the target region is smaller than the size of the first image.
[0157] The mask processing module 804 is used to obtain a mask image of the first image based on the position information of the target region in the first image and the size of the first image. The pixel values of the target region in the mask image are different from the pixel values of other regions in the mask image. The other regions are the regions in the mask image other than the target region.
[0158] The image fusion module 805 is used to fuse the first image, the second image, and the mask image to obtain the third image.
[0159] Image pairing module 806 is used to associate the first image with the third image to obtain image pairs with different image qualities.
[0160] In some alternative implementations, the mask processing module 804 includes:
[0161] The mask image acquisition unit is used to acquire an initial mask image with the same size as the first image.
[0162] The mapping unit is used to perform mapping in the initial mask image based on the position information of the target region in the first image, and to determine the mapping region in the initial mask image.
[0163] The pixel value setting unit is used to set the pixel value of the mapped area to the first pixel value and set the pixel value of the remaining area in the initial mask image to the second pixel value to obtain the processed mask image.
[0164] The determining unit is used to determine the optimization processing method based on the method of determining the target region. The method of determining the target region includes at least one of image element detection based on image information of the first image, detection of a specified target, and edge detection.
[0165] The optimization unit optimizes the edges of the processed mask image according to the optimization processing method to obtain the mask image of the first image.
[0166] In some alternative implementations, the region determination module 803 includes:
[0167] The corpus acquisition unit is used to acquire the target corpus, which includes image elements that need to be processed.
[0168] The image element determination unit is used to determine selectable image elements in the first image based on the language model and the first prompt information, wherein the first prompt information is used to instruct the language model to recognize the image elements in the first image.
[0169] The image element filtering unit is used to filter the selectable image elements using the target corpus to obtain the target image elements in the first image.
[0170] The first region determination unit is used to determine the target region where the target image element is located in the first image.
[0171] In some optional implementations, the first region determination unit includes:
[0172] The prompt information generation subunit is used to generate second prompt information based on the target image elements. The second prompt information is used to instruct the image segmentation model to segment the target image elements.
[0173] The image element segmentation subunit is used to segment the target image elements of the first image based on the image segmentation model and the second prompt information to obtain the target region.
[0174] In some alternative implementations, the region determination module 803 includes:
[0175] The target detection unit is used to detect a specified target in the first image based on the image information of the first image, and determine the detection result of the specified target.
[0176] The second region determination unit is used to determine the target region in the first image based on the location information in the detection results.
[0177] And / or, in some alternative implementations, the region determination module 803 includes:
[0178] The edge detection unit is used to perform edge detection on the first image based on the image information of the first image, so as to obtain the edge information of the first image and determine the target region in the first image.
[0179] In some alternative embodiments, the image processing apparatus further includes:
[0180] The description acquisition unit is used to acquire the target region and the processing description for global image quality processing of the first image.
[0181] The speech processing unit is used to perform semantic processing based on the target region and the processing description to obtain image processing information.
[0182] The information association module is used to associate image processing information with image pairs to obtain the association results of the image pairs.
[0183] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.
[0184] In this embodiment, the image processing device is presented in the form of a functional unit. Here, a unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0185] This disclosure also provides an electronic device having the above-described features. Figure 8 The image processing device shown.
[0186] Please see Figure 9 , Figure 9 This is a schematic diagram of the structure of an electronic device provided in an optional embodiment of this disclosure, such as... Figure 9As shown, the electronic device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise as required. The processors can process instructions executed within the electronic device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple electronic devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 9 Take a processor 10 as an example.
[0187] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.
[0188] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.
[0189] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0190] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0191] The electronic device also includes a communication interface 30 for communicating with other devices or communication networks.
[0192] This disclosure also provides a computer-readable storage medium in which the methods described in this disclosure can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium after being downloaded over a network. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium may be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium may also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code that, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.
[0193] A portion of this disclosure can be applied to computer program products, such as computer program instructions, which, when executed by a computer, can invoke or provide methods and / or technical solutions according to this disclosure through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, and installation package files. Accordingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions; the computer compiling the instructions and then executing the corresponding compiled program; the computer reading and executing the instructions; or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0194] Although embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. An image processing method, characterized in that, The method includes: Get the first image of the highest quality; The first image is processed to obtain a second image with a second quality, and the first quality is different from the second quality. Based on the image information of the first image, a target area in the first image that needs to be image quality processed is determined, and the size of the target area is smaller than the size of the first image. Based on the location information of the target region in the first image and the size of the first image, a mask image of the first image is obtained. The pixel value of the target region in the mask image is different from the pixel value of other regions in the mask image. The other regions are the regions in the mask image other than the target region. A third image is obtained by fusing the first image, the second image, and the mask image; The first image is associated with the third image to obtain image pairs with different image qualities.
2. The method according to claim 1, characterized in that, The step of obtaining a mask image for the first image based on the position of the target region in the first image and the size of the first image includes: Obtain an initial mask image with the same size as the first image; Based on the location information of the target region in the first image, mapping is performed in the initial mask image to determine the mapped region in the initial mask image; The pixel values within the mapped area are set to the first pixel value, and the pixel values of the remaining areas in the initial mask image are set to the second pixel value, thus obtaining the processed mask image; The optimization processing method is determined based on the method of determining the target region, wherein the method of determining the target region includes at least one of image element detection based on image information of the first image, detection of a specified target, and edge detection; The edges of the processed mask image are optimized according to the optimization method to obtain the mask image of the first image.
3. The method according to claim 1, characterized in that, The step of determining the target area in the first image that needs image quality processing based on the image information of the first image includes: Obtain the target corpus, which includes image elements that require image quality processing; Based on the language model and the first prompt information, selectable image elements in the first image are determined, and the first prompt information is used to instruct the language model to recognize image elements in the first image; The target image elements are filtered using the target corpus to obtain the target image elements in the first image; Determine the target region where the target image element is located in the first image.
4. The method according to claim 3, characterized in that, Determining the target region where the target image element is located in the first image includes: A second prompt message is generated based on the target image elements. The second prompt message is used to instruct the image segmentation model to segment the target image elements. Based on the image segmentation model and the second prompt information, the first image is segmented into target image elements to obtain the target region.
5. The method according to claim 1, characterized in that, The step of determining the target area in the first image that needs image quality processing based on the image information of the first image includes: Based on the image information of the first image, a specified target is detected in the first image, and the detection result of the specified target is determined; based on the location information in the detection result, a target region in the first image is determined; and / or, Based on the image information of the first image, edge detection is performed on the first image to obtain the edge information of the first image, so as to determine the target region in the first image.
6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: The target region is obtained, along with a description of the global image quality processing of the first image. Semantic processing is performed based on the target region and the processing description to obtain the image processing information; The image processing information is associated with the image pair to obtain the association result of the image pair.
7. An image processing apparatus, characterized in that, The device includes: The image acquisition module is used to acquire a first image of first quality. The quality processing module is used to perform image quality processing on the entire first image to obtain a second image with a second image quality, wherein the first image quality is different from the second image quality. The region determination module is used to determine, based on the image information of the first image, a target region in the first image that needs to undergo image quality processing, wherein the size of the target region is smaller than the size of the first image. A mask processing module is used to obtain a mask image of the first image based on the position information of the target region in the first image and the size of the first image. The pixel value of the target region in the mask image is different from the pixel value of other regions in the mask image. The other regions are regions in the mask image other than the target region. An image fusion module is used to obtain a third image by fusing the first image, the second image, and the mask image; The image pairing module is used to associate the first image with the third image to obtain image pairs with different image qualities.
8. An electronic device, characterized in that, include: A memory and a processor are communicatively connected, the memory stores computer instructions, and the processor executes the image processing method of any one of claims 1 to 6 by executing the computer instructions.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the image processing method according to any one of claims 1 to 6.
10. A computer program product, characterized in that, Includes computer instructions for causing a computer to perform the image processing method according to any one of claims 1 to 6.