Method, apparatus, device, storage medium and program product for image processing

By processing images through multiple units of the model, generating and fusing global, segmentation, and edge feature information, the problem of low image super-resolution quality in existing technologies is solved, achieving high-quality and efficient image super-resolution results.

CN121304449BActive Publication Date: 2026-03-31VASTAI TECH (SHANGHAI) INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing image super-resolution techniques produce super-resolution images of low quality and poor visual effects, making it difficult to meet the needs of high-resolution image processing.

Method used

The model processes images through multiple units, generating global feature information, segmentation feature information, and edge feature information. It then upsamples by fusing these feature information to construct a high-resolution image.

Benefits of technology

It improves the quality and efficiency of image processing, ensures visual coherence and the integrity of boundary structures, clearly distinguishes different content areas, and reduces conflicts between the output content of multiple units.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121304449B_ABST
    Figure CN121304449B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a method, apparatus, device, storage medium and program product for image processing. The method proposed herein includes: encoding a first image using a first unit of a model to generate global feature information, the first image corresponding to a first resolution; processing the first image using a second unit of the model to generate segmentation feature information; processing the first image using a third unit of the model to generate edge feature information; determining fused feature information by fusing the global feature information, the segmentation feature information and the edge feature information; and constructing a second image by up-sampling the fused feature information, the second image corresponding to a second resolution, the second resolution being higher than the first resolution. In this way, embodiments of the present disclosure can analyze the first image from different dimensions using multiple units of the model, so that the quality of image processing can be improved, and the efficiency of image processing can be improved by up-sampling the fused fused feature information after fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to methods, apparatus, devices, computer-readable storage media, and computer program products for image processing. Background Technology

[0002] With the development of computer technology, more and more image applications require higher resolution images. However, due to hardware limitations or transmission limitations, the actual images acquired are usually of limited resolution, making it difficult to meet the needs of image processing. Therefore, techniques such as image super-resolution have emerged. Image super-resolution technology aims to perform super-resolution processing on some lower-resolution images to reconstruct higher-resolution images. How to improve the quality and processing efficiency of image super-resolution is a key concern. Summary of the Invention

[0003] In a first aspect of this disclosure, an image processing method is provided. The method includes: encoding a first image using a first unit of a model to generate global feature information of the first image, the first image corresponding to a first resolution; processing the first image using a second unit of the model to generate segmentation feature information of the first image, the segmentation feature information indicating the distribution of multiple semantic regions in the first image; processing the first image using a third unit of the model to generate edge feature information of the first image, the edge feature information indicating at least one type of edge in the first image; determining fused feature information by fusing the global feature information, the segmentation feature information, and the edge feature information; and constructing a second image by upsampling the fused feature information, the second image corresponding to a second resolution, the second resolution being higher than the first resolution.

[0004] In a second aspect of this disclosure, an apparatus for image processing is provided. The apparatus includes: an encoding module configured to encode a first image using a first unit of a model to generate global feature information of the first image, the first image corresponding to a first resolution; a first processing module configured to process the first image using a second unit of the model to generate segmentation feature information of the first image, the segmentation feature information indicating the distribution of multiple semantic regions in the first image; a second processing module configured to process the first image using a third unit of the model to generate edge feature information of the first image, the edge feature information indicating at least one type of edge in the first image; a fusion module configured to determine fused feature information by fusing the global feature information, the segmentation feature information, and the edge feature information; and a construction module configured to construct a second image by upsampling the fused feature information, the second image corresponding to a second resolution, the second resolution being higher than the first resolution.

[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.

[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.

[0007] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.

[0008] The embodiments of this disclosure can utilize the global feature information of the first image to ensure the visual coherence of the processed image, and can utilize the segmentation feature information of the first image to protect the boundary structure, thereby improving the quality of image processing. Furthermore, the embodiments of this disclosure further distinguish different content regions and boundary regions more clearly by analyzing the distribution of multiple semantic regions in the first image indicated by the segmentation feature information, thereby further improving the quality of image processing. Further, the embodiments of this disclosure reduce conflicts between the output content of multiple units by fusing global feature information, segmentation feature information, and edge feature information, thereby taking into account global content, edge content, and semantic content, thereby further improving the quality of image processing and the quality of constructing the second image. Additionally, the embodiments of this disclosure process the image in multiple dimensions through multiple units of the model, and uniformly upsample after fusing the corresponding outputs, thereby improving image processing efficiency. Thus, the embodiments of this disclosure can utilize multiple units of the model to improve the quality of image processing, and can improve image processing efficiency by upsampling the fused feature information.

[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0011] Figure 1A schematic diagram is shown of an example environment in which embodiments of the present disclosure may be implemented;

[0012] Figure 2 Example structural block diagrams of models according to some embodiments of the present disclosure are shown;

[0013] Figure 3 A flowchart illustrating an example process of image processing according to some embodiments of the present disclosure is shown;

[0014] Figure 4 A flowchart illustrating an example process for edge recognition according to some embodiments of this disclosure is shown;

[0015] Figure 5 A flowchart illustrating an example process of model training according to some embodiments of this disclosure is shown;

[0016] Figure 6 A flowchart illustrating an example process for sample processing according to some embodiments of this disclosure is shown;

[0017] Figure 7 A schematic structural block diagram of an example apparatus for image processing according to some embodiments of the present disclosure is shown; and

[0018] Figure 8 A block diagram of an electronic device capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation

[0019] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0020] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0021] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0022] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.

[0023] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.

[0024] In the description of embodiments of this disclosure, the term "image super-resolution" and similar terms can be understood as an image processing technique designed to perform super-resolution processing on images to construct (or reconstruct) images of higher resolution. Other explicit and implicit definitions may also be included below.

[0025] As mentioned above, with the development of computer technology, more and more image applications require higher resolution images. However, due to hardware limitations or transmission limitations, the actual images acquired are usually of limited resolution, making it difficult to meet the needs of image processing. Therefore, techniques such as image super-resolution have emerged. Image super-resolution technology aims to perform super-resolution processing on some lower-resolution images to reconstruct higher-resolution images. How to improve the quality and processing efficiency of image super-resolution is a key concern.

[0026] Embodiments of this disclosure propose an image processing scheme. According to this scheme, a first image can be encoded using a first unit of a model to generate global feature information of the first image, which corresponds to a first resolution. Further, the first image can be processed using a second unit of the model to generate segmentation feature information of the first image, the segmentation feature information indicating the distribution of multiple semantic regions in the first image. Further, the first image can be processed using a third unit of the model to generate edge feature information of the first image, the edge feature information indicating at least one type of edge in the first image. Additionally, fused feature information can be determined by fusing global feature information, segmentation feature information, and edge feature information. Furthermore, a second image can be constructed by upsampling the fused feature information, the second image corresponding to a second resolution, which is higher than the first resolution.

[0027] Based on this approach, embodiments of this disclosure can utilize the global feature information of the first image to ensure the visual coherence of the processed image, and can utilize the segmentation feature information of the first image to protect the boundary structure, thereby improving the quality of image processing. Furthermore, embodiments of this disclosure also more clearly distinguish different content regions and boundary regions by analyzing the distribution of multiple semantic regions in the first image indicated by the segmentation feature information, thereby further improving the quality of image processing. Further, embodiments of this disclosure reduce conflicts between the output content of multiple units by fusing global feature information, segmentation feature information, and edge feature information, thereby taking into account global content, edge content, and semantic content, further improving the quality of image processing and the quality of constructing the second image. Additionally, embodiments of this disclosure process images in multiple dimensions through multiple units of the model and uniformly upsamples them after fusing the corresponding outputs, thereby improving image processing efficiency.

[0028] Therefore, embodiments of this disclosure can utilize multiple units of the model to improve the quality of image processing and can improve image processing efficiency by upsampling the fused feature information.

[0029] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.

[0030] Example Environment

[0031] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. For example... Figure 1 As shown, example environment 100 may include image processing device 110. Image processing device 110 is configured to perform image super-resolution processing on the image.

[0032] In this example environment 100, the image processing device 110 can use the model 120 to process the first image 130 corresponding to a first resolution to construct a second image 140 corresponding to a second resolution, which is higher than the first resolution. It should be understood that the image content shown in the first image 130 and the second image 140 is for illustrative purposes only and is not intended to be limiting.

[0033] Model 120 can be a suitable image processing model. Alternatively or additionally, model 120 can be a trained model. Model 120 can be deployed on image processing device 110 or other devices (different from image processing device 110), which will not be described in detail here.

[0034] Image processing device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, image processing device 110 can also support any type of user-facing interface (such as "wearable" circuitry).

[0035] The image processing device 110 can also be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. The image processing device 110 may include, for example, computing systems / servers such as mainframes, edge computing nodes, and computing devices in a cloud environment, etc.

[0036] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0037] Currently, some traditional image super-resolution techniques can perform super-resolution processing on images to reconstruct their corresponding higher resolution images. Examples include bicubic interpolation and general super-resolution methods based on convolutional neural networks. However, the super-resolution images constructed by these techniques are of low quality and have poor visual effects.

[0038] In view of this, embodiments of the present disclosure provide an image processing method to improve the quality of image super-resolution. The following will be described in conjunction with the appendix. Figure 2 This document systematically describes example image processing procedures for some embodiments of the present disclosure. Figure 2 An example structural block diagram of model 120 according to some embodiments of the present disclosure is shown. Reference is made below. Figure 1 This disclosure describes some example image processing procedures.

[0039] like Figure 2 As shown, model 120 may include, for example, an image input unit 211, a multi-branch processing unit 221, an original image branch 231 (also referred to as, the first unit), a semantic segmentation branch 241 (also referred to as, the second unit), an edge branch 251 (also referred to as, the third unit), an adaptive fusion unit 261, an image output unit 271, etc. It should be understood that in some examples, model 120 may include some of the units described above, but not necessarily all of them. Alternatively or additionally, model 120 may also include other units besides those described above. The embodiments of this disclosure are only intended to illustrate... Figure 2 The structural block diagram shown is used as an example to illustrate some image processing procedures of this disclosure, but is not intended to limit it.

[0040] In some embodiments, the image processing device 110 can input the first image 130 into the model 120 via the image input unit 211. Furthermore, the image processing device 110 can utilize the multi-branch processing unit 221 to distribute the input first image 130 to the original image branch 231, the semantic segmentation branch 241, and the edge branch 251 to support parallel computation of the original image branch 231, the semantic segmentation branch 241, and the edge branch 251, thereby improving data throughput and processing efficiency.

[0041] In some embodiments, the image processing device 110 may encode the first image 130 using the original image branch 231 to generate global feature information 232 corresponding to the first image 130. Additionally, the image processing device 110 may process the first image 130 using the semantic segmentation branch 241 to generate segmentation feature information 242 for the first image 130. Additionally, the image processing device 110 may process the first image 130 using the edge branch 251 to generate edge feature information 252 for the first image 130.

[0042] Furthermore, the image processing device 110 can utilize the adaptive fusion unit 261 to fuse global feature information 232, segmentation feature information 242, and edge feature information 252 to determine fused feature information. Additionally, the image processing device 110 can utilize the image output unit 271 to upsample such fused feature information to construct a higher resolution second image 140. Thus, embodiments of this disclosure can utilize multiple units (or branches) of the model to improve image processing quality and can improve image processing efficiency by uniformly upsampling the fused feature information.

[0043] The following description will continue with reference to the accompanying drawings to further describe some exemplary embodiments of this disclosure.

[0044] Example image processing procedure

[0045] The following will be referenced Figure 3 To describe some example processes of image processing according to embodiments of the present disclosure. Figure 3 A flowchart of an example process 300 for image processing according to some embodiments of the present disclosure is shown. Process 300 can be implemented at an image processing device 110. Reference is made below. Figure 1 Describe the process 300.

[0046] like Figure 3 As shown, in block 310, the image processing device 110 uses the first unit of the model to encode the first image to generate global feature information of the first image, which corresponds to the first resolution.

[0047] As an example, such a model could be like... Figure 1The model 120 shown is an example of a trained model. Its training process can be found in the example model training process described below, and will not be repeated here. As an example, the first image can be an image of any format, which may include any suitable image content. The first image may correspond to a first resolution, which may be smaller than the second resolution of the second image described below. For example, such a first resolution may also be called low resolution (LR), and such a second resolution may also be called high resolution (HR). Thus, embodiments of this disclosure can process low-resolution images (e.g., the first image) to construct high-resolution images (e.g., the second image).

[0048] In some examples, the first image can be an image with a preset style, such as an anime-style image. For example, the first image can include animation content, game content, and so on. Compared to images with a photorealistic style or natural images, such a first image can have higher contrast edges, smoother color blocks, more continuous lines, more complex textures, more repeated textures, and so on.

[0049] In some examples, such a first unit may be, for example, the original image branch 231 mentioned above, which may be any suitable coding network, such as a convolutional neural network structure or other suitable feature extraction network. For example, such a first unit can extract global feature information at multiple scales (or multiple levels) in the first image. Taking a convolutional neural network as an example, such a first unit may include multiple convolutional layers, residual connection layers, activation functions, etc., and can extract global feature information of such a first image through forward propagation.

[0050] As an example, such global feature information may include, for instance, the global feature information 232 mentioned above. This global feature information can indicate the overall contextual information of the first image, and may be in the form of feature vectors, feature maps, etc. As an example, such global feature information may include the color distribution, texture pattern, global structure, etc., of the first image to support the model in understanding the more basic image content of the first image from an overall perspective. Therefore, embodiments of this disclosure allow the model to refer to the overall information of the first image during the subsequent construction of the second image, thereby improving the color consistency and overall visual coherence of the constructed second image.

[0051] In box 320, the image processing device 110 uses the second unit of the model to process the first image to generate segmentation feature information of the first image, the segmentation feature information indicating the distribution of multiple semantic regions in the first image.

[0052] As an example, such a second unit could be the semantic segmentation branch 241 described above. Such a second unit could, for example, output multiple image regions based on the input image. In some examples, such a second unit could process such a first image based on semantic information.

[0053] As an example, the image processing device 110 can use the second unit to analyze semantic regions in the first image and can generate segmentation feature information. As an example, such segmentation feature information can be a semantic label map, which can indicate the distribution of multiple (or one) semantic regions in the first image.

[0054] As an example, such segmentation feature information can be used to identify the location, shape, and other information of regions containing image content with different semantic meanings in the first image. Therefore, embodiments of this disclosure can utilize such segmentation feature information to guide the model in distinguishing different semantic regions during image super-resolution, thereby improving the semantic consistency in the subsequently constructed second image.

[0055] Additionally, if the first image includes image content such as animation content or game content, such a second unit can also be configured to segment content such as color blocks with color differences below a threshold and edges with sharpness above a threshold, in order to improve segmentation accuracy and thereby further improve the quality of constructing the second image.

[0056] Alternatively or additionally, such a second unit may also segment such a first image based on other information. For example, such a second unit may also segment such a first image based on non-semantic information to obtain such segmentation feature information. As an example, such non-semantic information differs from the semantic information described above, and may include, for example, color information, entity information, depth information, etc. Its specific implementation can be referred to the specific implementation of semantic segmentation of the first image described above, and will not be repeated in the embodiments of this disclosure.

[0057] In some embodiments, such a second unit may include a semantic segmentation unit and a semantic recognition unit. Further, the image processing device 110 may utilize such a semantic segmentation unit to segment the first image to determine multiple segmented regions corresponding to the first image. Additionally, the image processing device 110 may utilize the semantic recognition unit to recognize semantic information of the multiple segmented regions to generate segmentation feature information of the first image.

[0058] In some examples, such semantic segmentation units can include any suitable segmentation network, such as the SAM (Segment Anything Model). To improve the quality of semantic segmentation, such semantic segmentation units can be obtained by fine-tuning a pre-trained segmentation network using a pre-defined dataset (e.g., a sample dataset including animation content, game content, etc.).

[0059] As an example, the image processing device 110 can utilize a fine-tuned semantic segmentation network to segment such a first image and obtain a semantic segmentation result (e.g., a region mask), which may include multiple segmented regions corresponding to the first image. As an example, different segmented regions among the multiple segmented regions may correspond to different categories. For instance, taking a first image comprising a river and a forest as an example, the region corresponding to the river and the region corresponding to the forest may correspond to different segmented regions. Thus, embodiments of this disclosure can further assist the model in understanding the image structure of the first image.

[0060] In some examples, such a semantic recognition unit may include any suitable classification network, which, for example, is capable of outputting multiple semantic labels based on multiple segmented regions of the input, and further labeling such multiple semantic labels onto the corresponding segmented regions. Continuing with the previous example, the segmented region where the river is located can correspond to the label "river," while the segmented region where the forest is located can correspond to the label "forest." Thus, the embodiments of this disclosure can help the model, based on its understanding of the image structure of the first image, further understand the different segmented regions in the first image, thereby further improving the quality of the subsequent construction of the second image.

[0061] In some embodiments, to further improve the accuracy of the model's understanding of the first image, the image processing device 110 can highlight specified semantic regions in such segmentation feature information. Specifically, the image processing device 110 can utilize a label classifier in the semantic recognition unit to identify multiple segmentation regions to determine multiple semantic labels corresponding to the multiple segmentation regions.

[0062] As an example, such multiple semantic labels can be used to determine the weight information of multiple segmentation regions. For instance, the image processing device 110 can map the weight information corresponding to the multiple semantic labels to the corresponding multiple segmentation regions to determine the weight information of the multiple segmentation regions.

[0063] In some examples, the weights of multiple semantic labels can be determined by any appropriate weight calculation network, or by preset weight ratios, weight scores, evaluation information, etc. For example, the weights of entities such as characters and objects can exceed the weights of the background, and the weights of image content with less depth can exceed the weights of image content with greater depth, and so on.

[0064] Additionally, the image processing device 110 can utilize multiple semantic tags to label multiple segmentation regions in the first image to generate segmentation feature information for the first image. As an example, such segmentation feature information can also indicate weight information corresponding to the multiple segmentation regions.

[0065] Therefore, the embodiments of this disclosure can assign different weight information to different image content in the first image, thereby improving the transition quality of different semantic regions during the construction of the second image and reducing the degree of content distortion.

[0066] In box 330, the image processing device 110 processes the first image using the third unit of the model to generate edge feature information of the first image, the edge feature information indicating at least one type of edge in the first image.

[0067] As an example, such a third unit could be the edge branch 251 mentioned above. Such a third unit could be any suitable edge processing network, capable of recognizing edges in the input image and outputting corresponding edge feature information. As an example, the edges of an image could be the edges between different regions presenting different content in the image, also known as boundaries, such as the outline of a person, the boundary of a lake, riverbank lines, etc. As an example, different styles of edges could be different types of edges; for example, straight edges and curved edges could be different types of edges.

[0068] In some examples, the image processing device 110 can utilize the third unit of the model to identify or extract edges between different image regions in the first image and generate edge feature information of the first image. Such edge feature information can indicate the boundary structure, edge distribution, etc., of the image regions. Therefore, embodiments of this disclosure can further improve the accuracy of the model in distinguishing different image content in the first image.

[0069] In some embodiments, the image processing device 110 may employ different edge recognition operators to identify or extract edges of different types. Specifically, the image processing device 110 may use at least one edge recognition operator to identify a first image to determine at least one edge type in the first image corresponding to the at least one edge recognition operator.

[0070] In some examples, the edge recognition operator can be any suitable feature detection tool capable of detecting or extracting edges of the corresponding edge type in the first image. For example, such an edge recognition operator can be in the form of a convolution kernel, capable of identifying whether edges of the corresponding edge type exist in the first image and outputting a corresponding processing result that reflects the edge of the corresponding edge type. In some examples, such at least one edge recognition operator can include at least one of the following: a straight line operator, an arc operator, a corner operator, a continuous curve operator, and a texture operator.

[0071] As an example, such a line operator can be configured to recognize line edges at a preset angle in a first image. For example, such a line operator may include horizontal operators, vertical operators, etc., which can, for example, recognize edges such as horizontal edges and vertical edges in the first image.

[0072] As an example, an arc operator can be configured to identify arc edges with a preset curvature in a first image. Such an arc operator can include, for example, a smooth arc operator, a small-radian curve operator, etc. For instance, a smooth arc operator can be used to identify arcs in the first image with curvature within a first curvature range, such as facial contours, pupil contours, etc. Similarly, a small-radian curve operator can be used to identify arcs in the first image with curvature within a second preset curvature range, such as the boundaries of hair tips, feather boundaries, etc. In some examples, the curvature of the arc to be identified by such a small-radian curve operator can be greater than the curvature of the arc to be identified by the smooth arc operator.

[0073] As an example, an angle operator can be configured to identify corner edges in the first image at preset angles, such as bends, sharp corners, etc. A continuous curve operator can be configured to identify curve edges in the first image that satisfy the continuous curve type. This can, for example, analyze the trend of the entire curve globally to reduce the probability of line breaks in the constructed second image. For example, curve edges satisfying the continuous curve type can include the overall edge of a mountain, the entire riverbank, etc. As an example, a texture operator can be configured to identify texture edges in the first image at preset frequencies. As an example, the preset frequency can indicate the frequency of local grayscale or color changes in the image. For example, texture edges with preset frequencies can include high-frequency textures (textures with frequencies higher than a third threshold), sparse textures (textures with frequencies lower than a fourth threshold), etc.

[0074] In some examples, the image processing device 110 can merge at least one edge recognition operator into a single efficient network through reparameterized fusion, thereby reducing computational overhead and improving image processing efficiency. Furthermore, the image processing device 110 can utilize this efficient network to determine the processing result of the at least one edge recognition operator.

[0075] Additionally, the image processing device 110 can generate edge feature information of the first image based on the processing result of at least one edge recognition operator. As an example, such edge feature information can indicate edges of at least one edge type recognized by the at least one edge recognition operator.

[0076] Taking at least one edge recognition operator, including straight line operator, smooth arc operator, small-radius curve operator, corner operator, continuous curve operator, high-frequency texture operator, and sparse texture operator, as an example, the following will refer to... Figure 4 The following describes some example edge recognition processes in embodiments of this disclosure. Figure 4 A flowchart of an example process 400 for edge recognition according to some embodiments of the present disclosure is shown. Process 400 can be implemented in Figure 2 There is a 251-point edge branch in the middle.

[0077] like Figure 4 As shown, the image processing device 110 can utilize the multi-operator branch 410 in the third unit to distribute the input first image 130 to multiple processing units to identify edges of different edge types, such as straight line operator processing unit 421, smooth arc operator processing unit 422, small-radian curve operator processing unit 423, angle operator processing unit 424, continuous curve operator processing unit 425, high-frequency texture operator processing unit 426, sparse texture operator processing unit 427, etc. Furthermore, the image processing device 110 can utilize the reparameterization fusion unit 430 to fuse the processing results of at least one such processing unit to generate edge feature information 252. Thus, the embodiments of this disclosure can obtain richer and more accurate edge feature information.

[0078] Based on this, embodiments of the present disclosure can improve the integrity of the edge structure in the constructed second image, thereby improving the quality of lines in the second image. Furthermore, for a first image including animation content and game content, through at least one such edge recognition operator, embodiments of the present disclosure can further ensure the edge sharpness and edge structure integrity of the constructed second image, thereby avoiding line breakage and structural distortion problems during the super-resolution process of images including animation content and game content, and further improving the quality of the constructed second image.

[0079] In frame 340, the image processing device 110 determines the fused feature information by fusing global feature information, segmentation feature information, and edge feature information.

[0080] As an example, the image processing device 110 can utilize any suitable fusion network (e.g., the adaptive fusion unit 261 described above) to fuse such global feature information, segmentation feature information, and edge feature information to determine fused feature information. As an example, such fused feature information can combine multiple dimensions such as detail, edges, and semantics. As an example, such a fusion network can fuse such global feature information, segmentation feature information, and edge feature information based on fixed weights.

[0081] As an example, such global feature information, segmentation feature information, and edge feature information can have the same spatial size, channel dimension, and other dimensions. Such spatial size can correspond to a first resolution, and the spatial size of the fused feature information determined therefrom can also correspond to such a first resolution.

[0082] Alternatively, such a fusion network may include, for example, an appropriate adaptive weight network (also known as a weight unit), which may fuse such global feature information, segmentation feature information, and edge feature information based on dynamic weights. Such dynamic weights may be determined by the appropriate adaptive weight network through analysis of such global feature information, segmentation feature information, and edge feature information.

[0083] Specifically, the image processing device 110 can use weight units to determine weight information corresponding to global feature information, segmentation feature information, and edge feature information. Additionally, the image processing device 110 can fuse global feature information, segmentation feature information, and edge feature information based on the weight information to determine fused feature information.

[0084] As an example, such weighting units can include convolutional neural networks, attention subnetworks, etc. They can, for example, compute weight information (e.g., weight maps) corresponding to global feature information, segmentation feature information, and edge feature information.

[0085] Furthermore, the image processing device 110 can utilize such weight units to analyze global feature information, segmentation feature information, and edge feature information, thereby dynamically adjusting the fusion ratio of global feature information, segmentation feature information, and edge feature information based on the corresponding weight information. For example, higher weight can be assigned to edge feature information for regions with rich edges. Or, for example, higher weight can be assigned to global feature information for flat color blocks, and so on.

[0086] Therefore, the embodiments of this disclosure can support the weight units in optimizing the fusion strategy (e.g., fusion ratio) based on image content through the learnability of the weight units, thereby outputting fused feature information. Such fused feature information can avoid conflicts between multiple feature information and enable the process of constructing the second image to take into account multiple dimensions of image information such as details, edges, and semantics, thereby further improving the quality of the constructed second image.

[0087] In frame 350, the image processing device 110 constructs a second image by upsampling and fusing feature information. The second image corresponds to a second resolution, which is higher than the first resolution.

[0088] As an example, such an upsampling process can be implemented in the image output unit 271 described above. As an example, the image processing device 110 can implement such an upsampling process using upsampling tools such as subpixel convolution and transposed convolution. For example, the image processing device 110 can use upsampling tools such as subpixel convolution and transposed convolution to construct a higher-resolution second image from the fused feature information. Therefore, embodiments of this disclosure can construct a higher-quality second image through more comprehensive fused feature information.

[0089] Based on this approach, embodiments of this disclosure can utilize the global feature information of the first image to ensure the visual coherence of the processed image, and can utilize the segmentation feature information of the first image to protect the boundary structure, thereby improving the quality of image processing. Furthermore, embodiments of this disclosure also more clearly distinguish different content regions and boundary regions by analyzing the distribution of multiple semantic regions in the first image indicated by the segmentation feature information, thereby further improving the quality of image processing. Further, embodiments of this disclosure reduce conflicts between the output content of multiple units by fusing global feature information, segmentation feature information, and edge feature information, thereby taking into account global content, edge content, and semantic content, further improving the quality of image processing and the quality of constructing the second image. Additionally, embodiments of this disclosure process images in multiple dimensions through multiple units of the model and uniformly upsamples them after fusing the corresponding outputs, thereby improving image processing efficiency.

[0090] Therefore, embodiments of this disclosure can utilize multiple units of the model to improve the quality of image processing and can improve image processing efficiency by upsampling the fused feature information.

[0091] The following description, with reference to the accompanying drawings, will continue to illustrate example training processes for models in some embodiments of this disclosure.

[0092] Example model training process

[0093] The following will be combined with the appendix Figure 5 The example model training process of a model according to some embodiments of the present disclosure will be further described. Figure 5 A flowchart illustrating an example process 500 for model training according to some embodiments of the present disclosure is shown. Process 500 may be implemented at image processing device 110 or other devices, and the model trained by process 500 may be, for example, Figure 1 The model 120 is shown below. The following will use training model 120 at image processing device 110 as an example, and refer to... Figure 1 The process 500 is described below. It should be understood that the embodiments disclosed herein are not intended to be limited thereto.

[0094] like Figure 5 As shown, at box 510, the image processing device 110 uses the first unit of the model to encode the first sample image to generate the predicted global feature information corresponding to the first sample image.

[0095] At box 520, the image processing device 110 uses the second unit of the model to process the first sample image to generate predictive segmentation feature information of the first sample image.

[0096] At box 530, the image processing device 110 uses the third unit of the model to process the first sample image to generate predicted edge feature information of the first sample image.

[0097] As an example, the specific implementation of using the first unit, second unit, and third unit of the model to process the first sample image and generate corresponding predictive feature information can be referred to the specific implementation of using the first unit, second unit, and third unit to process the first image to generate corresponding feature information in the above model application process. The embodiments of this disclosure will not be repeated here.

[0098] Furthermore, at box 540, the image processing device 110 trains a model based on predicted global feature information, predicted segmentation feature information, predicted edge feature information, and the second sample image.

[0099] In some examples, the image processing device 110 can independently train the first, second, and third units of the model based on predicted global feature information, predicted segmentation feature information, predicted edge feature information, and the second sample image, thereby reducing computational cost and improving processing efficiency. For example, the image processing device 110 can train the first unit of the model based on predicted global feature information and the second sample image. As another example, the image processing device 110 can train the second unit of the model based on predicted segmentation feature information and the second sample image. And yet another example, the image processing device 110 can train the third unit of the model based on predicted edge feature information and the second sample image.

[0100] In some embodiments, in order to improve the training quality of the model, the image processing device 110 may first train the second unit of the model, and refer to the segmentation results of the second unit during the training of the first and third units of the model, so as to optimize the model in a targeted manner through different semantic regions.

[0101] As an example, the image processing device 110 can adjust the parameters of the second unit based on a second loss function determined by the predicted segmentation feature information and the second sample image. As an example, such a second loss function can indicate the difference between the predicted segmentation feature information and the second sample image.

[0102] In some examples, the image processing device 110 can compare the differences between predicted segmentation feature information and the second sample image in the feature dimension. For example, the image processing device 110 can use the second unit to determine the ground truth segmentation feature information corresponding to the downsampling result of the second sample image, and adjust the parameters of the second unit by the difference between the predicted segmentation feature information and the ground truth segmentation feature information.

[0103] In other examples, the image processing device 110 may also compare the differences between the predicted segmentation feature information and the second sample image in the image dimension. For example, the image processing device 110 may use the predicted segmentation feature information to upsample and reconstruct the predicted image, so as to adjust the parameters of the second unit based on the differences between the predicted image and the second sample image. Thus, embodiments of this disclosure can improve the model's semantic understanding ability of the input image by training the second unit of the model.

[0104] Additionally, the image processing device 110 can adjust the parameters of the first unit based on a first loss function determined by predicting global feature information, predicting segmentation feature information, and the second sample image. Additionally, the image processing device 110 can adjust the parameters of the third unit based on a third loss function determined by predicting edge feature information, predicting segmentation feature information, and the second sample image.

[0105] As an example, the image processing device 110 can utilize predicted segmentation feature information to guide the training process of the first unit and / or the third unit, enabling the first unit and / or the third unit to possess region-aware capabilities. Taking the image processing device 110 using predicted segmentation feature information to guide the training process of the first unit as an example, the image processing device 110 can divide different semantic regions based on the predicted segmentation feature information and calculate independent loss functions in different semantic regions to independently train the first unit. Therefore, embodiments of this disclosure can further improve the model's ability to distinguish between different semantic regions, thereby further reducing the degree of edge blurring, color bleeding, etc., during image super-resolution, and improving the quality of the trained model.

[0106] Alternatively or additionally, the image processing device 110 can holistically adjust the model's parameters (also known as joint fine-tuning) based on predicted global feature information, predicted segmentation feature information, predicted edge feature information, and the second sample image to improve the model's overall consistency and overall training quality. Specifically, the image processing device 110 can co-train the first, second, and third units in the model based on the target loss function determined by the first, second, and third loss functions.

[0107] As an example, such a target loss function can be determined based on a weighted sum of a first loss function, a second loss function, and a third loss function, for use in adjusting the parameters of the overall model (e.g., adjusting the parameters of the first unit, the second unit, and / or the third unit). Therefore, embodiments of this disclosure can further improve the quality of model training.

[0108] In some embodiments, the image processing device 110 may further train the fusion unit (e.g., Figure 2 The adaptive fusion unit 261 in the model improves the accuracy of the weight information referenced by the subsequent model in the process of fusing global feature information, edge feature information and segmentation feature information.

[0109] As an example, the image processing device 110 can input the predicted global feature information, the predicted edge feature information, and the predicted segmentation feature information into the weight unit to obtain the predicted weight information. Then, the image processing device 110 can fuse the predicted global feature information, the predicted edge feature information, and the predicted segmentation feature information based on the predicted weight information to obtain the predicted fusion feature.

[0110] Furthermore, the image processing device 110 can adjust the parameters of the weight unit based on the loss between the predicted fusion features and the second sample image. As an example, the image processing device 110 can adjust the parameters of the weight unit using the loss between the predicted image obtained from the upsampled predicted fusion features and the second sample image. Alternatively, the image processing device 110 can adjust the parameters of the weight unit using the loss between the predicted fusion features and the ground truth fusion features corresponding to the image obtained from the downsampled second sample image.

[0111] Based on this approach, the embodiments of this disclosure can improve the training quality of the model and enable the trained model to have image understanding capabilities in multiple dimensions, including semantics, regions, and edges, thereby improving the image processing quality of subsequent models.

[0112] In some embodiments, the trained model can be used for image processing, specifically for image super-resolution processing. The specific implementation can be referred to the example image processing process described above, and the embodiments of this disclosure will not be repeated here.

[0113] In some embodiments, such a model can be trained using a preset dataset, which may include pairs of sample images, each pair corresponding to a different resolution. For example, such a pair of sample images may consist of a low-resolution image (meeting a preset low-resolution range) and a high-resolution image (meeting a preset high-resolution range). For instance, such a pair of sample images could be the first sample image and the second sample image mentioned above.

[0114] To improve model training quality, in some embodiments, such a pre-processed dataset is preprocessed for model training based on the pre-processed dataset. As an example, the paired sample images in the pre-processed dataset may be a first sample image and a third sample image, where the first sample image has a lower resolution than the third sample image. Furthermore, the pre-processed dataset may include the aforementioned first sample image and a second sample image obtained by processing the third sample image.

[0115] In some embodiments, the image processing device 110 may determine fitted edges corresponding to edge content in the third sample image. Additionally, the image processing device 110 may use the fitted edges to update the edge content in the third sample image to determine the second sample image.

[0116] As an example, such edge content can indicate the boundaries between different image regions in the third sample image, which can, for example, separate different image regions in the third sample image. As an example, the fitted edge can be a fitted edge content corresponding to the edge content, which, compared to the edge content in the third sample image, can, for example, be more continuous and smoother. Based on this, the image processing device 110 can use the higher-quality fitted edge to update the edge content in the third sample image, thereby obtaining a second sample image with clearer edges.

[0117] In some embodiments, to improve the accuracy of edge processing, the image processing device 110 may segment such a third sample image based on semantic information to obtain such edge content. Specifically, the image processing device 110 may segment the third sample image using a semantic segmentation model to determine multiple regions corresponding to the third sample image. Additionally, the image processing device 110 may fit the edge content of the multiple regions to determine fitted edges corresponding to the edge content.

[0118] As an example, such a semantic segmentation model may include, for example, the segmentation network described above, or other suitable image processing networks, which may segment the input image and output multiple corresponding regions (also known as image regions).

[0119] As an example, image processing device 110 can utilize any suitable fitting prediction tool to fit the edge content of such multiple regions to determine the fitted edges corresponding to the edge content. As an example, such semantic segmentation networks and fitting prediction tools can be implemented using any suitable model or tool, and embodiments of this disclosure are not intended to limit such segmentation networks or fitting prediction tools. Thus, image processing device 110 can determine more reasonable multiple regions by more accurately segmenting such a third sample image, thereby improving the accuracy of edge content recognition.

[0120] In some embodiments, the image processing device 110 may use fitted edges to directly replace such edge content to determine such a second sample image. For example, the image processing device 110 may delete edge content in a third sample image and fill in such fitted edges at corresponding positions to determine such a second sample image.

[0121] To improve the fitting quality, in some embodiments, the image processing device 110 may first blur such edge content and then further fill such fitted edges into the corresponding positions. Specifically, the image processing device 110 may blur a preset region in the third sample image to determine the fourth sample image. Additionally, the image processing device 110 may fill the fitted edges into a preset region in the fourth sample image to determine the second sample image.

[0122] As an example, such a preset region can be configured to represent edge content in a third sample image. Thus, the image processing device 110 can erase image content in the preset region of the third sample image. Further, the image processing device 110 can generate blurred content using a neighborhood blurring method and fill such blurred content into the preset region to determine the fourth sample image. Therefore, embodiments of this disclosure can ensure content coherence between different regions.

[0123] Furthermore, the image processing device 110 can fill the fitted edges into a preset region in the fourth sample image to further improve the continuity and smoothness of the edges in the determined second sample image. Thus, embodiments of this disclosure can improve the quality of edges in the second sample image, thereby enhancing the edge structure understanding capability of the model trained based on the second sample image.

[0124] The following will combine Figure 6 This example describes the processing procedure for such a preset dataset. Figure 6A flowchart of an example process 600 for sample processing according to some embodiments of the present disclosure is shown. Process 600 can be implemented at image processing device 110 or other devices, and the second sample image processed by process 600 can be used to train the model described above (e.g., model 120). The following example uses process 600 implemented at image processing device 110 as an illustration, and references... Figure 1 The process 600 is described below. It should be understood that the embodiments disclosed herein are not intended to be limited thereto.

[0125] In some embodiments, the image processing device 110 may perform segmentation processing on the third sample image 610 to obtain a segmentation result 620, thereby dividing the third sample image 610 into different semantic regions. For example, the third sample image 610 may include region 621, region 622, and region 623. Further, the image processing device 110 may perform fitting prediction on the edges at the boundaries between different regions in the third sample image 610 to obtain more continuous and smoother curves as fitting edges.

[0126] Additionally, the image processing device 110 can use fitted edges to update the third sample image 610 to determine the second sample image. In some examples, the image processing device 110 can fill the fitted edges into the third sample image 610 to determine the second sample image.

[0127] In other examples, image processing device 110 may erase the original edge pixels (e.g., image content in a preset area) in the third sample image 610 and fill the erased gaps (e.g., the preset area) using a neighborhood blurring method to obtain a fourth sample image 630. Further, image processing device 110 may fill the fourth sample image 630 with fitted edges to determine such a second sample image (e.g., second sample image 640). Thus, embodiments of this disclosure can improve the continuity and smoothness of edges in the second sample image.

[0128] In this way, the embodiments of this disclosure can improve the quality of the model, thereby improving the quality of subsequent image processing based on the model.

[0129] Example devices and equipment

[0130] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 7 A schematic structural block diagram of an example apparatus 700 for image processing according to certain embodiments of the present disclosure is shown. Apparatus 700 may be implemented as or included in image processing device 110. Various modules / components in apparatus 700 may be implemented by hardware, software, firmware, or any combination thereof.

[0131] like Figure 7 As shown, the apparatus 700 includes an encoding module 710 configured to encode a first image using a first unit of a model to generate global feature information of the first image, the first image corresponding to a first resolution; a first processing module 720 configured to process the first image using a second unit of a model to generate segmentation feature information of the first image, the segmentation feature information indicating the distribution of multiple semantic regions in the first image; a second processing module 730 configured to process the first image using a third unit of a model to generate edge feature information of the first image, the edge feature information indicating at least one type of edge in the first image; a fusion module 740 configured to determine fused feature information by fusing global feature information, segmentation feature information, and edge feature information; and a construction module 750 configured to construct a second image by upsampling the fused feature information, the second image corresponding to a second resolution, the second resolution being higher than the first resolution.

[0132] In some embodiments, the second unit of the model includes a semantic segmentation unit and a semantic recognition unit, and the first processing module 720 is further configured to: segment the first image using the semantic segmentation unit to determine multiple segmentation regions corresponding to the first image; and recognize the semantic information of the multiple segmentation regions using the semantic recognition unit to generate segmentation feature information of the first image.

[0133] In some embodiments, the first processing module 720 is further configured to: identify multiple segmentation regions using a label classifier in the semantic recognition unit to determine multiple semantic labels corresponding to the multiple segmentation regions; and label multiple segmentation regions in the first image using the multiple semantic labels to generate segmentation feature information of the first image, wherein the segmentation feature information also indicates weight information corresponding to the multiple segmentation regions, and the weight information is determined based on the multiple semantic labels.

[0134] In some embodiments, the third unit of the model corresponds to at least one edge recognition operator, and the second processing module 730 is further configured to: recognize the first image using at least one edge recognition operator to determine at least one edge type in the first image corresponding to at least one edge recognition operator; and generate edge feature information of the first image based on the processing result of at least one edge recognition operator.

[0135] In some embodiments, at least one edge recognition operator includes at least one of the following: a straight line operator configured to recognize straight line edges at a preset angle in a first image; an arc operator configured to recognize arc edges with a preset curvature in a first image; an angle operator configured to recognize angle edges with a preset angle in a first image; a continuous curve operator configured to recognize curved edges in a first image that satisfy the continuous curve type; and a texture operator configured to recognize texture edges with a preset frequency in a first image.

[0136] In some embodiments, the fusion module 740 is further configured to: use a weighting unit to determine weight information corresponding to global feature information, segmentation feature information and edge feature information; and based on the weight information, fuse global feature information, segmentation feature information and edge feature information to determine fused feature information.

[0137] In some embodiments, the apparatus 700 further includes a data processing module configured to: determine a third sample image corresponding to a first sample image, wherein the resolution of the first sample image is lower than the resolution of the third sample image; determine a fitted edge corresponding to edge content in the third sample image; and update the edge content in the third sample image using the fitted edge to determine a second sample image.

[0138] In some embodiments, the data processing module is further configured to: segment a third sample image using a semantic segmentation model to determine multiple regions corresponding to the third sample image; and fit the edge content of the multiple regions to determine the fitted edges corresponding to the edge content.

[0139] In some embodiments, the data processing module is further configured to: blur a preset region in a third sample image to determine a fourth sample image, wherein the preset region is configured to present edge content in the third sample image; and fill the preset region in the fourth sample image with fitted edges to determine a second sample image.

[0140] In some embodiments, the apparatus 700 further includes a model training module configured to: encode a first sample image using a first unit of the model to generate predicted global feature information corresponding to the first sample image; process the first sample image using a second unit of the model to generate predicted segmentation feature information of the first sample image; process the first sample image using a third unit of the model to generate predicted edge feature information of the first sample image; and train the model based on the predicted global feature information, the predicted segmentation feature information, the predicted edge feature information, and the second sample image.

[0141] In some embodiments, the model training module is further configured to: adjust the parameters of the first unit based on a first loss function determined by predicted global feature information, predicted segmentation feature information, and the second sample image; adjust the parameters of the second unit based on a second loss function determined by predicted segmentation feature information and the second sample image; and adjust the parameters of the third unit based on a third loss function determined by predicted edge feature information, predicted segmentation feature information, and the second sample image.

[0142] In some embodiments, the model training module is further configured to: co-train the first unit, the second unit, and the third unit in the model based on a target loss function determined by the first loss function, the second loss function, and the third loss function.

[0143] In some embodiments, the first image corresponds to animated content or game content.

[0144] The modules included in device 700 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the modules in device 700 can be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chips (SoCs), complex programmable logic devices (CPLDs), and so on.

[0145] Figure 8 A block diagram of an electronic device 800 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 8 The electronic device 800 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 8 The electronic device 800 shown can be used to achieve Figure 1 Image processing device 110 or Figure 7 Device 700.

[0146] like Figure 8As shown, electronic device 800 is in the form of a general-purpose electronic device. Components of electronic device 800 may include, but are not limited to, at least one processor 810 or processing unit, memory 820, storage device 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860. Processor 810 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 820. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 800.

[0147] Electronic device 800 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 820 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 830 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 800.

[0148] Electronic device 800 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 8 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 820 may include computer program product 825 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0149] The communication unit 840 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 800 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 800 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or another network node.

[0150] Input device 850 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 860 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 800 can also communicate with one or more external devices (not shown) via communication unit 840 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 800, or with any device that enables electronic device 800 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0151] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0152] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0153] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0154] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0155] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0156] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method of image processing, characterized by, The method comprises: encoding a first image by a first unit of a model to generate global feature information of the first image, the first image corresponding to a first resolution; processing the first image by a second unit of the model to generate segmentation feature information of the first image, the segmentation feature information indicating distribution of a plurality of semantic regions in the first image; processing the first image by a third unit of the model to generate edge feature information of the first image, the edge feature information indicating at least one type of edge in the first image; determining fusion feature information by fusing the global feature information, the segmentation feature information and the edge feature information; and constructing a second image by up-sampling the fusion feature information, the second image corresponding to a second resolution, the second resolution being higher than the first resolution, wherein the model is obtained by training based on a first sample image and a second sample image, the second sample image being determined by: determining a third sample image corresponding to the first sample image, the first sample image having a resolution lower than that of the third sample image; determining a fitted edge corresponding to edge content in the third sample image; obtaining a fourth sample image by performing blur processing on a preset region in the third sample image, wherein the preset region is configured to present the edge content in the third sample image; and filling the fitted edge into the preset region in the fourth sample image to determine the second sample image.

2. The method of claim 1, wherein, The second unit of the model comprises a semantic segmentation unit and a semantic recognition unit, and processing the first image by the second unit of the model to generate the segmentation feature information of the first image comprises: segmenting the first image by the semantic segmentation unit to determine a plurality of segmentation regions corresponding to the first image; and recognizing semantic information of the plurality of segmentation regions by the semantic recognition unit to generate the segmentation feature information of the first image.

3. The method of claim 2, wherein, Recognizing semantic information of the plurality of segmentation regions by the semantic recognition unit to generate the segmentation feature information of the first image comprises: recognizing the plurality of segmentation regions by a label classifier in the semantic recognition unit to determine a plurality of semantic labels corresponding to the plurality of segmentation regions; and labeling the plurality of segmentation regions in the first image by the plurality of semantic labels to generate the segmentation feature information of the first image, wherein the segmentation feature information further indicates weight information corresponding to the plurality of segmentation regions, the weight information being determined based on the plurality of semantic labels.

4. The method of claim 1, wherein, The third unit of the model corresponds to at least one edge recognition operator, and processing the first image by the third unit of the model to generate the edge feature information of the first image comprises: recognizing the first image by the at least one edge recognition operator to determine at least one type of edge in the first image corresponding to the at least one edge recognition operator; and generate the edge feature information of the first image based on a processing result of the at least one edge recognition operator.

5. The method of claim 4, wherein, The at least one edge recognition operator comprises at least one of: a straight line operator configured to recognize a straight line edge of a preset angle in the first image; an arc line operator configured to recognize an arc line edge of a preset curvature in the first image; an angle operator configured to recognize an angle edge of a preset angle in the first image; a continuous curve operator configured to recognize a curve edge of a continuous curve type in the first image; a texture operator configured to recognize a texture edge of a preset frequency in the first image.

6. The method of claim 1, wherein, The determining of the fusion feature information by fusing the global feature information, the segmentation feature information and the edge feature information comprises: determining weight information corresponding to the global feature information, the segmentation feature information and the edge feature information by using a weight unit; and fusing the global feature information, the segmentation feature information and the edge feature information based on the weight information to determine the fusion feature information.

7. The method of claim 1, wherein, The determining of the fitted edge corresponding to the edge content in the third sample image comprises: segmenting the third sample image by using a semantic segmentation model to determine a plurality of regions corresponding to the third sample image; and fitting edge content of the plurality of regions to determine the fitted edge corresponding to the edge content.

8. The method of claim 1, wherein, The model is obtained by training based on the following process: encoding a first sample image by using the first unit of the model to generate predicted global feature information corresponding to the first sample image; processing the first sample image by using the second unit of the model to generate predicted segmentation feature information of the first sample image; processing the first sample image by using the third unit of the model to generate predicted edge feature information of the first sample image; and training the model based on the predicted global feature information, the predicted segmentation feature information, the predicted edge feature information and the second sample image. The training of the model based on the predicted global feature information, the predicted segmentation feature information, the predicted edge feature information and the second sample image comprises:

9. The method of claim 8, wherein, adjusting parameters of the first unit based on a first loss function determined by the predicted global feature information, the predicted segmentation feature information and the second sample image; adjusting parameters of the second unit based on a second loss function determined by the predicted segmentation feature information and the second sample image; and adjusting parameters of the third unit based on a third loss function determined by the predicted edge feature information, the predicted segmentation feature information and the second sample image. The training of the model based on the predicted global feature information, the predicted segmentation feature information, the predicted edge feature information and the second sample image further comprises:

10. The method of claim 9, wherein, ​ The first unit, the second unit, and the third unit in the model are collaboratively trained based on a target loss function determined based on the first loss function, the second loss function, and the third loss function.

11. The method of claim 1, wherein, The first image corresponds to animation content or game content.

12. An apparatus for image processing, characterized by The apparatus comprises: An encoding module configured to encode, by a first unit of a model, a first image to generate global feature information of the first image, the first image corresponding to a first resolution; A first processing module configured to process, by a second unit of the model, the first image to generate segmentation feature information of the first image, the segmentation feature information indicating a distribution of a plurality of semantic regions in the first image; A second processing module configured to process, by a third unit of the model, the first image to generate edge feature information of the first image, the edge feature information indicating at least one type of edge in the first image; A fusion module configured to determine fusion feature information by fusing the global feature information, the segmentation feature information, and the edge feature information; and A construction module configured to construct, by upsampling the fusion feature information, a second image, the second image corresponding to a second resolution, the second resolution being higher than the first resolution, wherein the model is trained based on a first sample image and a second sample image, the second sample image being determined by: determining a third sample image corresponding to the first sample image, a resolution of the first sample image being lower than a resolution of the third sample image; determining a fitted edge corresponding to edge content in the third sample image; performing blur processing on a preset region in the third sample image to determine a fourth sample image, wherein the preset region is configured to present the edge content in the third sample image; and filling the fitted edge into the preset region in the fourth sample image to determine the second sample image.

13. An electronic device, comprising: The electronic device comprises: at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, causing the electronic device to perform the method according to any one of claims 1 to 11.

14. A computer-readable storage medium having stored thereon computer- executable instructions, wherein, The computer-executable instructions are executable by a processor to implement the method according to any one of claims 1 to 11.

15. A computer program product, the computer program product being tangibly stored in a computer storage medium and comprising computer-executable instructions, the computer program product being characterized in that, The computer-executable instructions, when executed by a device, cause the device to perform the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Edge perception image semantic segmentation method based on adaptive feature fusion

    CN113658200A

  • Super-resolution image reconstruction method and system, electronic equipment and storage medium

    CN119963416A