Image processing method, electronic equipment, chip system and storage medium
The method uses dual image processing models to enhance clarity and suppress pseudo-textures in images by fusing models that prioritize texture preservation and high-frequency detail restoration, addressing the issue of pseudo-textures in deep learning networks.
Patent Information
- Application Number
- CN202410036200.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-08
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-01-08
AI Technical Summary
In the prior art, when deep learning networks process images, due to the limited distribution range of training samples, pseudo-textures may appear in the generated target images, which reduces image quality.
Two image processing models are used to process the original image. The first model is used to retain more texture features and color space, and the second model is used to improve clarity and suppress pseudo-textures through image fusion technology to generate high-definition target images.
While ensuring clarity, more texture features and color space are retained, effectively suppressing pseudo-textures and improving the quality of the target image.
Smart Images

Figure CN120318086A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to an image processing method, an electronic device, a chip system, and a storage medium. Background Art
[0002] After an electronic device obtains the original image captured by a camera, the original image can be processed through a deep learning network to obtain a high-definition target image, thereby improving the user's photo-taking experience. However, due to the limited learning ability of the deep learning network and the limited distribution range of the training samples during the training of the deep learning network, there may be a situation of pseudo-textures in the generated target image, that is, areas without textures in the original image appear with textures in the target image, reducing the quality of the obtained target image. Summary of the Invention
[0003] This application provides an image processing method, an electronic device, a chip system, and a storage medium, which solve the problem of pseudo-textures existing in the target image obtained after processing the original image in the prior art.
[0004] To achieve the above object, this application adopts the following technical solutions:
[0005] In a first aspect, an image processing method is provided, which is executed on an electronic device. The method includes:
[0006] In response to a first instruction, obtain the original image collected by the camera of the electronic device;
[0007] If it is determined that there is a preset area in the original image, perform image fusion processing on a first image and a second image to obtain a target image; the first image is the image obtained after the original image is processed by a first image processing model, and the second image is the image obtained after the original image is processed by a second image processing model; the first image processing model is trained according to a first training sample, and the second image processing model is trained according to a second training sample, and the image degradation degree of the second training sample is higher than that of the first training sample; the texture features of the preset area in the second image are inconsistent with those in the original image;
[0008] Output the target image.
[0009] In the above embodiments, since the image degradation degree of the second training samples used to train the second image processing model is higher than that of the first training samples used to train the first image processing model, the image output by the second image processing model has higher image quality relative to the image input to the second image processing model. That is, the second image obtained by using the second image processing model has better image quality and higher clarity relative to the original image. The first image obtained by using the first image processing model has poorer quality and weaker detail restoration ability compared to the second image, but retains more texture features and color spaces. That is, more texture features and color spaces of the preset area are retained in the first image. Therefore, the target image obtained by fusing the first image and the second image can retain more texture features and color spaces of the preset area while ensuring clarity, thereby suppressing the appearance of pseudo-textures in the preset area and improving the image quality of the target image.
[0010] In one embodiment, the image fusion processing of the first image and the second image to obtain a target image includes: generating a first image to be fused according to the position of the preset area and the first image, where the first image to be fused is used to represent the pixel information of the preset area; generating a second image to be fused according to the position of the preset area and the second image, where the second image to be fused is used to represent the pixel information of the area outside the preset area; and performing image fusion processing on the first image to be fused and the second image to be fused to obtain the target image.
[0011] Since the first image to be fused is generated according to the first image and is used to represent the pixel information of the preset area, and the second image to be fused is generated according to the second image and is used to represent the pixel information of the area outside the preset area, when the first image to be fused and the second image to be fused are subjected to image fusion processing, the preset area in the first image can replace the preset area in the second image, and the area outside the preset area in the second image can be retained. Therefore, the clarity of the target image can be improved while suppressing the appearance of pseudo-textures in the preset area.
[0012] In one embodiment, before the image fusion processing of the first image and the second image, the method further includes: determining whether the preset area exists in the original image according to the pixel difference between the corresponding local areas of the second image and the original image. If the pixel difference between a local area in the second image and in the original image is large, it indicates that there is a difference between the corresponding local area in the second image and in the original image, and further indicates that the second image has lost more image details relative to the original image. Then it is determined that the preset area exists in the second image.
[0013] In one embodiment, determining whether the preset region exists in the original image according to the pixel differences between the corresponding local regions in the second image and the original image includes: determining a difference image between the second image and the original image, where the difference image represents the pixel differences between the corresponding pixels in the second image and the original image; determining the variance of the pixel values of each local region in the difference image; if there is a local region with a variance greater than the threshold, it indicates that the pixel difference between the corresponding local regions in the original image and the second image is large, and it is determined that the preset region exists in the original image; if there is no local region with a variance greater than the threshold, it is determined that the preset region does not exist in the original image.
[0014] In one embodiment, the method further includes: if it is determined that the preset region does not exist in the original image, using the second image as the target image, so as to obtain a target image with higher clarity.
[0015] In one embodiment, before obtaining the original image captured by the camera of the electronic device, the method further includes: training a first classification model according to a first loss function and the first training sample to obtain the first image processing model, where the first loss function is determined according to the first pixel difference between the first prediction image output by the first classification model and the corresponding first label image, and the first pixel difference is the difference between the pixel values of the corresponding pixels in the first prediction image and the first label image. Training the first classification model with the first loss function to obtain the first image processing model can improve the closeness between the image output by the first image processing model and the original image, so that the image output by the first image processing model retains more texture features of the original image.
[0016] In one embodiment, before obtaining the original image captured by the camera of the electronic device, the method further includes: training a second classification model according to a second loss function and the second training sample to obtain the second image processing model, where the second loss function is determined according to the second pixel difference between the second prediction image output by the second classification model and the corresponding second label image, and the frequency domain difference of the preset region; the second pixel difference is the difference between the pixel values of the corresponding pixels in the second prediction image and the second label image, and the frequency domain difference of the preset region includes the difference in amplitude and the difference in phase between the first frequency domain image and the second frequency domain image, the first frequency domain image is an image obtained by performing frequency domain conversion on the preset region in the second prediction image, and the second frequency domain image is an image obtained by performing frequency domain conversion on the preset region in the second label image. Since the second loss function includes the frequency domain difference of the preset region, training the second classification model with the second loss function to obtain the second image processing model can improve the image restoration ability of the second image processing model in the high-frequency texture region.
[0017] In one embodiment, the method further includes:
[0018] Determine a first standard deviation of pixel values of each local region in the second predicted image; determine a second standard deviation of pixel values of each local region in the second labeled image; and determine the location of the preset region according to the difference between each of the first standard deviations and the corresponding second standard deviation. Specifically, for the same local region, if the difference between the first standard deviation and the second standard deviation is greater than a threshold, it is determined that there is a pixel difference between the local region in the second labeled image and the second predicted image, and the corresponding local region is the region with texture difference, and the region composed of all local regions is the preset region.
[0019] In a second aspect, there is provided an image processing apparatus, including:
[0020] An acquisition module, configured to acquire an original image collected by a camera of the electronic device in response to a first instruction;
[0021] A processing module, configured to perform image fusion processing on a first image and a second image to obtain a target image if it is determined that there is a preset region in the original image; the first image is an image obtained by processing the original image through a first image processing model, the second image is an image obtained by processing the original image through a second image processing model; the first image processing model is trained according to a first training sample, the second image processing model is trained according to a second training sample, and the image degradation degree of the second training sample is higher than that of the first training sample; the texture feature of the preset region in the second image is inconsistent with the texture feature in the original image;
[0022] An output module, configured to output the target image.
[0023] In one embodiment, the processing module is specifically configured to:
[0024] Generate a first image to be fused according to the location of the preset region and the first image, where the first image to be fused is used to represent pixel information of the preset region;
[0025] Generate a second image to be fused according to the location of the preset region and the second image, where the second image to be fused is used to represent pixel information of the region other than the preset region;
[0026] Perform image fusion processing on the first image to be fused and the second image to be fused to obtain the target image.
[0027] In one embodiment, the processing module is further configured to:
[0028] Determine whether the preset region exists in the original image according to the pixel differences between the corresponding local regions in the second image and the original image.
[0029] In one embodiment, the processing module is further configured to:
[0030] Determine a difference image between the second image and the original image, where the difference image characterizes the pixel differences between the corresponding pixels in the second image and the original image;
[0031] Determine the variance of the pixel values of each local region in the difference image;
[0032] If there is a local region where the variance is greater than a threshold, determine that the preset region exists in the original image;
[0033] If there is no local region where the variance is greater than the threshold, determine that the preset region does not exist in the original image.
[0034] In one embodiment, the processing module is further configured to:
[0035] If it is determined that the preset region does not exist in the original image, use the second image as the target image.
[0036] In one embodiment, the processing module is further configured to:
[0037] Train a first classification model according to a first loss function and the first training samples to obtain the first image processing model, where the first loss function is determined according to the first pixel difference between the first predicted image output by the first classification model and the corresponding first labeled image, and the first pixel difference is the difference between the pixel values of the corresponding pixels in the first predicted image and the first labeled image.
[0038] In one embodiment, the processing module is further configured to:
[0039] Train a second classification model according to a second loss function and the second training samples to obtain the second image processing model, where the second loss function is determined according to the second pixel difference between the second predicted image output by the second classification model and the corresponding second labeled image, and the frequency domain difference of the preset region; the second pixel difference is the difference between the pixel values of the corresponding pixels in the second predicted image and the second labeled image, and the frequency domain difference of the preset region includes the difference in amplitude and the difference in phase between a first frequency domain image and a second frequency domain image, the first frequency domain image is an image obtained by performing frequency domain conversion on the preset region in the second predicted image, and the second frequency domain image is an image obtained by performing frequency domain conversion on the preset region in the second labeled image.
[0040] In one embodiment, the processing module is further configured to:
[0041] Determine a first standard deviation of pixel values of each local region in the second predicted image;
[0042] Determine a second standard deviation of pixel values of each local region in the second label image;
[0043] Determine the location of the preset region according to the difference between each of the first standard deviations and the corresponding second standard deviation.
[0044] In a third aspect, an electronic device is provided, including a processor, and the processor is configured to execute a computer program stored in a memory to implement the image processing method as described in the first aspect above.
[0045] In a fourth aspect, a chip system is provided, including a processor, the processor is coupled to a memory, the processor includes a central processing unit and an image processor, and the central processing unit or the image processor executes a computer program or instruction stored in the memory to implement the image processing method as described in the first aspect above.
[0046] Wherein, the image processor may be an image signal processor, a digital signal processor, a neural network processor, etc.
[0047] In a fifth aspect, a computer-readable storage medium is provided, and the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the image processing method as described in the first aspect above is implemented.
[0048] In a sixth aspect, a computer program product is provided, and when the computer program product runs on an electronic device, the electronic device is enabled to execute the image processing method described in the first aspect above.
[0049] It can be understood that the beneficial effects of the second aspect to the sixth aspect above can refer to the relevant descriptions in the first aspect above, and will not be elaborated here. Description of Drawings
[0050] Figure 1 It is a software architecture diagram of an electronic device provided by an embodiment of the present application;
[0051] Figure 2 It is a photographing scene diagram provided by an embodiment of the present application;
[0052] Figure 3 It is a schematic diagram of an image with a texture region provided by an embodiment of the present application;
[0053] Figure 4 It is a schematic diagram of an image with a pseudo-texture region provided by an embodiment of the present application;
[0054] Figure 5 Schematic diagram of the process for training the first image processing model provided by an embodiment of the present application;
[0055] Figure 6 Schematic diagram of the process for training the second image processing model provided by an embodiment of the present application;
[0056] Figure 7 Schematic diagram of the method for determining the position of the pseudo - texture region in the second predicted image provided by an embodiment of the present application;
[0057] Figure 8 Schematic diagram of the mask of the pseudo - texture region provided by an embodiment of the present application;
[0058] Figure 9 Schematic diagram of the process of the image processing method provided by an embodiment of the present application;
[0059] Figure 10 Schematic diagram of the image without pseudo - texture region provided by an embodiment of the present application;
[0060] Figure 11 Schematic diagram of the image with pseudo - texture region provided by an embodiment of the present application;
[0061] Figure 12 Schematic diagram of the image fusion method provided by an embodiment of the present application;
[0062] Figure 13 Flowchart of the image processing method provided by an embodiment of the present application;
[0063] Figure 14 Schematic diagram of the structure of the electronic device provided by an embodiment of the present application. Detailed implementation manners
[0064] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures, technologies, etc. are set forth in order to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well - known systems, devices, circuits, and methods are omitted so as not to obscure the description of the present application with unnecessary details.
[0065] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0066] It should also be understood that the term "and / or" as used in the specification and claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0067] As used in the specification and claims of this application, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrases "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" depending on the context.
[0068] In addition, in the description of this application, the terms "first", "second", etc. are only used for distinguishing descriptions and should not be construed as indicating or implying relative importance.
[0069] Reference to "one embodiment" or "some embodiments" etc. described in the specification of this application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0070] Exemplarily, the electronic device described in the embodiments of this application can be a device that can be held / operated with one hand, such as a mobile phone, a tablet computer, a handheld computer, a personal digital assistant (PDA), an augmented reality (AR) / virtual reality (VR) device, a media player, a wearable device, etc. The embodiments of this application do not impose special restrictions on the specific form / type of this electronic device. The above-mentioned electronic device includes but is not limited to devices equipped with Harmony OS or other operating systems.
[0071] The software system of the electronic device can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In the embodiments of the present invention, the Android system with a layered architecture is taken as an example to exemplarily illustrate the software structure of the electronic device.
[0072] Figure 1It is a software structure block diagram of an electronic device according to an embodiment of the present invention.
[0073] The layered architecture divides the software into several layers, and each layer has clear roles and divisions of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into five layers, from top to bottom are the application layer, the application framework layer, the Android runtime and system libraries, the Hardware Abstraction Layer (HAL), and the kernel layer.
[0074] The application layer may include a series of application packages.
[0075] Such as Figure 1 shown, the application packages may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc.
[0076] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.
[0077] Such as Figure 1 shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, etc.
[0078] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.
[0079] The content provider is used to store and obtain data, and make this data accessible to applications. The data may include video, image, audio, dialed and received calls, browsing history and bookmarks, phone book, etc.
[0080] The view system includes visible controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build applications. The display interface can be composed of one or more views. For example, a display interface including a short message notification icon may include a view for displaying text and a view for displaying pictures.
[0081] The phone manager is used to provide the communication function of the electronic device. For example, the management of call states (including connection, disconnection, etc.).
[0082] The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, etc.
[0083] The notification manager enables an application to display notification information in the status bar. It can be used to convey messages of the notification type, and can automatically disappear after a short stay without user interaction. For example, the notification manager is used to inform that a download is completed, message reminders, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as a notification of a background running application, or a notification that appears in the form of a dialogue window on the screen. For example, it can prompt text information in the status bar, emit a prompt tone, vibrate the electronic device, blink the indicator light, etc.
[0084] The Android Runtime includes core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.
[0085] The core libraries consist of two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core libraries of Android.
[0086] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0087] The system libraries can include multiple functional modules. For example: Surface Manager, Media Libraries, 3D graphics processing library (such as: OpenGL ES), 2D graphics engine (such as: SGL), etc.
[0088] The Surface Manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications.
[0089] The Media Libraries support the playback and recording of various common audio and video formats, as well as static image files, etc. The Media Libraries can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0090] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc.
[0091] The 2D graphics engine is the drawing engine for 2D drawing.
[0092] HAL is used to provide a unified hardware access interface and resource management, enabling upper-layer applications to be developed and run independently of the specific hardware platform. HAL includes a camera interface, a sensor interface, an audio interface, etc.
[0093] The kernel layer is the layer between hardware and software. The kernel layer at least includes a display driver, a camera driver, an audio driver, and a sensor driver.
[0094] The following takes the photographing scenario as an example to exemplarily illustrate the working processes of the software and hardware of the electronic device.
[0095] When the touch sensor receives a touch operation, the corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including information such as touch coordinates and the timestamp of the touch operation). The raw input event is stored in the kernel layer. The application framework layer obtains the raw input event from the kernel layer and identifies the control corresponding to the input event. Taking the touch operation as a touch click operation and the control corresponding to the click operation being the control of the camera application icon as an example, the camera application calls the interface of the application framework layer to start the camera application, and then starts the camera driver by calling the kernel layer to capture a static image or video through the camera.
[0096] The image processing method provided by the embodiments of the present application is applied to the photographing scenario.
[0097] Exemplarily, as shown in (a) of Figure 2 , in an application scenario, the electronic device displays a preview interface of the camera application program on the display interface in response to a user's instruction to open the camera application program. When detecting an operation of the user clicking the "take a photo" control 21, the electronic device acquires the raw image collected by the camera, performs image processing on the raw image to obtain a target image, and displays the target image as shown in (b) of Figure 2 on the display interface. The target image is a high-definition image, and the image clarity and resolution of the target image are higher than those of the raw image.
[0098] In another application scenario, when the electronic device displays a preview interface of the camera application program on the display interface, if it detects an instruction of the user to select a preset photographing mode (such as a high-pixel photographing mode), it performs image processing on the raw image collected by the camera to obtain a target image and displays the target image on the preview interface. Or when the camera application program of the electronic device is in a preset photographing mode, in response to a photographing instruction, it acquires the raw image collected by the camera, performs image processing on the raw image to obtain a target image and displays the target image on the display interface.
[0099] In other application scenarios, the electronic device can also directly display the target image on the preview interface of the camera application program in response to a user's instruction to open the camera application program. Or, when the electronic device displays a preview interface of the camera application program on the display interface, in response to the user's video recording instruction, it processes each frame of the raw image in the raw video to obtain a target video composed of multiple frames of target images and displays the target video on the display interface.
[0100] Among them, the electronic device performs image processing on the original image through an image processing model to obtain a target image. The resolution and clarity of the target image are higher than those of the original image. However, due to the limited image processing ability of the image processing model, the image processing model will lose image details during the process of extracting the image details of the original image, resulting in the appearance of pseudo-texture regions in the target image.
[0101] Among them, the texture region of an image refers to a region composed of some repeatedly appearing texture units, and the texture unit can be a line, a point, or a pattern with a specific shape, etc. The appearance rule of the texture unit can be regular appearance, random appearance, or semi-regular appearance. In the texture region, the pixels or brightness of the image show regular changes. For example, Figure 3 In the region 31 of the image shown, the stripes of two kinds of pixels appear repeatedly, and the region 31 is the texture region.
[0102] The pseudo-texture region refers to a region that is not a texture region in the original image but is recognized as a texture region in the target image. For example, in the original image shown in (a) of Figure 4 , there is no change in pixels or brightness in the region 41, so it is not a texture region. Performing image processing on the original image shown in (a) of Figure 4 , the target image shown in (b) of Figure 4 is obtained. There is a texture (i.e., a region where pixels show regular changes) in the region 42 corresponding to the region 41, and the region 42 is the pseudo-texture region. The region 41 in the original image corresponding to the region 42 is also called the pseudo-texture region.
[0103] To solve the problem of pseudo-textures appearing in the image, the method of suppressing pseudo-textures can be adopted by increasing training samples during the training stage of the image processing model or adjusting the loss function of the region prone to pseudo-textures. However, these methods are only applicable to the scenario of processing images close to the training samples. In the actual photo-taking scenario, the photo-taking environment, photo-taking distance, and photo-taking object vary greatly, and the types of original images obtained are also diverse. As a result, the regions prone to pseudo-textures are also inconsistent. Therefore, the image processing model trained using the above methods is not applicable to the actual photo-taking scenario.
[0104] The present application provides an image processing method. After an electronic device obtains an original image, it first determines whether there is a preset area in the original image. The preset area is a pseudo-texture area, that is, after the second image model processes the original image to obtain a second image, the texture features in the second image are inconsistent with the texture features in the original image. That is, there is no texture in the preset area in the original image, but there is texture in the second image. When it is determined that there is a preset area, the electronic device performs image fusion processing on the first image and the second image to obtain a target image. Among them, the first image is the image obtained after the first image processing model processes the original image. Since the image degradation degree of the second training samples used to train the second image processing model is higher than that of the first training samples used to train the first image processing model, the second image obtained by using the second image processing model has better image quality and higher clarity compared to the original image. The first image obtained by using the first image processing model has poorer quality and weaker detail recovery ability compared to the second image, but retains more texture features and color spaces. That is, more texture features and color spaces of the preset area are retained in the first image. Therefore, the target image obtained by fusing the first image and the second image can retain more texture features and color spaces of the preset area while ensuring clarity, thereby suppressing the appearance of pseudo-textures in the preset area and improving the image quality of the target image.
[0105] The image processing method provided by the embodiments of the present application will be described in detail below.
[0106] In the image processing method provided by the embodiments of the present application, the original image needs to be processed by a first image processing model and a second image processing model. First, the training methods of the first image processing model and the second image processing model will be introduced.
[0107] The first image processing model is trained according to the first training samples. Specifically, the first training samples include multiple image groups, and each image group includes two images with the same content but different image qualities. Image quality reflects the clarity of an image. The higher the image quality, the higher the clarity of the image, and the lower the image quality, the lower the clarity of the image. One of the two images can be an initial image, and the other image can be an image obtained after performing image degradation processing on the initial image. Image degradation processing refers to blurring an image to obtain an image with lower clarity. The two images can also be images obtained by taking pictures of the same object using different shooting modes, and the images obtained by the two shooting modes have different clarities. Each image group is input into the first classification model, and the parameters of the first classification model are optimized to obtain the first image processing model. Among them, the first classification model can be a neural network model.
[0108] Specifically, as Figure 5As shown, in each image group of the first training sample, the low-definition image is used as the first input image, and the high-definition image is used as the first label image. The first classification model is used to process the first input image to obtain a first predicted image, and the parameters of the first classification model are optimized according to the difference between the first predicted image and the first input image.
[0109] In one embodiment, the first pixel difference between the first predicted image and the corresponding first label image is used as the first loss function, and the parameters of the first classification model are optimized according to the loss function.
[0110] Specifically, the first loss function is determined according to the first differences of the corresponding pixels in the first predicted image and the corresponding first label image.
[0111] In one embodiment, by determining the first differences of the corresponding pixels in the first predicted image and the corresponding first label image, multiple first differences can be obtained, and the mean value of the squares of the first differences is the first loss function.
[0112] In other embodiments, the first loss function can also be the sum of the first differences, the mean value of the first differences, the sum of the squares of the first differences, etc.
[0113] The electronic device optimizes the parameters of the first classification model according to the values of the first loss functions corresponding to the respective image groups of the first training sample. The parameters corresponding to the minimum value of the first loss function are the optimal parameters, and the first classification model with the optimal parameters is the first image processing model.
[0114] By using the first pixel difference between the first predicted image and the corresponding first label image as the first loss function to optimize the first classification model, the closeness between the image output by the first image processing model and the input image can be ensured, and richer texture features and color spaces of the input image can be retained.
[0115] The second image processing model is trained according to the second training samples, and the image degradation degree of the second training samples is higher than that of the first training samples. Specifically, the second training samples include multiple image groups, each image group includes two images, the contents of the two images are the same, but the image quality is different, that is, the sharpness of the two images is different. Compared with each image group of the first training samples, in each image group of the second training samples, the sharpness difference between the two images is greater. Exemplarily, for any initial image, perform a first image degradation process on the initial image to obtain a first sample image, and perform a second image degradation process on the initial image to obtain a second sample image. The degradation degree of the second image degradation process is higher than that of the first image degradation process, that is, the image degradation degree of the second sample image is higher than that of the first sample image. Take the initial image and the first sample image as an image group of the first training samples, and take the initial image and the second sample image as an image group of the second training samples. Input each image group into the second classification model, and optimize the parameters of the second classification model to obtain the second image processing model. Among them, the second classification model can be a neural network model.
[0116] Specifically, in each image group, take the low - sharpness image as the second input image, and take the high - sharpness image as the second label image. Use the second classification model to process the second input image to obtain a second predicted image, and optimize the parameters of the second classification model according to the difference between the second predicted image and the second input image. The second classification model with optimized parameters is the second image processing model.
[0117] In each image group of the second training samples, the sharpness difference between the two images is greater, that is, the sharpness difference between the second input image and the second label image is greater. Correspondingly, the sharpness difference between the second input image and the second predicted image is greater. When the second classification model generates the second predicted image according to the second input image, it is easy to cause pseudo - texture regions to appear in the second predicted image.
[0118] In one embodiment, as Figure 6 shown, the electronic device determines the pseudo - texture region in the second predicted image according to the difference between the second predicted image and the corresponding second label, and then determines the second loss function according to the second pixel difference between the second predicted image and the corresponding second label image and the frequency - domain difference of the pseudo - texture region. The frequency - domain difference of the pseudo - texture region includes the difference in amplitude and the difference in phase between the first frequency - domain image and the second frequency - domain image. The first frequency - domain image is an image obtained by performing frequency - domain conversion on the pseudo - texture region in the second predicted image, and the second frequency - domain image is an image obtained by performing frequency - domain conversion on the pseudo - texture region in the second label image.
[0119] Specifically, the electronic device inputs the second input image into the second classification model to obtain the corresponding second predicted image. Determine the first standard deviation of the pixel values of each local region in the second predicted image, and determine the second standard deviation of the pixel values of each local region in the second label image corresponding to the second input image. Among them, for any image, including multiple local regions with a pixel count of N*N, according to Formula 1:
[0120]
[0121] Determine the standard deviation of the pixel values of each local region. Among them, i represents the pixel in the i-th row, j represents the pixel in the j-th column,
[0122] represents the row to the between, the column to the column between the pixel values of the pixels, Var represents the calculation of the standard deviation.
[0123] The positions of each local region in the second predicted image and each local region in the second label image correspond one by one. For the second predicted image, calculate the first standard deviation of each local region according to the above Formula 1. For the second label image, calculate the second standard deviation of each local region according to the above Formula 1.
[0124] After that, the electronic device determines the positions of the regions with texture differences between the second label image and the second predicted image according to the difference between each first standard deviation and the corresponding second standard deviation. Specifically, the electronic device determines the difference between each first standard deviation and the corresponding second standard deviation, and determines whether each difference is greater than the threshold. If the difference is greater than the threshold, the corresponding local region has a texture difference between the second predicted image and the second label image. If the difference is less than the threshold, the corresponding local region has no texture difference between the second predicted image and the second label image.
[0125] In one embodiment, after the electronic device obtains each first standard deviation and the corresponding second standard deviation, according to Formula 2: Normalize the difference between each first standard deviation and the corresponding second standard to between 0 and 1. Among them, represents the normalized difference, σ pred represents the first standard deviation, σ lobel represents the second characterization difference, and C represents a constant.
[0126] After obtaining the normalized difference according to Formula 2, determine the pseudo-texture region according to the relationship between the normalized difference and the threshold, thereby reducing the calculation amount.
[0127] After obtaining the positions of the regions with texture differences, in the second predicted image, the pixel values of the regions with texture differences are set to 1, and the pixel values of the regions without texture differences are set to 0 to obtain a binary image. Then, the binary image is processed according to a morphological algorithm to obtain a pseudo-texture region. Among them, the morphological algorithm is used to extract specified shapes or features in the image to facilitate target recognition. For example, the morphological algorithm can be an erosion operation. The erosion operation is used to slide a custom structuring element (such as a rectangle or a circle) on the image, compare the pixel points in the image with the pixel points in the structuring element, and the intersection obtained is the pixel of the image after the erosion operation. After the erosion operation, the edges of the image can be smoothed and the pixels in the image can be reduced. The morphological algorithm can be a dilation operation. The dilation operation is used to slide a custom structuring element (such as a rectangle or a circle) on the image, compare the pixel points in the image with the pixel points in the structuring element, and the union obtained is the pixel of the image after the dilation operation. After the dilation operation, the edges of the image can be smoothed and the pixels in the image can be increased. The morphological algorithm can be a hole filling operation. The hole filling operation is used to determine holes and fill the holes in the image with a structuring element to remove the interfering pixels in the image. Among them, a hole is a background region surrounded by the boundaries connected by foreground pixels.
[0128] For example, according to Figure 7 the second predicted image shown in (a) of Figure 7 and the second label image shown in (b) of Figure 7 it is possible to obtain the image shown in (c) of
[0129] where the region with pixels of 1 in the image is the pseudo-texture region, and further the position of the pseudo-texture region in the second predicted image can be determined.
[0130] After obtaining the position of the pseudo-texture region, the pseudo-texture region in the second predicted image is subjected to a frequency domain transformation to obtain a first frequency domain image, and the pseudo-texture region in the second label image is subjected to a frequency domain transformation to obtain a second frequency domain image.
[0131] Among them, the frequency domain transformation refers to performing a two-dimensional Fourier transform on the image to convert the gray distribution function of the image into a frequency distribution function. The frequency of the image characterizes the severity of the gray change in the image. After the frequency domain transformation, the corresponding amplitude and phase of the image can be obtained.
[0132] After obtaining the first frequency-domain image and the second frequency-domain image, determine the frequency-domain difference of the pseudo-texture region according to the difference in amplitude and the difference in phase between the first frequency-domain image and the second frequency-domain image, and then determine the loss function according to the second pixel difference between the second predicted image and the corresponding second label image and the frequency-domain difference of the pseudo-texture region.
[0133] Among them, the frequency-domain difference of the pseudo-texture region can be the sum of the first difference and the second difference. The first difference includes the difference in corresponding amplitudes of the first frequency-domain image and the second frequency-domain image, and the second difference includes the difference in corresponding phases of the first frequency-domain image and the second frequency-domain image. The second pixel difference between the second predicted image and the second label image is determined according to the second difference of the corresponding pixels in the second predicted image and the second label image. In one embodiment, the second pixel difference between the second predicted image and the second label image is the mean of the squares of the second differences. In other embodiments, the second pixel difference between the second predicted image and the second label image can also be the sum of the second differences, or the mean of the second differences, or the sum of the squares of the second differences.
[0134] In one embodiment, the second loss function Loss is Loss = λ1L2Loss + λ2SSIM Loss + λ3Freq1Loss, where λ1, λ2, and λ3 are weights, L2Loss represents the second pixel difference between the second predicted image and the corresponding second label image, Freq 1Loss represents the frequency-domain difference of the pseudo-texture region corresponding to the second predicted image and the second label image, and SSIM Loss is a structural loss function representing the texture difference between the second predicted image and the corresponding second input image.
[0135] Among them, according to the formula determine SSIM Loss, μ x represents the mean of the pixel values of the second predicted image, μ y represents the mean of the pixel values of the second input image, σ xy represents the covariance of the pixel values of the second predicted image and the second input image, σ x represents the variance of the pixel values of the second predicted image, σ y represents the variance of the pixel values of the second input image, and C1 and C2 represent constants.
[0136] In another embodiment, the second loss function Loss is
[0137] Loss = λ3L2Loss + λ4SSIM Loss + λ5Freq2Loss + λ6L3Loss, where λ3, λ4, λ5, and λ6 are all weights. Freq 2Loss represents the frequency domain difference between the third frequency domain image obtained by performing frequency domain transformation on the second predicted image and the fourth frequency domain image obtained by performing frequency domain transformation on the second labeled image, that is, it represents the difference in amplitude and the difference in phase corresponding to the third frequency domain image and the fourth frequency domain image. For example, Freq 2Loss includes a third difference and a fourth difference. The third difference includes the difference in amplitude corresponding to each of the third frequency domain image and the fourth frequency domain image, and the fourth difference includes the difference in phase corresponding to each of the third frequency domain image and the fourth frequency domain image. It can be understood that the first frequency domain image is a part of the third frequency domain image, and the second frequency domain image is a part of the fourth frequency domain image. L3Loss represents the pixel difference between the pseudo-texture region in the second predicted image and the pseudo-texture region in the second labeled image. Exemplarily, calculate the difference image between the second predicted image and the second labeled image. Each pixel in the difference image corresponds to a position. For any pixel, there are pixels with the same position in both the second predicted image and the second labeled image. For any pixel in the difference image, the pixel value of this pixel is the difference between the pixel value of the second predicted image and the pixel value of the second labeled image at the corresponding position. After obtaining the difference image, the product of the difference image and the mask of the pseudo-texture region is L3Loss. Exemplarily, the mask of the pseudo-texture region is determined according to the position where the pseudo-texture region is located. For example, after obtaining the pseudo-texture region, set the pixel value of the pseudo-texture region to 0 and the pixel value of the region outside the pseudo-texture region to 1, then the mask of the pseudo-texture region as shown in Figure 8 is obtained.
[0138] After determining the second loss function, optimize the parameters of the second classification model according to the values of the second loss function corresponding to each image group in the second training sample. The parameters corresponding to the minimum value of the second loss function are the optimal parameters, and the second classification model with the optimal parameters is the second image processing model.
[0139] In the above embodiment, adding the second pixel difference between the second predicted image and the corresponding second labeled image to the second loss function is used to make the pixel value of the second predicted image close to the pixel value of the second labeled image. Adding the texture difference between the second predicted image and the corresponding second input image to the second loss function is used to make the texture difference between the second predicted image and the second input image smaller. Adding the frequency domain difference of the pseudo-texture region to the second loss function is used to make the amplitude and phase of the second predicted image close to the amplitude and phase of the second labeled image, effectively improving the overall image restoration ability of the second image processing model and the image restoration ability in the pseudo-texture region, and further improving the details and clarity of the image output by the second image processing model.
[0140] In one embodiment, the model structures of the first classification model and the second classification model are the same. First, the first classification model can be trained according to the first loss function and the first training samples to obtain the first image processing model. Then, the parameters of the first image processing model are used as the initial parameters of the second classification model, and the second classification model is trained according to the second loss function and the second training samples, so that the second image processing model can be obtained by adjusting the parameters on the basis of the first image processing model, improving the training efficiency of the second image processing model.
[0141] Since the clarity difference between the second input image and the second label image is greater, the second image processing model can output an output image with higher clarity relative to the input image. Therefore, compared with the first image processing model, an image with higher clarity can be obtained after being processed by the second image processing model. That is, both the first image processing model and the second image processing model can perform image enhancement processing on the image, and the image enhancement degree of the first image processing model is weak, while the image enhancement ability of the second image processing model is strong. After being processed by the first image processing model, a smoother image can be obtained, with low image clarity, but it can retain the rich texture features and color space of the input image and is not prone to pseudo-texture regions. The second image processing model can perform more refined feature extraction on the high-frequency texture details of the input image. The image processed by the second image processing model has higher clarity, but pseudo-texture regions are prone to appear while enhancing the high-frequency texture details.
[0142] As Figure 9 shown, the image processing method provided by an embodiment of the present application includes:
[0143] S901: In response to a first instruction, obtain the original image collected by the camera of the electronic device.
[0144] In one embodiment, as Figure 2 shown, the first instruction is a photographing instruction, and the original image is an unprocessed and low-clarity image collected by the camera. In other embodiments, the first instruction may also be a video recording instruction or an instruction to open the camera application.
[0145] S902: If it is determined that there is a preset area in the original image, perform image fusion processing on the first image and the second image to obtain a target image; the first image is the image obtained after the original image is processed by the first image processing model, and the second image is the image obtained after the original image is processed by the second image processing model; the texture features of the preset area in the second image are inconsistent with those in the original image.
[0146] Specifically, the original image is input into the first image processing model to obtain the first image output by the first image processing model, and the original image is input into the second image processing model to obtain the second image output by the second image processing model. The clarity of the second image is higher than that of the first image.
[0147] As Figure 3 and Figure 4 shown, the position of the preset area in the original image is the same as the position of the preset area in the second image. The preset area is a pseudo-texture area. The electronic device determines whether there is an area with inconsistent texture features in the second image and the original image according to the pixel difference between the second image and the original image, and this area is the preset area.
[0148] In one embodiment, the electronic device determines whether there is a preset area in the original image according to the pixel difference between the corresponding local areas in the second image and the original image. Among them, according to the positions of the pixels in the original image, the original image is composed of multiple local areas with a pixel number of N*N, where N is a positive integer, and the value of N can be set according to actual needs. The local areas in the original image correspond one-to-one with the local areas in the second image.
[0149] In one embodiment, the specific method for determining whether there is a preset area in the original image is as follows. The electronic device first determines the difference image between the second image and the original image, and the difference image represents the pixel difference between the corresponding pixels in the second image and the original image. Specifically, the second image includes multiple pixels, each pixel corresponding to a position. There are two pixels with the same position in the original image and the second image, and the difference between the pixel values of the two pixels with the same position is the pixel value of the pixel at the same position in the difference image.
[0150] After obtaining the difference image, the electronic device determines the variance of the pixel values of each local area in the difference image. Exemplarily, according to Equation 3: the standard deviation of the pixel values of each local area in the difference image is determined. Where i represents the pixel in the i-th row, j represents the pixel in the j-th column, represents the pixels from the row to the column to the column, and Var represents the calculation of the standard deviation. After obtaining the standard deviation, the square of the standard deviation is the corresponding variance.
[0151] For each local region, if the variance is greater than the threshold, it is determined that the texture features of the local region in the original image and in the second image are inconsistent, that is, there is a texture difference. If the variance value is less than the threshold, it is determined that the texture features of the local region in the original image and in the second image are consistent, that is, there is no texture difference. The local regions in the original image and the second image with inconsistent texture features are the local regions of the pseudo-textures in the original image. For example, as Figure 10 shown in (a) of Figure 10 and (b) of Figure 11 shown in (a) of Figure 11 and (b) of
[0152] In the image, if each local region is a flat region or an image with a regular and clearly defined edge structure, the variance of each local region in the corresponding difference image is less than the threshold, then there is no pseudo-texture region. As shown in (a) of
[0153] and (b) of
[0154] In the image, if there is texture or the texture edge is not clear, the variance of each local region in the corresponding difference image is greater than the threshold, then there is a pseudo-texture region.
[0155] Correspondingly, if there is a local region with a variance greater than the threshold, it is determined that there is a preset region in the original image. If there is no local region with a variance greater than the threshold, it is determined that there is no preset region in the original image.
[0156] When it is determined that there is a preset region in the original image, the electronic device performs a fusion process on the first image and the second image to obtain a target image. Figure 12As shown, image a represents the original image, image b represents the first image, image c represents the second image, image d represents the pseudo-texture region image, and image e represents the mask of the pseudo-texture region. Multiply the first image and the pseudo-texture region image to obtain the first fused image. Since in the pseudo-texture image, the pixel values of the pseudo-texture region are 1 and the pixel values of the regions outside the pseudo-texture region are 0, in the first fused image obtained by multiplication, the pixel values of the regions outside the pseudo-texture region are 0, and only the pixel values of the pseudo-texture region in the first image are retained. Multiply the second image and the mask of the pseudo-texture region to obtain the second fused image. Since in the mask of the pseudo-texture region, the pixel values of the pseudo-texture region are 0 and the pixel values of the regions outside the pseudo-texture region are 1, in the second fused image obtained by multiplication, the pixel values of the pseudo-texture region are 0, and only the pixel values of the regions outside the pseudo-texture region in the second image are retained. Fuse the first image and the second image to obtain the complete target image f, and the preset region in the target image is obtained from the first image, retaining more texture features, so that there is no pseudo-texture region in the target image.
[0157] In other embodiments, after determining the position where the preset region is located, the preset region may also be segmented from the first image, and the preset region in the second image may be replaced with the preset region segmented from the first image to obtain the target image.
[0158] S903: Output the target image.
[0159] Specifically, if the first instruction is a photographing instruction, after the electronic device processes the original image, the target image is displayed on the display interface. If the first instruction is an instruction to open the camera application, after the electronic device opens the camera application, the target image obtained by processing the original image is displayed on the preview interface.
[0160] In one embodiment, the flow of the image processing method is as Figure 13 shown. The electronic device inputs the original image into the second image processing model, and determines whether there is a preset region in the original image according to the pixel difference between the second image and the original image. If there is a preset region, the first image and the second image are fused to obtain the target image. If there is no preset region, the second image is used as the target image and the target image is output, so that in the case where there is a preset region in the original image, a target image with higher clarity is obtained, and pseudo-textures do not appear in the target image. In the case where there is no preset region in the original image, a target image with higher clarity is obtained. At the same time, according to whether there is a preset region, different image processing methods are used to obtain the target image, which can reduce the computational amount of image post-processing and improve the overall performance of the algorithm.
[0161] In the above embodiments, after the electronic device acquires the original image, it determines whether there is a preset area in the original image. When there is a preset area in the original image, the electronic device inputs the original image into the first image processing model to obtain a first image, inputs the original image into the second image processing model to obtain a second image, and performs a fusion process on the first image and the second image to obtain a target image. Since the second image has better image quality and higher clarity than the original image, and the first image retains more texture features and color spaces than the second image, the target image obtained by fusing the first image and the second image can retain more texture features and color spaces in the preset area while ensuring clarity, thereby suppressing the appearance of pseudo-textures in the preset area and improving the image quality of the target image.
[0162] In one embodiment, the computer program for implementing the above image processing method is stored in the HAL layer. After the data of the original image collected by the camera is uploaded to the HAL layer, the processor calls the computer program for implementing the above image processing method to perform image processing on the data of the original image, and uploads the processed image to the application framework layer for further image processing.
[0163] Among them, the processor can be the central processing unit (CPU) of the electronic device or the image signal processor (ISP).
[0164] In another embodiment, the computer program for implementing the above image processing method can also be stored in the ISP of the electronic device or in a specified memory on the electronic device.
[0165] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0166] Exemplarily, Figure 14 shows a schematic structural diagram of the electronic device 100.
[0167] The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0168] It can be understood that the structure schematically shown in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In some other embodiments of the present application, the electronic device 100 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0169] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modulation and demodulation processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0170] The controller may generate operation control signals according to the instruction operation code and the timing signal to complete the control of fetching instructions and executing instructions.
[0171] A memory can also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0172] In some embodiments, the processor 110 may include one or more interfaces.
[0173] It can be understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are only illustrative descriptions and do not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0174] The electronic device 100 realizes the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, and is connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.
[0175] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. In some embodiments, the electronic device 100 may include 1 or N display screens 194, where N is a positive integer greater than 1.
[0176] The electronic device 100 can realize the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc.
[0177] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In some embodiments, the ISP may be provided in the camera 193.
[0178] The camera 193 is used to capture static images or videos. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transfers the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard formats such as RGB and YUV. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0179] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function.
[0180] The internal memory 121 can be used to store computer-executable program code, and the executable program code includes instructions. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area can store data created during the use of the electronic device 100 (such as audio data, phone book, etc.). In addition, the internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121 and / or the instructions stored in the memory provided in the processor.
[0181] The touch sensor 180K, also known as a "touch control device". The touch sensor 180K can be disposed on the display screen 194, and the touch sensor 180K and the display screen 194 form a touch screen, also known as a "touch control screen". The touch sensor 180K is used to detect touch operations acting on or near it. The touch sensor can transfer the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In some other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device 100, at a different position from the display screen 194.
[0182] It should be noted that, regarding the information interaction, execution process, etc. between the above-mentioned devices / units, since they are based on the same concept as the method embodiments of this application, for their specific functions and the technical effects brought about, reference can be made to the method embodiment section, and details will not be elaborated here.
[0183] In the above embodiments, the descriptions of each embodiment have their own focuses. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0184] Those skilled in the art can clearly understand that, for the sake of convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments, and details will not be elaborated here.
[0185] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of this application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / electronic device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc.
[0186] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0187] In the embodiments provided in this application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the coupling or direct coupling or communication connection shown or discussed with each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0188] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0189] Finally, it should be noted that the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any change or replacement within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. An image processing method, which is executed on an electronic device, characterized in that The method includes: In response to a first instruction, obtaining an original image captured by a camera of the electronic device; If it is determined that there is a preset area in the original image, performing image fusion processing on a first image and a second image to obtain a target image; the first image is an image obtained by processing the original image through a first image processing model, and the second image is an image obtained by processing the original image through a second image processing model; the first image processing model is trained according to a first training sample, and the second image processing model is trained according to a second training sample, and the image degradation degree of the second training sample is higher than that of the first training sample; the texture feature of the preset area in the second image is inconsistent with the texture feature in the original image; Outputting the target image.
2. The method according to claim 1, characterized in that, The performing image fusion processing on the first image and the second image to obtain a target image includes: Generating a first image to be fused according to the position where the preset area is located and the first image, and the first image to be fused is used to represent the pixel information of the preset area; Generating a second image to be fused according to the position where the preset area is located and the second image, and the second image to be fused is used to represent the pixel information of the area other than the preset area; Performing image fusion processing on the first image to be fused and the second image to be fused to obtain the target image.
3. The method according to claim 1, characterized in that Before the performing image fusion processing on the first image and the second image, the method further includes: Determining whether there is the preset area in the original image according to the pixel difference between corresponding local areas in the second image and the original image.
4. The method according to claim 3, wherein The determining whether there is the preset area in the original image according to the pixel difference between corresponding local areas in the second image and the original image includes: Determining a difference image between the second image and the original image, and the difference image represents the pixel difference between corresponding pixels in the second image and the original image; Determining the variance of the pixel values of each local area in the difference image; If there is a local area with a variance greater than a threshold, determining that there is the preset area in the original image; If there is no local area with a variance greater than the threshold, determining that there is no preset area in the original image.
5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: If it is determined that there is no preset area in the original image, using the second image as the target image.
6. The method according to claim 1, wherein Before the obtaining the original image captured by the camera of the electronic device, the method further includes: Training a first classification model according to a first loss function and the first training sample to obtain the first image processing model, where the first loss function is determined according to the first pixel difference between a first predicted image output by the first classification model and a corresponding first labeled image, and the first pixel difference is the difference between the pixel values of corresponding pixels in the first predicted image and the first labeled image.
7. The method according to claim 1, characterized in that, Before the obtaining the original image captured by the camera of the electronic device, the method further includes: The second classification model is trained according to the second loss function and the second training samples to obtain the second image processing model, where the second loss function is determined according to the second pixel difference between the second predicted image output by the second classification model and the corresponding second labeled image, and the frequency domain difference of the preset region; the second pixel difference is the difference between the pixel values of the corresponding pixels of the second predicted image and the second labeled image, and the frequency domain difference of the preset region includes the difference in amplitude and the difference in phase between the first frequency domain image and the second frequency domain image. The first frequency domain image is an image obtained by performing frequency domain conversion on the preset region in the second predicted image, and the second frequency domain image is an image obtained by performing frequency domain conversion on the preset region in the second labeled image.
8. The method according to claim 7, wherein The method further includes: Determining a first standard deviation of the pixel values of each local region in the second predicted image; Determining a second standard deviation of the pixel values of each local region in the second labeled image; Determining the position where the preset region is located according to the difference between each of the first standard deviations and the corresponding second standard deviations.
9. An electronic device, characterized in that, It includes a processor, and the processor is configured to execute a computer program stored in a memory to implement the method according to any one of claims 1 to 8.
10. A chip system, characterized in that, It includes a processor, the processor is coupled to a memory, the processor includes a central processing unit and an image processor, and the central processing unit or the image processor executes a computer program or instruction stored in the memory to implement the method according to any one of claims 1 to 8.
11. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Artifact removal model training method and device, equipment, medium and program product
CN115330615A
Data processing method and device
CN115731115A
Image processing method and related equipment thereof
CN116055895A
Image processing method and device, electronic equipment and computer readable storage medium
CN116452416A
Image splicing method and device, equipment and storage medium
CN116912467A