Image processing method and electronic device
By using the defuzzy model in the image processing method, the images collected by the camera with a fixed focus distance are processed based on the distance information of the target object, which solves the image blur problem and improves the sharpness and accuracy of the image.
Patent Information
- Application Number
- CN202411178577.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-27
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-08-27
AI Technical Summary
In the prior art, images captured by cameras with fixed focus distances may be blurred, resulting in low clarity.
By using an image processing method, the defuzzing model uses the defuzzing model to defuzz the image acquired by the camera with a fixed focus distance to determine the target image.
Improves the accuracy and clarity of the images obtained by processing and improves the user experience.
Smart Images

Figure CN118714466B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and more specifically, to an image processing method and an electronic device. Background Art
[0002] In order to balance factors such as photo effects, device thickness, weight, and cost, the cameras in some electronic devices can be designed with a fixed focus distance. When using an electronic device with a fixed focus distance camera, the user cannot adjust the camera focus distance (i.e., zoom in or out the picture), and can only move the electronic device itself closer to or farther away from the subject. This design simplifies the structure of the camera module, contributes to the lightweight design of electronic devices, and also reduces manufacturing costs. However, the images captured by a camera with a fixed focus distance may be blurry.
[0003] How to improve the clarity of images collected by a camera with a fixed focus distance is an urgent problem to be solved. Summary of the invention
[0004] The present application provides an image processing method, which can improve the clarity of processed images.
[0005] In a first aspect, an image processing method is provided, which is applied to an electronic device, the method comprising: acquiring a first object image through a first camera of the electronic device, the first object image recording a target object, the first camera being a camera with a fixed focus distance; acquiring distance information of the target object, the distance information being used to represent the distance between the target object and the first camera; determining a target image based on the distance information and the first object image through a deblurring model, the deblurring model being capable of deblurring an arbitrary image based on the distance information of an object in the arbitrary image.
[0006] The image processing method provided in the embodiment of the present application processes a first object image recording the target object acquired by a first camera at a fixed focus distance based on the distance information of the target object through a deblurring model to determine the target image. The influence of the distance information of the target object is taken into account during the processing of the deblurring model, so that the accuracy and clarity of the target image are higher, thereby improving the user experience.
[0007] In some possible implementations, the electronic device includes multiple cameras, all of which are cameras with fixed focusing distances, and the focusing distances of the multiple cameras are different; before deblurring the first object image based on the distance information using the deblurring model, the method also includes: determining a candidate model corresponding to the first camera according to a first corresponding relationship, the deblurring model is the candidate model corresponding to the first camera, and the first corresponding relationship represents a corresponding relationship between multiple candidate models and the multiple cameras.
[0008] Different candidate models are set for cameras with different focus distances in the electronic device, and the image captured by each camera can be deblurred using the candidate model corresponding to the camera. In the process of image processing, the influence of the focus distances of different cameras is fully considered, and the first object image is deblurred by selecting the candidate model corresponding to the camera that captures the first object image as the deblur model, thereby improving the accuracy and clarity of the processed image.
[0009] In some possible implementations, the target object belongs to a target object category, and before determining the target image through the deblurring model based on the distance information and the first object image, the method further includes: determining a candidate model corresponding to the target object category according to a second correspondence relationship, the deblurring model is a candidate model corresponding to the target object type, and the second correspondence relationship represents a correspondence between multiple candidate models and multiple object categories.
[0010] Objects belonging to different object categories have different characteristics. Candidate models corresponding to the object categories are set for different object categories. The first object image recording the object is deblurred using the candidate model corresponding to the corresponding category to which each object belongs, taking into account the influence of the characteristics of different object categories, and improving the accuracy and clarity of the processed image.
[0011] In some possible implementations, the target object category of the target object is a face.
[0012] Among the various objects recorded in an image, the object that users are most concerned about is generally the face. For objects belonging to the face category, that is, for face images, deblurring is performed through a deblurring model to improve user satisfaction with the processed image.
[0013] In some possible implementations, the distance information is an object depth image of the target object, and the object depth image is obtained through a second camera of the electronic device.
[0014] The object depth image obtained by the second camera in the electronic device includes a large amount of distance information related to the target object, which can reflect the characteristics of the target object. Through the deblurring model, the first object image is deblurred based on the object depth image to improve the accuracy and clarity of the processed image.
[0015] Compared with the object depth image obtained by performing depth estimation on the first object image, the object depth image obtained by the second camera has higher accuracy, so that the image obtained by performing the deblurring model on the object depth image and the first object image has higher accuracy.
[0016] In some possible implementations, acquiring a first object image through a first camera of the electronic device includes: identifying the first image acquired by the first camera to determine an object area, wherein the object area is an area where the target object in the first image is located; acquiring an image located in the object area in the first image to obtain the first object image; acquiring distance information of the target object includes: determining a depth image of the object located in the object area in a depth image corresponding to the first image, wherein the depth image is acquired through acquisition by the second camera.
[0017] By identifying the first image captured by the first camera, the image area where the target object is located in the first image is obtained, and the image located in the image area in the first image is used as the first object image, and the image located in the image area in the second depth image captured by the second camera is used as the object depth image. The area proportions of the target object recorded in the first object image and the object depth image are both larger, and the image obtained by deblurring model processing is more accurate.
[0018] In some possible implementations, determining the target image based on the distance information and the first object image through a deblurring model includes: processing the distance information and the first object image through the deblurring model to determine a second object image; and fusing the second object image with an image in the first image that is outside the object area to obtain the target image.
[0019] By fusing the second object image with the image outside the object area in the first image, the target image is clearer than the first image.
[0020] The number of target objects obtained by identifying the first image may be one or more. Different target objects correspond to different object areas, that is, different target objects are recorded in different first object images. After the first object area recording each target object is processed by the deblurring model to obtain the second object image of the target object, the multiple object images can be fused with the image outside each object area in the first image to obtain the target image. The object depth information of the multiple target objects may be different. Different target objects are processed separately by the deblurring model, so that each second object image is clearer, thereby the clarity of the target image is higher.
[0021] In some possible implementations, the deblurring model is obtained by training an initial deblurring model based on multiple groups of first training data, each group of first training data in the multiple groups of first training data includes a first training sample image of a first training object, a first reference image and training distance information, wherein the training distance information represents the distance between the first training object and a first training camera that captures the first reference image, the focus distance corresponding to the first training sample image is equal to the focus distance of the first camera, and the first reference image is an image captured after the first training camera focuses on the first training object.
[0022] In some possible implementations, the first training sample image is obtained by blurring the first reference image using a target blur kernel, wherein the target blur kernel is a blur kernel determined from a plurality of blur kernels corresponding to a focus distance of the first camera and corresponding to a target distance range to which the training distance information belongs, and different distance ranges correspond to different blur kernels.
[0023] The image processed by the blur kernel corresponding to a certain distance range can accurately reflect the blur degree of the image acquired by the camera pair with the same focus distance as the first camera and the object whose distance from the camera is within the certain distance range. Different blur kernels are set for different distance ranges, and the image obtained by processing the image with the blur kernel corresponding to a certain distance range is used to simulate the image acquired by the camera pair with the same focus distance as the first camera and the object whose distance from the camera is within the distance range. The image processed by the blur kernel corresponding to a certain distance range can accurately reflect the blur degree of the image acquired by the camera pair with the same focus distance as the first camera and the object whose distance from the camera is within the certain distance range. Therefore, the deblurring model obtained by training based on multiple sets of first training data is more accurate.
[0024] The first training sample image in the first training data is determined by using a fuzzy kernel, thereby reducing the number of images that need to be manually collected during the construction of multiple groups of first training data, thereby reducing labor costs.
[0025] In some possible implementations, any blur kernel among the multiple blur kernels is obtained by training an initial blur kernel based on multiple groups of second training data, each group of second training data among the multiple groups of second training data includes a second training object image and a second reference image of a second training object, the second training object image and the second reference image are respectively obtained by a second training camera and a third training camera that are equal to the distance from the second training object, the training distance between the third training camera and the second training object belongs to the distance range corresponding to any blur kernel, when acquiring the second training object image, the focus distance of the second training camera is equal to the focus distance of the first camera, and the first reference image is an image acquired after the third training camera focuses on the second training object.
[0026] In some possible implementations, determining the target image based on the distance information and the first object image through the deblurring model includes: determining the target image based on the distance information and the first object image through the deblurring model when the difference between the distance represented by the distance information and the focus distance of the first camera is greater than or equal to a first difference threshold; wherein the training distance information in the multiple sets of first training data includes first training distance information and second training distance information, the difference between the distance represented by the first training distance information and the focus distance of the first camera is greater than or equal to the first difference threshold, and the difference between the distance represented by the second training distance information and the focus distance of the first camera is less than the first difference threshold.
[0027] When the distance between the target object and the first camera is greater than or equal to the first difference threshold, the clarity of the first object image may be low, and the first object image is deblurred. When the distance between the target object and the first camera is less than the first difference threshold, the first object image is clear, and deblurring is no longer required, thereby obtaining a clear image and reducing the resource occupation of image processing.
[0028] During the model inference process, when the distance between the target object and the first camera is less than the first difference threshold, deblurring may not be performed. In the multiple sets of first training data used in the training process of the deblurring model, first training distance information may be included, in which the difference between the distance represented and the focus distance of the first camera is greater than or equal to the first difference threshold, and second training distance information may be included, in which the difference between the distance represented and the focus distance of the first camera is less than the first difference threshold. During the training process, the deblurring model is trained using the first training data including the second training distance information and the first training data including the first training distance information, so that the deblurring model is more accurate and the accuracy of the image processed by the deblurring model is improved.
[0029] In a second aspect, a training method for a deblurring model is provided, comprising: obtaining multiple groups of first training data, each group of first training data comprising a first training sample image of a first training object, a first reference image and training distance information, wherein the training distance information represents the distance between the first training object and a first training camera that captures the first reference image, the camera focus distance corresponding to the first training sample image is equal to the focus distance of the first camera, and the first reference image is an image captured after the first training camera focuses on the first training object; using an initial deblurring model, based on the training distance information, deblurring the first training object image to obtain a third training object image; and adjusting the parameters of the initial deblurring model according to the difference between the third training object image and the first reference image.
[0030] In a third aspect, an image processing apparatus is provided, comprising a unit for executing the method of the first aspect and / or the second aspect. The apparatus may be a terminal device or a chip in the terminal device.
[0031] In a fourth aspect, an electronic device is provided, comprising a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the electronic device executes the method of the first aspect and / or the second aspect.
[0032] In a fifth aspect, a chip is provided, comprising a processor and a data interface, wherein the processor reads instructions stored in a memory through the data interface to implement the method of the first aspect and / or the second aspect.
[0033] In a sixth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program code, and the computer program code is used to implement the method of the first aspect and / or the second aspect.
[0034] In a seventh aspect, a computer program product is provided, the computer program product comprising: a computer program code, wherein the computer program code is used to implement the method of the first aspect and / or the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figures 1 to 3 is a schematic diagram of a graphical user interface provided in an embodiment of the present application;
[0036] Figure 4 is a schematic structural diagram of an image processing model provided in an embodiment of the present application;
[0037] Figure 5 is a schematic flow chart of a deblurring model training method provided in an embodiment of the present application;
[0038] Figure 6 is a schematic diagram of a training process of a blur kernel and a deblurring model provided in an embodiment of the present application;
[0039] Figure 7 is a schematic flow chart of an image processing method provided in an embodiment of the present application;
[0040] Figure 8 is a schematic flow chart of another image processing method provided in an embodiment of the present application;
[0041] Fig. 9 is a schematic diagram of a hardware system of an electronic device applicable to the present application;
[0042] Fig.10 is a schematic diagram of a software system for an electronic device applicable to the present application;
[0043] Fig.11 It is a schematic structural diagram of an image processing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0044] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below in conjunction with the accompanying drawings.
[0045] It should be understood that the "at least one" mentioned in this application refers to one or more, and "multiple" refers to two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. In the description of this application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in order to facilitate the clear description of the technical solution of this application, the words "first" and "second" are used to distinguish the same or similar items with basically the same functions and effects. Those skilled in the art can understand that the words "first" and "second" do not limit the quantity and execution order, and the words "first" and "second" do not limit them to be different.
[0046] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0047] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties (for example, the user's explicit consent, effective notification to the user, etc.), and the collection, use and processing of relevant data must comply with the relevant regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0048] The technical solutions in the embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0049] Artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0050] Machine learning is an important branch of artificial intelligence, and deep learning is an important branch of machine learning. Deep learning refers to the use of multi-layer neural network structures to learn from big data the representation of various things in the real world that can be directly used for computer calculations (for example, things in images, sounds in audio, etc.).
[0051] The neural network model is a mathematical method that simulates the actual human neural network. The neural network model is a complex network system formed by a large number of simple processing units (called neurons or neural units) that are widely connected to each other. The neural network model is a network formed by connecting many of the above-mentioned single neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the characteristics of the local receptive field. The local receptive field can be an area composed of several neural units.
[0052] A deep neural network (DNN) can also be called a multi-layer neural network, which can be understood as a neural network with multiple hidden layers. According to the position of different layers, the neural network inside the DNN can be divided into three categories: input layer, hidden layer, and output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the layers in between are hidden layers. The layers are fully connected, that is, any neuron in the i-th layer is connected to any neuron in the i+1-th layer. In a deep neural network, more hidden layers allow the network to better depict complex situations in the real world. In theory, the more parameters a model has, the higher its complexity and the greater its "capacity", which means it can complete more complex learning tasks. Training a deep neural network is the process of learning the weight matrix, and its ultimate goal is to obtain the weight matrix of all layers of the trained deep neural network (the weight matrix formed by many layers of vectors).
[0053] Neural network models can be applied in the process of image deblurring.
[0054] In order to balance factors such as photo quality, device thickness, weight, and cost, the cameras in some electronic devices can be designed with a fixed focus distance. Cameras with a fixed focus distance can also be called fixed focus cameras or fixed focus (FF) cameras.
[0055] The focus distance refers to the distance between the object and the image. When the distance between the object and the image sensor in the camera is the focus distance, the object recorded in the image captured by the camera is clear. In other words, if the distance between the position of the object and the image sensor is the focus distance, the light from the object can form a clear image point on the sensor, thereby forming a clear image on the image sensor.
[0056] The focusing distance may be understood as the distance between the image sensor and the object at which the object recorded in the image captured by the image sensor is clearest.
[0057] A point whose distance from the image sensor is the focusing distance and which is located on the axis of the lens can be called the focusing point of the camera.
[0058] Points within a certain range in front of and behind the focus point along the lens axis can also form clearer images on the image sensor that are acceptable to the eye.
[0059] Fixed focus distance means that the focal length of the camera is set to a specific, non-adjustable value, thereby limiting the range of distances of objects that can be clearly imaged when shooting.
[0060] This design simplifies the structure of the camera module, helps to make the electronic device thin and light, and also reduces manufacturing costs.
[0061] In a fixed focus camera, the focal length of the camera is fixed. Focal length is a measure of the convergence or divergence of light in an optical system. It refers to the distance from the optical center of the lens to the focus point where the light converges when parallel light is incident. Focal length is an important characteristic of a camera. In addition, a fixed focus camera cannot change the focal length to achieve focus by moving the lens position like an autofocus camera.
[0062] When using an electronic device with a fixed-focus lens, users cannot change the focus distance by adjusting the position of the lens in the camera. They can only move the electronic device itself closer to or farther away from the subject so that the subject is clear in the image captured by the camera. If the distance between the subject and the camera is too large or too small, the image captured by the camera will be out of focus and blurred. Out of focus and blur will directly affect the visual effect of the image, making the edges of the objects in the image blurred and the details lost.
[0063] In order to improve the clarity of the image displayed to the user, the image captured by the camera can be processed by a deblurring model to obtain a clearer image. The deblurring model can be a trained neural network model.
[0064] Alternatively, the blur degree of the image captured by the camera is estimated, and based on the blur degree, the image can be deblurred to obtain a clearer image.
[0065] Regardless of whether the image captured by the camera is processed through a deblurring model or the image is deblurred based on the blurriness of the image captured by the camera, the clarity of the resulting image is relatively low, and users are less satisfied with the processed image.
[0066] In order to further improve the clarity of the image obtained by deblurring, the present application embodiment provides an image processing method. The image processing method provided by the present application embodiment can be applied in an electronic device. Figures 1 to 3 The application scenario of the image processing method provided in the embodiment of the present application is described.
[0067] Figure 1 (a) in FIG. 1 shows a graphical user interface (GUI) of an electronic device, which is a desktop 1110 of the electronic device. The GUI provides a way for a user to interact with a computer application through graphical elements (such as windows, icons, menus, and buttons). When the electronic device detects that a user clicks on a camera control 1111 of a camera application (application, APP) on the desktop 1110, the camera application can be started and a screen such as a video screen can be displayed. Figure 1 Another GUI shown in (b). Figure 1 The GUI shown in (b) may be referred to as a photographing interface 1120 .
[0068] The shooting interface 1120 may include a viewfinder 1121 and a camera control 1122. A preview image may be displayed in real time in the viewfinder 1121. The preview image may be a visible image determined based on an original image file output in real time by a camera in the electronic device.
[0069] When it is detected that the user clicks the photo control 1122, the camera in the electronic device can capture an image. The captured image can be, for example, Figure 2 The electronic device can process the first image 1211 by the method provided in the embodiment of the present application to obtain the target image 1132, and store the target image 1132 in the album.
[0070] The shooting interface 1120 may also include a photo control 1123. When detecting that the user clicks the photo control 1123, the electronic device may display the following Figure 2 FIG. 1 is a first album interface 1130 shown in (b) of FIG. 1 . The first album interface 1130 includes a target image 1132 stored in the album.
[0071] In some cases, the user may want to save the image captured by the camera. In order to meet the user's needs, when it is detected that the user clicks the photo control 1122 in the shooting interface 1120, the electronic device captures the first image through the camera and stores the first image in the album.
[0072] When it is detected that the user clicks the photo control 1123 in the shooting interface 1120, or when it is detected that the user clicks the album control 1112 in the desktop 1110, the electronic device may display the following Figure 3 The second album interface 1210 is shown. The second album interface 1210 may include a first image 1211 and a deblur control 1212 .
[0073] When it is detected that the user clicks the deblur control 1212, the electronic device can process the first image 1211 by the method provided in the embodiment of the present application to obtain the target image 1132. After processing the first image 1211, the electronic device can display the target image 1132 as shown in FIG. Figure 2 The second album interface 1130 shown in (b) includes a target image 1132. The electronic device can store the processed target image 1132 in the album.
[0074] Figure 4 It is a schematic structural diagram of an image processing model provided in an embodiment of the present application.
[0075] The image processing model 400 includes a face detection model 410 , a cropping module 420 , a deblurring model 430 and a fusion module 440 .
[0076] The face detection model 410 is used to perform face detection on the first image to obtain a first face image.
[0077] Face detection refers to determining whether an image records a face, and if so, determining the area where the face is located.
[0078] The first face image may be an image in a region of the first image where the face is located.
[0079] The first image may be an image captured by the first camera. The focus distance of the first camera is fixed. The first camera may be configured in an electronic device provided with the image processing model 400. Alternatively, the first camera may also be configured in other electronic devices other than the electronic device provided with the image processing model 400.
[0080] The cropping module 420 is used to crop the depth image according to the face area where the first face image is located in the first image to obtain a depth face image.
[0081] The first image and the depth image may be acquired by capturing the same scene. The depth face image may represent the depth information of the face recorded in the first face image.
[0082] The depth image may be acquired by a second camera. The second camera and the first camera may be configured in the same or different electronic devices.
[0083] The deblurring model 430 is used to perform deblurring processing on the first facial image based on the depth facial image of the first facial image to obtain a second facial image.
[0084] The network structure of the deblurring model 430 may be a visual geometry group network (VGG), a residual network (ResNet), a generative adversarial network (GAN), a transformer, a diffusion model, etc.
[0085] VGG is a deep convolutional neural network model that builds a deep network by stacking multiple convolutional layers (usually 3×3 convolution kernels) and pooling layers. The VGG network uses small convolution kernels (such as 3×3) and strides (usually 1 or sometimes 2 for pooling layers) to replace large convolution kernels (such as 5×5 or 7×7) by stacking multiple such convolution layers. This design reduces the number of parameters while increasing the nonlinearity of the decision function and has the same receptive field as large convolution kernels. After each convolution block, the VGG network uses a maximum pooling layer (usually a 2×2 window and a stride of 2) to reduce the dimension of the feature map, thereby reducing the amount of calculation and the number of parameters while retaining important information. After the convolutional and pooling layers, the VGG network uses multiple fully connected layers or global average pooling layers to transform the feature map into the final category score. The VGG network is often used in conjunction with other techniques (such as GAN, variational autoencoder VAE, etc.) to achieve more complex image generation tasks.
[0086] GAN is a deep learning model. The model includes at least two modules: one is the generative model, and the other is the discriminative model. These two modules learn from each other through game, so as to produce better output. Both the generative model and the discriminative model can be neural networks, specifically deep neural networks or convolutional neural networks. Taking the generation of pictures as an example, in the process of training the generative adversarial network, the goal of the generative model G is to generate real pictures as much as possible to deceive the discriminative model D, and the goal of the discriminative model D is to try to distinguish the pictures generated by the generative model G from the real pictures. In this way, the generative model G and the discriminative model D constitute a dynamic "game" process, which is the "confrontation" in the "generative adversarial network". The final result of the game is that, under ideal conditions, the generative model G can generate pictures that are "fake enough to be mistaken for real", while the discriminative model D is difficult to determine whether the pictures generated by the generative model G are real. In this way, an excellent generative model G is obtained, which can be used to generate pictures.
[0087] The Transformer model consists of an encoder and a decoder. In image processing, the image data is first converted into a series of feature vectors. These feature vectors can be viewed as a special type of sequence in which each position contains part of the image information. The encoder uses a self-attention mechanism to capture the dependencies between different positions in the image. The decoder can predict and generate the next part of the image based on the features extracted by the encoder and the generated image parts, thereby gradually building a complete image.
[0088] The use of diffusion models involves forward diffusion process and reverse diffusion process. The forward diffusion process can also be called forward diffusion process or diffusion process, etc. The reverse diffusion process can also be called reverse diffusion process, abstraction process, denoising process or reverse sampling process, etc. In the forward diffusion process, a Markov chain of samples is generated by slowly adding noise. In the reverse diffusion process, a Markov chain of samples is generated by slowly removing noise. The data obtained from the forward diffusion process can be used to train the diffusion model. The reverse diffusion process can be understood as the inference process of the diffusion model.
[0089] The fusion module 440 is used to fuse the second face image with the image of the first image outside the face area to obtain a target image.
[0090] The fusion module 440 fuses the second face image with the image outside the face area of the first image by splicing the second face image with images of other areas outside the face area of the first image to obtain a target image.
[0091] The second face image is different from the first face image. When the second face image is directly spliced with an image outside the face region of the first image, the spliced image may have obvious seams or unnatural transitions.
[0092] The fusion module 440 may determine the pixels in the edge region of the target image based on the fusion weight, the pixels in the edge region of the face region in the second face image, and the pixels in the edge region in the first image. The fusion weight may be preset.
[0093] By adjusting the pixels in the edge area of the face area, the target image after stitching can achieve a seamless connection effect, thereby improving the quality and viewing experience of the image.
[0094] The deblurring model 430 in the image processing model 400 may be obtained through training. Figures 5 to 7 , the training process of the deblurring model is explained.
[0095] Figure 5 It is a schematic flowchart of a deblurring model training method provided in an embodiment of the present application. Figure 5 The training method shown includes steps S510 to S530.
[0096] Step S510, obtaining multiple groups of first training data, each group of first training data includes a first training sample image of a first training object, a first reference image and training distance information, wherein the training distance information represents the distance between the first training object and a first training camera that captures the first reference image, the camera focus distance corresponding to the first training sample image is equal to the focus distance of the first camera, and the first reference image is an image captured after the first training camera focuses on the first training object.
[0097] The first reference image may also be referred to as a ground truth (GT) in the first training data.
[0098] The first reference image and the first training sample image may be acquired from the same scene, and the first training object is located in the scene. That is, the first reference image and the first training sample image both record the first training object. The first reference image and the first training sample image may both be understood as images of the first training object.
[0099] The training distance information indicates the distance between the first training object and the first training camera that collects the first reference image. Therefore, the training distance information can also be understood as the training distance information of the first training object.
[0100] It should be understood that, for different groups of first training data, the first training objects may be the same or different, and the first training cameras may be the same or different cameras.
[0101] The first training camera focuses on the first training object and then collects the first reference image. The first training camera focuses on the first training object by adjusting the distance between the lens and the image sensor in the first training camera.
[0102] By focusing the first training camera on the first training object, the first reference image can be the clearest image of the first training object recorded. In other words, by focusing the first training camera on the first training object, the focusing distance of the first training camera can be equal to the distance between the first training camera and the first training object.
[0103] Alternatively, by focusing the first training object through the first training camera, the first training object can be located within the distance range represented by the depth of field of the first training camera. The distance range represented by the depth of field of the first training camera can be understood as the range of distances from the first training camera.
[0104] Depth of field (DOF) refers to the distance range in front of and behind the subject that the camera can obtain an acceptable and clear image. Simply put, a clear image can be formed in the range before and after the focus, and this distance range before and after is called the depth of field. Among them, the clear range between the subject and the camera is the front depth of field, and the clear range between the subject and the background is called the back depth of field.
[0105] When the object is within this depth of field, that is, the distance between the object and the camera is within the depth of field of the camera, the camera can clearly capture the image of the object. However, if the object is too close or too far away from the camera, the distance between the object and the camera exceeds the depth of field, then the image captured by the camera in the area where the object is located may become blurred.
[0106] The range of depth of field is affected by the permissible circle of confusion. In photography, when light is focused through the lens, due to the limitations of the lens' optical performance and differences in shooting conditions, the actual image on the image sensor is not a perfect point, but a circle of a certain size. If the diameter of this circle is smaller than the human eye's ability to distinguish, then within a certain range, this blur cannot be visually recognized, and this unrecognizable blur is called the permissible circle of confusion. The diameter of the permissible circle of confusion is the permissible circle of confusion diameter. In photography, the permissible circle of confusion diameter is an important indicator for determining whether an image is clear. The range of depth of field can be such that the radius of the circle formed by imaging any point on the subject on the image sensor is less than or equal to the permissible circle of confusion diameter.
[0107] Depth of field includes two parts: front depth of field and back depth of field.
[0108] The front depth of field refers to the part of the depth of field that is in front of the focus target (i.e., the side close to the camera). The back depth of field refers to the part of the depth of field that is behind the focus target (i.e., the side far from the camera). The focus target can be understood as an object whose distance from the image sensor in the camera is equal to the object distance of the camera.
[0109] That is to say, there is a certain distance between the camera lens and the focus target where the scene is clear, and this range is the foreground depth of field. The foreground depth of field represents the distance between the focus point of the camera and the position corresponding to the minimum acceptable clarity on the side close to the camera. The focus point of the camera and the position corresponding to the minimum acceptable clarity are both located on the axis of the camera lens. The minimum acceptable clarity is the clarity that cannot be recognized visually, and can be a value set based on experience.
[0110] There is a certain distance between the focus target and the background where the objects are clear. This range is the depth of field. The depth of field represents the distance between the focus point of the camera and the position corresponding to the minimum acceptable clarity on the side away from the camera.
[0111] Generally, the front depth of field is smaller than the back depth of field. That is to say, after precise focusing, in the image captured by the camera, the objects within a shorter distance in front of the focus point can be clearly imaged, while the objects within a longer distance behind the focus point are all clear. For a camera with a fixed focus distance, the focus point can be understood as a point whose distance from the image sensor in the camera is the focus distance.
[0112] That is to say, by focusing on the first training object through the first training camera, the focus point of the first training camera can be adjusted so that the first training object is located at the focus point of the first training camera, or, when the first training object is located between the focus point and the training camera, the distance between the first training object and the focus point of the first training camera is less than or equal to the foreground depth of field of the first training camera, and when the first training object is located on the side of the focus point away from the training camera, the distance between the first training object and the focus point of the first training camera is less than or equal to the back depth of field of the training camera.
[0113] It should be understood that the depth of field of the first training camera may change as the focus point of the first training camera changes. That is, different focus points may correspond to different depths of field.
[0114] Step S520 : Deblurring the first training object image based on the training distance information by using the initial deblurring model to obtain a third training object image.
[0115] The training distance information may be determined based on a training object depth image. The training object depth image may be obtained by performing depth estimation on the first training sample image. Depth estimation refers to estimating the distance of each pixel in an image relative to a shooting source through one or more images from different viewing angles.
[0116] The depth estimation of the first training sample image can be understood as monocular depth estimation. Monocular depth estimation refers to depth estimation using a monocular camera, that is, depth estimation through a color image. This method relies on clues in the image, such as perspective deformation, occlusion relationship, texture change, etc., to infer depth information. Although monocular depth estimation is relatively simple in data acquisition, its accuracy and robustness are relatively low due to the lack of multi-view information.
[0117] Alternatively, the training object depth image may be obtained through a fourth training camera.
[0118] The fourth training camera and the first training camera used to collect the first reference image respectively collect the scene including the first training object. Through the fourth training camera, a depth image of the training object can be obtained.
[0119] The image captured by the fourth training camera may be a color image or a depth image.
[0120] When the image captured by the fourth training camera is a color image, the depth image may be determined by performing binocular depth estimation on the color image captured by the fourth training camera and the first training image captured by the first training camera. The first training image is a color image.
[0121] Binocular depth estimation uses two images taken by a binocular camera (or two cameras on the same horizontal line at a certain distance from each other) to estimate depth. By calculating the disparity of corresponding points in the two images (i.e. the horizontal distance difference between the two points) and combining the internal and external parameters of the camera, the depth information of each point in the scene can be accurately calculated.
[0122] When the image captured by the fourth training camera is a depth map, the fourth training camera may be a camera in a time of flight (TOF) module. The depth map determined by the TOF module may be understood as the depth map captured by the fourth training camera.
[0123] Based on TOF technology calculation, the TOF module emits light pulses of specific wavelengths through the transmitter. These lights will be reflected by the object. The brightness of the reflected light depends on the reflectivity of the object surface and the angle and intensity of the light. The reflected light is received by the camera of the TOF module. By measuring the time (or phase difference) required for these lights to be emitted and reflected back to the camera by the object, the depth information of the object can be calculated.
[0124] Step S530: adjusting parameters of the initial deblurring model according to the difference between the third training object image and the first reference image.
[0125] The difference between the third training object image and the first label object image can be expressed as a loss value. The loss value can be calculated according to a loss function. In other words, the loss value is the output value of the loss function.
[0126] In the process of training deep neural networks, because we hope that the output of the deep neural network is as close as possible to the value we really want to predict, we can compare the predicted value of the current network with the target value we really want, and then update the weight vector of each layer of the neural network according to the difference between the two (of course, there is usually an initialization process before the first update, that is, pre-configuring parameters for each layer in the deep neural network). For example, if the predicted value of the network is high, adjust the weight vector to make it predict lower, and keep adjusting until the deep neural network can predict the target value we really want or a value very close to the target value we really want. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or objective function, which are important equations used to measure the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference, so the training of the deep neural network becomes a process of minimizing this loss as much as possible.
[0127] The parameters of the initial deblurring model can be adjusted by using the back propagation (BP) algorithm. Through the BP algorithm, the size of the parameters of the initial deblurring model is corrected during the training process, so that the reconstruction error loss of the initial deblurring model becomes smaller and smaller. Specifically, the forward transmission of the input signal to the output will generate error loss, and the parameters in the initial deblurring model are updated by back propagating the error loss information, so that the error loss converges. The back propagation algorithm is a back propagation movement dominated by error loss, which aims to obtain the optimal model parameters, such as the weight matrix.
[0128] The initial deblurring model after parameter adjustment can be used as the deblurring model. Exemplarily, when multiple cameras are provided in the electronic device, the initial deblurring model after parameter adjustment can be used as a candidate model corresponding to the first lens of the multiple cameras.
[0129] In the case where the second training object of each first training data belongs to a certain object category, the initial deblurring model after parameter adjustment can be used as a candidate model corresponding to the certain object category, or can be used as a candidate model corresponding to a combination of the first lens and the certain object category. Exemplarily, in the case where the second training object of each first training data is a face, the initial deblurring model after parameter adjustment can be used as the deblurring model 430 in the image processing model 400.
[0130] In the case where the second training objects of multiple first training data belong to multiple object categories, the initial deblurring model with adjusted parameters can be used as a deblurring model to deblur the first object image that records the target object belonging to any one of the multiple object categories based on the distance information of the first object image.
[0131] When multiple cameras are provided in the electronic device, the focus distances of the multiple cameras may be different. The multiple sets of first training data acquired in step S510 may be multiple sets of training data corresponding to the first camera. The first camera may be any one of the multiple cameras.
[0132] That is, for each camera, a candidate model corresponding to the camera can be trained. Multiple cameras in the electronic device can correspond to multiple candidate models.
[0133] The candidate model corresponding to any one of the multiple cameras may be obtained by training the initial deblurring model according to the multiple sets of first training data corresponding to the any one of the cameras. Each set of first training data in the multiple sets of first training data corresponding to the any one of the cameras includes a first training sample image of a first training object, a first reference image, and training distance information. The focus distance of the any one of the cameras corresponding to the first training sample image is equal to the focus distance of any one of the cameras. The training distance information indicates the distance between the first training object and the first training camera that captures the first reference image. When the first training camera captures the first reference image, the distance indicated by the training distance information may belong to the distance range indicated by the depth of field of the first training camera.
[0134] It should be understood that when the focus distance of the first training camera is equal to the focus distance of any camera, the depth of field of the first training camera may be equal to or different from the depth of field of any camera.
[0135] In some embodiments, the first training sample image may be obtained by capturing the scene recorded in the first reference image through a fifth training camera. The focus distance of the fifth training camera may be equal to the focus distance of the first camera, so that the focus distance corresponding to the first training sample image captured by the fifth training camera may be understood as the focus distance of the first camera.
[0136] The fifth training camera and the first training camera may be the same or different cameras. Exemplarily, the focus distance of the fifth training camera may be fixed. Alternatively, the focus distance of the fifth training camera may be adjustable, and the focus distance of the fifth training camera is equal to the focus distance of the first camera during acquisition to obtain the first training sample image.
[0137] When the first training sample image is acquired by the fifth training camera, the first training sample image may be an image acquired by the fifth training camera, or an image of a training object area located at the first training object in the first training image acquired by the fifth training camera.
[0138] The first training camera and the fifth training camera are used to capture images of the same scene, the image captured by the fifth training camera is the first training image, and the image captured by the first training camera is the second training image. The first training object is located in the scene.
[0139] The first training image may be used as a first training sample image, and the second training image may be used as a first reference image.
[0140] Alternatively, the first reference image may be identified or the first reference image may be manually marked to obtain the training object region where the first training object is located. The image in the first training image located in the training object region may be used as the first training sample image. The image in the second training image located in the training object region may be used as the first reference image.
[0141] The first reference image can be understood as a quasi-focus image. In the imaging process, by adjusting the focal length or using other technical means, the object is made to appear the clearest and sharpest in the captured image. In this case, the captured image is a quasi-focus image.
[0142] When the distance between the first training object and the first training camera is not equal to the focusing distance of the fifth training camera, the first training sample image acquired by the fifth training camera can be understood as an out-of-focus image.
[0143] When the first training camera and the fourth training camera respectively capture images of a scene including the first training object, and the first reference image is an image located in the training object area in the second training image, a training depth image can be determined according to the capture result of the fourth training camera. The training depth image is used to represent the distance (depth) value of each point in the scene. Among the training depth images, the image located in the training object area can be used as the training object depth image.
[0144] In some other embodiments, the first training sample image may be obtained by blurring the first reference image using a target blur kernel.
[0145] The target blur kernel is a blur kernel corresponding to the target distance range to which the training distance information belongs, determined from among the multiple blur kernels corresponding to the focus distance of the first camera. Different distance ranges correspond to different blur kernels.
[0146] That is to say, according to the correspondence between the multiple distance ranges and the multiple fuzzy kernels, the fuzzy kernel corresponding to the target distance range to which the training distance information belongs can be determined as the target fuzzy kernel.
[0147] Blur kernel is a convolution kernel. The parameters of blur kernel can be expressed as a matrix. The processing of image by blur kernel can be understood as convolution of the image with the parameters of blur kernel. Blur kernel can obtain blurred image after processing clear image.
[0148] The blur kernel corresponding to each distance range processes an image, and the distance between the object recorded in the image and the camera that captured the image belongs to the distance range. The image obtained by processing the image with the blur kernel corresponding to a distance range is used to simulate the image acquired by a camera pair with a focus distance equal to the focus distance of the first camera and an object with a distance from the camera within the distance range. The image obtained by processing the blur kernel corresponding to a distance range can accurately reflect the blur degree of the image acquired by a camera pair with a focus distance equal to the focus distance of the first camera and an object with a distance from the camera within the distance range. The blur degree can be understood as the degree of out-of-focus blur.
[0149] The multiple blur kernels may all be obtained through training. The training process of the blur kernel may be implemented through iterative optimization. Any blur kernel among the multiple blur kernels is obtained by training the initial blur kernel based on multiple sets of second training data, and the training of the initial blur kernel is to adjust the parameters of the initial blur kernel. The multiple sets of second training data used to train any blur kernel may be understood as data corresponding to the any blur kernel.
[0150] Each set of second training data includes a second training object image and a second reference image of a second training object. The second training object image and the second reference image are respectively obtained by a second training camera and a third training camera that are at equal distances from the second training object. The training distance between the third training camera and the second training object belongs to the distance range corresponding to any blur kernel. When acquiring the second training object image, the focus distance of the second training camera is equal to the focus distance of the first camera. The first reference image is an image acquired after the third training camera focuses on the second training object.
[0151] The range sizes of the multiple distance ranges corresponding to the multiple blur kernels may be equal or unequal, and the multiple distance ranges may not overlap.
[0152] The second training camera and the third training camera may be the same or different cameras. For example, the second training camera may be a camera with a fixed focus distance, and the third training camera may be a camera with an adjustable focus distance.
[0153] Exemplarily, both the second training camera and the third training camera may be cameras with adjustable focus distances. When the focus distance of the third training camera with adjustable focus distance is the focus distance of the first camera, the third training camera may serve as the second camera. The second training camera and the third training camera may be the same or different cameras.
[0154] like Figure 6 As shown in (a), the second reference image is acquired for an object within a certain distance range. and the second training object image , the second reference image is blurred by the initial blur kernel 611 Perform deblurring to obtain the fourth training object image , and according to the fourth training object image and the second training object image The difference between the values of θ and θ is used to adjust the parameters of the initial blur kernel 611. The adjusted initial blur kernel can be used as the blur kernel corresponding to the certain distance range.
[0155] The difference between the fourth training object image and the second training object image may be represented as a loss value.
[0156] The parameters of the initial blur kernel are expressed as , then the fourth training object image It can be expressed as ,in, represents a convolution operation. The process of adjusting the parameters of the initial blur kernel according to the difference between the fourth training object image and the second training object image can be expressed as: ,in, represents the norm, Indicates when the function reaches its minimum , It can be understood as a loss function.
[0157] The initial blur kernel can be understood as a blur kernel obtained by parameter initialization. The parameters of the initial blur kernel can obey a Gaussian distribution. The standard deviation (σ) of the parameters of the initial blur kernel can be 1.
[0158] In the process of training a deblurring model corresponding to a certain object category, the object category to which the second training object in the training process belongs may be the same as or different from the certain object category.
[0159] When the second training objects in multiple groups of second training samples used in the training process of the blur kernel corresponding to each distance range belong to multiple object categories, the blur kernel corresponding to the distance range can blur the first reference image to obtain the first training sample image, which can be applied in the candidate model training process corresponding to any object category.
[0160] That is to say, at least one first training sample image used in the training process of candidate models corresponding to different object categories may be obtained by processing with the same blur kernel.
[0161] When the second training objects in the multiple groups of second training samples used in the training process of the blur kernel corresponding to each distance range belong to a certain object category, the blur kernel corresponding to the distance range can blur the first reference image to obtain the first training sample image which can be applied in the training process of the candidate model corresponding to the certain category.
[0162] That is to say, the first training sample images used in the candidate model training process corresponding to different object categories are all obtained by processing with different blur kernels. The blur kernel used to process the first training sample images in the candidate model training process corresponding to any object category, the second training objects in the multiple sets of second training data used in the training process all belong to the any object category.
[0163] The blur caused by the deviation of the focus distance between the object and the camera is almost unaffected by the object analogy to which the object belongs. Therefore, at least one first training sample image used in the candidate model training process corresponding to different object categories can be obtained by processing with the same blur kernel, which has little impact on the inference accuracy of the candidate models corresponding to each object category. And for each object category, the blur kernel is no longer trained separately, which can simplify the generation process of the first training data required for candidate model training and improve training efficiency.
[0164] The training process from step S810 to step S830 can be understood as the ability of the model to remove out-of-focus blur for a camera with a fixed focus distance and objects at different shooting distances. Therefore, the trained model can remove the out-of-focus blur of the object in a targeted manner according to the distance information, thereby improving the image clarity.
[0165] Combine the following Figure 7 , taking the category of the first training object being face as an example, the training process of the deblurring model provided in the embodiment of the present application is explained.
[0166] Figure 7 The image processing method shown includes steps S710 to S740.
[0167] Step S710, collecting multiple groups of second training data.
[0168] Each set of second training data includes a second training object image having a second training object recorded therein and a second reference image having a second training object recorded therein. The category to which the second training object belongs is a human face.
[0169] Through the camera with the same first focusing distance and the camera with adjustable focusing distance, multiple scenes including people can be captured respectively. In the multiple scenes, the distances between the people and the camera can be various. The multiple distances between the people and the camera can belong to multiple preset distance ranges. The first focusing distance can be the focusing distance of the first camera. The first camera can be Figure 7 The model trained by the method shown is deployed in the same electronic device. Figure 7 The method shown can train a model corresponding to the first focusing distance.
[0170] The face image in the image captured by the camera at the first focusing distance is used as the second training sample image, and the face image in the image captured by the camera with adjustable focusing distance after focusing is used as the second reference image.
[0171] Exemplarily, when the distances between a person and the camera are respectively a plurality of preset distances, the second training sample image is acquired by acquiring the camera with a first focusing distance, and the second reference image is acquired after focusing the camera with an adjustable focusing distance.
[0172] The plurality of preset distances may belong to a plurality of distance ranges. Each distance range may include at least one preset distance.
[0173] The difference between any two adjacent preset distances may be equal or unequal. The maximum preset distance may be greater than the distance between the far point of the camera of the first focusing distance and the first camera. The minimum preset distance may be less than or equal to the first focusing distance. The far point of the camera is the farthest distance that the camera can clearly capture.
[0174] Fixed focus distance cameras are usually designed to work best within a certain distance range. This specific distance range may be more for medium or long distance objects, but will also cover close objects to some extent. Since the camera has a fixed focal length, it is designed to match the focal length with the shooting distance, so when shooting close distances, as long as the object is within this specific distance range, it is unlikely to be seriously out of focus. Therefore, the minimum preset distance is set to be equal to the first focus distance.
[0175] Different preset distances may belong to different distance ranges. For example, each preset distance may be the midpoint of the distance range to which it belongs. The midpoint of a certain distance range may be understood as the central value or average value of all possible distance values within the certain distance range.
[0176] For example, at the first focusing distance D f When the distance between the person and the camera is 40 cm, the image can be captured at intervals of 10 cm when the distance between the person and the camera is between 40 centimeters (cm) and 120 cm, or when the distance between the person and the camera is between 20 cm and 120 cm, to obtain the second reference image and the second training sample image.
[0177] When the person is at the position corresponding to each preset distance, the camera focuses on the person and then collects the image, and the image in the area where the face is located in the image collected after focusing is used as the second reference image. The focus distance of the camera is set to the first focus distance for image collection, and the image in the area where the face is located in the image collected under the first focus distance is used as the second training sample image.
[0178] It should be understood that the focus of the camera can be achieved through auto focus (AF) or manual focus.
[0179] Step S720, training the blur kernel.
[0180] Based on multiple sets of second training data, the blur kernel can be trained. The blur kernel training process can be found in Figure 6 Description of (a) in FIG.
[0181] It should be understood that the second training sample image and the second reference image obtained when the person is located at a position corresponding to each preset distance range can be used to train blur kernels corresponding to different distance ranges.
[0182] Step S730, construct multiple groups of first training data.
[0183] The blur kernel obtained through training is applied in the process of acquiring the first training data used in the deblurring model training process.
[0184] The face image in the public image set is used as the first reference image. The face image in the public image set can be a clear image. In other words, the face image in the public image set can be an image acquired by focusing a camera on a person.
[0185] The public image set may be a face image set. The public image set may include or exclude training face depth images corresponding to the first reference image. In the case where the image set does not include training face depth images corresponding to the first reference image, depth estimation may be performed on the first reference image to obtain training depth face images corresponding to each first reference image.
[0186] like Figure 6 As shown in (b), according to the correspondence between multiple distance ranges and multiple blur kernels, the blur kernel corresponding to the target distance range to which the depth represented by the training depth face image belongs is determined among the multiple blur kernels as the target blur kernel 612, thereby realizing the selection of the target blur kernel 612 among the multiple blur kernels.
[0187] The first reference image is blurred by the target blur kernel 612 After blurring, the first training face image can be obtained. . The first training face image Can be expressed ,in, are the parameters of the target blur kernel 612 .
[0188] The target blur kernel 612 is used for the first reference image The blur processing may also be understood as a defocused blur fitting of the first reference image recording the object whose depth belongs to the target distance range.
[0189] Out-of-focus blur is a blur effect caused by the object not being on the focal plane of the camera lens. When light passes through the lens, if the object is outside the focal plane, the light cannot form a clear image point on the sensor, but forms a light spot with an area, resulting in blurred imaging.
[0190] Blur processing can also be called blur degradation. Blur degradation is the phenomenon that image quality deteriorates (gets worse) due to various factors.
[0191] Based on each first reference image in the public image set , a set of first training data can be obtained. Each set of first training data includes a first reference image , the first training face image The first training face image obtained by processing the target blur kernel 612 corresponding to the target distance range The degree of defocus blur of the image acquired by the camera at the first focusing distance for an object whose distance from the camera is within the target distance range can be accurately reflected. The model trained with the first training data has higher inference accuracy.
[0192] In the first training data set, the training depth face image may be a normalized training depth face image.
[0193] Normalization is an important step in data preprocessing, which refers to scaling the data so that it falls into a small specific interval, such as [0, 1]. The normalized result d can be expressed as: . Where d represents the normalized training depth face image, The preset maximum distance.
[0194] For example, the preset maximum distance may be equal to the farthest distance between the person and the camera when the second training object image and the second reference image are collected during the process of collecting multiple sets of second training data. That is, when the distance between the person and the camera in the second training sample image and the second reference image is the preset distance, the preset maximum distance may be the largest preset distance among the multiple preset distances.
[0195] Each pixel value of the normalized training depth face image is a normalized depth value. It should be understood that in the training depth face image The original depth value represented by a pixel in is less than or equal to In the case of , the normalized depth value represented by the pixel in the normalized training depth face image can be the original depth value of the pixel and In training deep face images The original depth value represented by a pixel in is greater than or equal to In the case of , the normalized depth value represented by the pixel in the normalized training depth face image is 1.
[0196] Step S740, training the candidate model.
[0197] like Figure 6 As shown in (b), the initial deblurring model 621 is used to deblur the first training face image based on the training depth face image. Perform deblurring to obtain the second training face image .
[0198] Second training face image It can be expressed as .in, represents the processing of the initial deblurring model 621, Represents the normalized training deep face image.
[0199] According to the second training face image and the first reference image The difference between the two is used to adjust the parameters of the initial deblurring model 621. The adjusted initial deblurring model can be used as a candidate model corresponding to the camera with the first focusing distance.
[0200] When the first focusing distance is the focusing distance of the first camera, the adjusted initial deblurring model is the candidate model corresponding to the first camera.
[0201] For different focus distances, you can Figure 7 The image processing method shown determines the candidate model corresponding to the focus distance.
[0202] pass Figure 7 The candidate models identified by the method shown can be used as Figure 4 The deblurring model 430 in the image processing model 400 can be applied to Figure 8 In the image processing method shown.
[0203] Combine the following Figure 8 The image processing method provided in the embodiment of the present application is described in detail. The execution subject of the method provided in the present application can be an electronic device, or a software / hardware module capable of image processing in the electronic device. For the sake of convenience, the electronic device is used as an example in the following embodiments.
[0204] Figure 8 is a schematic flow chart of an image processing method provided in an embodiment of the present application. The method may include steps S810 to S830. Figure 8 The image processing method shown can be understood as a method for processing an image based on distance information through a deblurring model. These steps are described in detail below.
[0205] Step S810: acquiring a first object image through a first camera of an electronic device, wherein the first object image records a target object, and the first camera is a camera with a fixed focus distance.
[0206] The first object image may be an image captured by the first camera, or may be an image of an object area where the target object is located in the image captured by the first camera. The first object image records the target object, which may also be understood as the first object image including the target object.
[0207] The first object image is a color image or a grayscale image. The pixel value of a color image can be a red, green, blue (RGB) color value, and the pixel value can be a long integer representing the color. For example, the pixel value is 256×Red+100×Green+76×Blue, where Blue represents the blue component, Green represents the green component, and Red represents the red component. In each color component, the smaller the value, the lower the brightness, and the larger the value, the higher the brightness. For a grayscale image, the pixel value can be a grayscale value.
[0208] After acquiring the first image captured by the first camera, object recognition can be performed on the first image to determine the object area where the target object is located in the first image. In the first image, the image located in the object area can be used as the first object image.
[0209] The shape of the object region may be preset or not. For example, the shape of the object region may be a rectangle or a circle. When the shape of the object region is a rectangle, the object region may be referred to as an identification frame or a detection frame. For another example, the edge of the object region may be the edge of the target object in the first image.
[0210] In image processing, regions can be represented in many ways, one of which is a common and intuitive way is through a coordinate system. The coordinate system allows us to accurately locate each point in the image, thereby describing and defining the region. In two-dimensional images, the most commonly used coordinate system is the Cartesian coordinate system, where the image is viewed as a two-dimensional grid, and each grid point (pixel) can be represented as a unique coordinate in the Cartesian coordinate system. In the case where the shape of the object region is rectangular, the object region can be represented by the coordinates of the upper left corner and the coordinates of the lower right corner.
[0211] Object recognition, which can also be called object detection or target detection, is an important task in the field of computer vision. Object recognition mainly includes two core tasks: determining whether an image contains one or more categories of objects, and locating and marking the specific position and size of the one or more categories of objects in the image if the image contains the one or more categories of objects. The position and size of an object can be understood as the area where the object is located.
[0212] When the recognized object category is a face, the recognition of the first image can also be understood as face detection of the first image.
[0213] In some embodiments, the recognition of the first image may be recognition of a specific object category. The first image is recognized to obtain at least one object region. The target object in each object region in the first image belongs to the object category.
[0214] The recognition process of the first image may include feature extraction and regression. By performing feature extraction on the first image, image features of the first image may be obtained. By performing regression on the image features of the first image, an object region of the target object may be obtained.
[0215] In order to recognize a plurality of preset object categories, the first image may be recognized respectively by using different recognition models, and different recognition models correspond to different object categories.
[0216] In some other embodiments, the recognition of the first image further includes classification, that is, determining the object category to which the target object recorded in each object region belongs. The object category to which each target object belongs can be determined based on the image features of the first image.
[0217] Step S820: Acquire distance information of the target object, where the distance information is used to indicate the distance between the target object and the first camera.
[0218] The distance information of the target object may be determined based on the object depth image. The distance information of the target object may be the object depth image. The distance information of the target object may also be the average value of the depths represented by multiple pixels in the object depth image. The distance information of the target object may also be a distance value collected by a distance sensor in an electronic device.
[0219] Exemplarily, the average value of the depths represented by multiple pixels in the object depth image or the distance value acquired by the distance sensor is expanded into a matrix having a size equal to that of the first object image, and the matrix can be used as the distance information of the target object.
[0220] Depth images, also known as range images, are a special way of representing images. Their core feature is to record the distance (depth) from the image collector (such as a camera, lidar, etc.) to each point in the scene as a pixel value. Object depth images can be understood as images that record the depth of the target object.
[0221] The object depth image may be obtained by performing depth estimation on the first object image. The depth estimation on the first object image may be understood as monocular depth estimation.
[0222] Alternatively, the object depth image may be obtained through a second camera of the electronic device.
[0223] In the case where the first object image is an image captured by the first camera, the object depth image may be acquired by the second camera of the electronic device within the scene recorded by the first object image.
[0224] Compared with the average value of the depths represented by multiple pixels in the object depth image as the distance information of the target object, the object depth image obtained by the second camera in the electronic device includes a large amount of distance information related to the target object, which can reflect the characteristics of the target object. Through the deblurring model, the first object image is deblurred based on the object depth image to improve the accuracy and clarity of the processed image.
[0225] Compared with the object depth image obtained by performing depth estimation on the first object image, the object depth image obtained by the second camera has higher accuracy, so that the image obtained by performing the deblurring model on the object depth image and the first object image has higher accuracy.
[0226] In the case where the first object image is an image located in the object area in the first image captured by the first camera, the object depth image may be an image located in the object area in the depth image. The depth image may be obtained by a second camera of the electronic device. The depth image and the first image may record the same scene.
[0227] The second camera may be another camera other than the first camera, or may be a camera in a TOF module of the electronic device. The TOF module is a distance sensor.
[0228] When the second camera is a camera other than the first camera and the second camera is used to capture color images, the first camera and the second camera have different shooting angles for the scene. Based on the image captured by the first camera and the image captured by the second camera, the depth information of each point in the scene can be calculated based on the parallax principle to obtain a depth image.
[0229] When the second camera is a camera in a TOF module, the second camera can capture a depth image of the object.
[0230] Step S830: determining a target image based on the distance information and the first object image by using a deblurring model. The deblurring model can be used to perform deblurring processing on any image based on the distance information of the object in the image.
[0231] The deblurring model performs deblurring on the input image, which can also be understood as image enhancement. Image enhancement is a method of improving image quality and visual effects through specific image processing techniques. The purpose of image enhancement is to make the image more suitable for human observation or computer analysis and processing.
[0232] The process of the deblurring model processing the distance information and the first object image is the reasoning process of the deblurring model.
[0233] In the case where the first object image is an image captured by the first camera, the second object image obtained by processing the distance information and the first object image by the deblurring model can be used as the target image.
[0234] When the first object image is an image located in the object area of the first image captured by the first camera, the second object image obtained by processing the distance information and the first object image by the deblurring model and the image outside the object area in the first image are fused to obtain the target image.
[0235] Exemplarily, when the distance information is an image, the distance information and the first object image may be spliced, and the image obtained after the splicing may be used as an input of the deblurring model.
[0236] Of course, the distance information and the first object image may not be concatenated, but the distance information and the first object image may be used as inputs of the deblurring model respectively.
[0237] The target image may be obtained by splicing the second object image with an image outside the object area in the first image.
[0238] Alternatively, after splicing the second object image with the image outside the object area in the first image, the pixels at the junction of the second object image and the image outside the object area in the first image can be adjusted to make the splicing more natural, so as to obtain a more beautiful target image.
[0239] Adjustment of pixels at the junction may include edge feathering or image fusion algorithm.
[0240] In the first image, the number of target objects recognized can be multiple, and different target objects are located in different object areas. In other words, different first object images record different target objects.
[0241] For each first object image and the distance information of the target object recorded in the first object image, after being processed by the deblurring model corresponding to the target object, a second object image recording the target object can be obtained. In other words, for each target object, the first object image recording the target object and the distance information of the target object can be processed by the deblurring model to obtain the second object image of the target object.
[0242] The target image can be obtained by fusing the plurality of second object images with images outside the object regions where the target objects are located in the first image.
[0243] It should be understood that the object regions where multiple target objects are located may or may not overlap.
[0244] Before performing step S830, a deblurring model may be determined. In some embodiments, for images of objects belonging to different object categories, the images may be deblurred using different candidate models. The correspondence between multiple candidate models and multiple object categories may be represented according to the second correspondence.
[0245] The target object belongs to the target object category. According to the second corresponding relationship, a candidate model corresponding to the target object category can be determined, and the defuzzification model is a candidate model corresponding to the target object type.
[0246] Objects belonging to different object categories have different characteristics. Setting candidate models corresponding to different object categories and using the candidate models corresponding to the corresponding categories of each object to deblur the object can make the processing results more accurate and clearer.
[0247] When there are multiple target objects in the first image, the deblurring model used to process the first object image of each target object and the distance information of the target object may be a candidate model corresponding to the object category to which the target object belongs.
[0248] In some other embodiments, the electronic device may be provided with a plurality of cameras. The plurality of cameras are all cameras with fixed focus distances, and the focus distances of the plurality of cameras are different. The plurality of cameras may correspond to a plurality of candidate models.
[0249] The first correspondence relationship represents the correspondence relationship between multiple candidate models and multiple cameras. According to the first object relationship, the candidate model corresponding to the first camera can be determined. The deblurring model is the candidate model corresponding to the first camera.
[0250] Each candidate model may be used to deblur the input image based on distance information of objects in the input image.
[0251] The blur degree of images captured by cameras with different focus distances for objects at the same depth or the same distance from focus is different. Setting different candidate models for cameras with different focus distances in electronic devices and deblurring images captured by a camera using the candidate model corresponding to the camera can make the processing result more accurate and the processed image clearer.
[0252] In the case where the electronic device may be provided with multiple cameras, the deblurring model may still be determined according to the object category.
[0253] According to the third correspondence, the model group corresponding to the first camera can be determined. The third correspondence represents the correspondence between multiple cameras and multiple model groups. According to the fourth correspondence of the model group corresponding to the first camera, the candidate model corresponding to the target object category to which the target object belongs in the model group corresponding to the first camera can be determined. The fourth correspondence of each model group represents the correspondence between multiple object categories and multiple candidate models in the model group. The deblurring model can be the candidate model corresponding to the target object category in the model group corresponding to the first camera.
[0254] Alternatively, the deblurring model may be determined according to the object relationship between the combination of multiple cameras and object categories and multiple candidate models. The deblurring model may be a candidate model corresponding to the combination of the first camera and the target object category.
[0255] The multiple object categories may include the object categories that people pay most attention to. For example, the multiple object categories may include human faces. The multiple object categories may also include animals such as cats, dogs, rabbits, etc., and may also include plants, vehicles, buildings, homes, clothing, etc.
[0256] The first object image recording the object of the object category that people pay most attention to is deblurred to obtain a clearer image, thereby improving the user's satisfaction with the processed image.
[0257] Before performing step S830, it may also be determined whether a difference between the distance represented by the distance information and a focus distance of the first camera is greater than or equal to a first difference threshold.
[0258] In the case where the distance represented by the distance information is less than the first difference threshold, step S830 may not be performed. If the first object image is an image captured by the first camera, the first object image may be used as the target image. If the number of first object images in the first image is one, the first image may also be used as the target image. If the number of first object images in the first image is multiple, the first object image may be used as the second object image, and the target image may be determined based on other second object images obtained by deblurring other first object images and images outside each object area in the first image.
[0259] When the distance indicated by the distance information is greater than or equal to the first difference threshold, step S830 may be performed.
[0260] When the distance between the target object and the first camera is greater than or equal to the first difference threshold, the clarity of the first object image may be low, and the first object image is deblurred. When the distance between the target object and the first camera is less than the first difference threshold, the first object image is clear, and deblurring is no longer required, thereby obtaining a clear image and reducing the resource occupation of image processing.
[0261] The distance information represents the distance between the target object and the first camera, and specifically may be the distance between the target object and the image sensor in the first camera.
[0262] The first difference threshold is preset and can be determined according to the depth of field of the first camera.
[0263] The first difference threshold may be equal to the size of the range of the foreground depth representation.
[0264] Alternatively, when the distance represented by the distance information is greater than or equal to the focus distance of the first camera, the first difference threshold may be equal to the size of the range represented by the back depth of field. When the distance represented by the distance information is less than the focus distance of the first camera, the first difference threshold may be equal to the size of the range represented by the front depth of field.
[0265] The deblurring model can be trained. The training process of the deblurring model can be found in Figure 5 Description.
[0266] The candidate model corresponding to the combination of any one of the multiple cameras and any one of the multiple object categories can be obtained by training the initial deblurring model according to multiple sets of first training data corresponding to the combination of any one of the multiple cameras and any one of the multiple object categories. The first training object belongs to the any one of the object categories.
[0267] During the model inference process, when the distance between the target object and the first camera is less than the first difference threshold, deblurring may not be performed. In the multiple sets of first training data used in the training process of the deblurring model, first training distance information may be included, in which the difference between the distance represented and the focus distance of the first camera is greater than or equal to the first difference threshold, and second training distance information may be included, in which the difference between the distance represented and the focus distance of the first camera is less than the first difference threshold. During the training process, the deblurring model is trained using the first training data including the second training distance information and the first training data including the first training distance information, so that the deblurring model is more accurate and the accuracy of the image processed by the deblurring model is improved.
[0268] In some embodiments, for any first camera in the electronic device, the multiple candidate models corresponding to the multiple candidate models set in the electronic device may include a candidate model corresponding to the object category of human face. According to the processing of the first object image belonging to the object category of human face by the electronic device, the parameters of the candidate model of the object category of human face may also be adjusted to obtain a candidate model corresponding to at least one person. Different people may be used as different object categories.
[0269] The electronic device may store at least one character image of a character. The distance between the character recorded in each character image and the camera that captures the character image is less than a first difference threshold of the camera. The first difference threshold of each camera may be determined according to the depth of field of the camera.
[0270] For a first object image in which the distance between the recorded target object and the first camera is greater than or equal to the first difference threshold, when the target object is a face, the first object image is processed based on the distance information of the target object by a candidate model corresponding to the face to obtain a second object image.
[0271] The second object image is subjected to face recognition to determine a candidate person corresponding to the target object. Alternatively, the person corresponding to the target object may be determined according to a user operation. The user operation may indicate the candidate person corresponding to the target object.
[0272] Feature extraction is performed on the second object image, and according to the difference between the extracted features and the features of the candidate character, the candidate model corresponding to the face is adjusted to obtain the candidate model corresponding to the face of the candidate character.
[0273] The features of the candidate character can be obtained by extracting features from an image located in the face region of the character image of the candidate character. The feature extraction model for extracting features from the second object image and the feature extraction model for extracting features from the image located in the face region of the character image of the candidate character can be the same model. The feature extraction model can be applied in the process of face recognition.
[0274] Face recognition is the process of determining a person's identity based on their facial features. Each person's facial features are unique, including face shape, facial features, skin texture, etc. By extracting these features and comparing them with a known face database, the system can determine whether the input face matches a face in the database, thereby determining the identity. The feature extraction model is used to extract features from face images. Based on the extracted features, the person corresponding to the face image can be determined. Different people have different identities.
[0275] After determining the candidate model corresponding to the candidate character, the second correspondence relationship representing the correspondence between the multiple candidate models and the multiple object categories can be updated. The updated second correspondence relationship includes the object relationship between the face of the candidate character and the candidate model, and the candidate models corresponding to the faces of other characters other than the candidate character. The candidate models corresponding to the faces of other characters are the candidate models corresponding to the faces in the second object relationship before the update.
[0276] After acquiring the first image captured by the first camera again, the first image is recognized to obtain at least one new first object image. A new target object is recorded in each new first object image. For each new first object image, when it is determined through recognition that the category to which the recorded new target object belongs is not a face, the category to which the target object belongs determined through recognition is the target category to which the target object belongs. When it is determined through recognition that the category to which the recorded new target object belongs is a face, face recognition can be performed on the first object image to determine the person corresponding to the target object. The person corresponding to the target object is the target category to which the target object belongs.
[0277] After determining the target category to which the new target object belongs, the deblurring model may be determined according to the updated second corresponding relationship, and the target image may be determined by processing the deblurring model.
[0278] Through steps S810 to S560, the deblurring model processes the first object image recording the target object acquired by the first camera at a fixed focus distance based on the distance information of the target object to determine the target image. The influence of the distance information of the target object is taken into account during the processing of the deblurring model, so that the clarity of the target image is higher, thereby improving the user experience.
[0279] It should be understood that Figure 8 The method shown can be processed by a CPU, or by a CPU and a GPU together, or by using other processors suitable for neural network calculations without using a GPU, and this application does not impose any restrictions.
[0280] Figure 8 The method described above can be specifically performed by an execution device. Figure 8 The first object image and distance information in the method may be determined based on data collected by the execution device, or the first object image and / or distance information may be input data given by the client device. The execution device and the client device may be different electronic devices.
[0281] Both the execution device and the client device can be terminals, such as mobile terminals, tablet computers, laptop computers, AR / VR, vehicle terminals, etc., or servers or clouds, etc.
[0282] Figure 8 The deblurring model used in the method shown can be a training device execution Figure 5 The training device may be a server or a cloud, etc. The training device may be an electronic device that is the same as or different from the execution device.
[0283] It should be understood that the above examples are intended to help those skilled in the art understand the embodiments of the present application, rather than to limit the embodiments of the present application to the specific numerical values or specific scenarios illustrated. Those skilled in the art can obviously make various equivalent modifications or changes based on the above examples, and such modifications or changes also fall within the scope of the embodiments of the present application.
[0284] Combination of the above Figures 1 to 8 , describes in detail the image processing model, the training method of the deblurring model and the image processing method of the embodiment of the present application, and will be combined with Figures 9 to 11 , describes the device embodiment of the present application in detail. It should be understood that the image processing device in the embodiment of the present application can execute the various data processing methods and training methods in the aforementioned embodiments of the present application, that is, the specific working processes of the following various products can refer to the corresponding processes in the aforementioned method embodiments.
[0285] Fig. 9 A hardware system of an electronic device suitable for the present application is shown.
[0286] The method provided in the embodiments of the present application can be applied to various electronic devices capable of image acquisition, such as mobile phones, tablet computers, wearable devices, laptop computers, netbooks, personal digital assistants (PDAs), and vehicle-mounted devices. The embodiments of the present application do not impose any restrictions on the specific types of electronic devices.
[0287] Fig. 9The structure diagram of the electronic device 100 is shown. The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0288] It is to be understood that the structure illustrated in the embodiment of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown in the figure, or combine some components, or split some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0289] The processor 110 may include one or more processing units, for example, the processor 110 may include an application processor (AP), a modem processor, a graphics processor (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0290] The controller may be the nerve center and command center of the electronic device 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.
[0291] The processor 110 may also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory may store instructions or data that the processor 110 has just used or cyclically used. If the processor 110 needs to use the instruction or data again, it may be directly called from the memory. This avoids repeated access, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0292] The electronic device 100 can realize the shooting function through ISP, camera 193, video codec, GPU, display screen 194 and application processor.
[0293] ISP is used to process the data fed back by camera 193. For example, when taking a photo, the shutter is opened, and the light is transmitted to the camera photosensitive element through the lens. The light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to ISP for processing and converts it into an image visible to the naked eye. ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. ISP can also optimize the exposure, color temperature and other parameters of the shooting scene. In some embodiments, ISP can be set in camera 193.
[0294] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then passes the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device 100 may include 1 or N cameras 193, where N is a positive integer greater than 1.
[0295] The digital signal processor is used to process digital signals, and can process not only digital image signals but also other digital signals. For example, when the electronic device 100 is selecting a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.
[0296] NPU is a neural network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission mode between neurons in the human brain, it can quickly process input data and can also continuously self-learn. NPU can realize applications such as intelligent cognition of electronic device 100, such as image recognition, face recognition, voice recognition, text understanding, etc.
[0297] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement a data storage function, such as storing music, video and other files in the external memory card.
[0298] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area may store data created during the use of the electronic device 100 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0299] The distance sensor 180F is used to measure the distance. The electronic device 100 can measure the distance by infrared or laser. In some embodiments, when shooting a scene, the electronic device 100 can use the distance sensor 180F to measure the distance to achieve fast focusing.
[0300] The software system of the electronic device 100 may adopt a layered architecture, an event-driven architecture, a micro-core architecture, a micro-service architecture, or a cloud architecture. The embodiment of the present application takes the Android system of the layered architecture as an example to exemplify the software structure of the electronic device 100.
[0301] Fig.10 1 is a software structure diagram of the electronic device 100 of the embodiment of the present application. The layered architecture divides the software into several layers, each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, namely, the application layer, the application framework layer, the system library of the Android runtime (Android runtime), and the kernel layer. The application layer can include a series of application packages.
[0302] like Fig.10 As shown, the application layer may include applications such as camera, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, wallpaper, photo album, multimedia editor, etc.
[0303] The application framework layer provides an application programming interface (API) and programming framework for applications in the application layer. The application framework layer includes some predefined functions.
[0304] like Fig.10 As shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like.
[0305] The window manager is used to manage window programs. The window manager can obtain the size of the display screen, determine whether there is a status bar, lock the screen, capture the screen, etc. The content provider is used to store and obtain data and make the data accessible to applications. The view system includes visual controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build applications. The display interface can be composed of one or more views. For example, a display interface including a text message notification icon can include a view for displaying text and a view for displaying pictures. The phone manager is used to provide communication functions for the electronic device 100. The resource manager provides various resources for applications. The notification manager enables applications to display notification information in the status bar, which can be used to convey notification-type messages and can disappear automatically after a short stay without user interaction.
[0306] Android runtime includes core libraries and virtual machines. Android runtime is responsible for scheduling and management of the Android system.
[0307] The core library consists of two parts: one part is the function that needs to be called by the Java language, and the other part is the Android core library.
[0308] The application layer and the application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as object life cycle management, stack management, thread management, security and exception management, and garbage collection.
[0309] The system library may include multiple functional modules, such as surface manager, media libraries, 3D graphics processing library, 2D graphics engine, etc. For example, the system library may also include image processing models, etc.
[0310] The surface manager is used to manage the display subsystem and provide fusion of 2D and 3D layers for multiple applications. The media library supports playback and recording of multiple common audio and video formats, as well as static image files. The media library can support multiple audio and video encoding formats. The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing. The 2D graphics engine is a drawing engine for 2D drawing.
[0311] The kernel layer is the layer between hardware and software. The kernel layer can include driver modules and hardware abstraction layer (HAL). Driver modules in the kernel layer can include display drivers, camera drivers, audio drivers, and sensor drivers. The hardware abstraction layer (HAL) is an interface layer between the operating system kernel and the hardware circuit, and its purpose is to abstract the hardware.
[0312] When the camera application detects that the user clicks the photo control in the photo interface, the camera application can send photo control information to the camera driver in the kernel layer. The camera driver can control the first camera to capture the first image and control the second camera to capture the depth image when receiving the photo control information.
[0313] The HAL may receive a first image captured by a first camera and a target image captured by a second camera, and transmit the first image and the target image to an image processing model in a system library.
[0314] The image processing model processes the first image and the depth image to obtain a target image. The image processing system can transmit the target image to the photo album application. The target image can be transmitted to the photo album application through the application framework layer.
[0315] Fig.11 It is a schematic diagram of an image processing device provided in an embodiment of the present application.
[0316] The image processing apparatus 1000 includes an acquisition unit 1010 and a processing unit 1020 .
[0317] In some embodiments, the image processing apparatus 1000 may be used to perform Figure 8 The image processing method shown in the figure. The image processing apparatus 1000 may be arranged in an electronic device.
[0318] The acquisition unit 1010 is used to acquire a first object image through a first camera of the electronic device, where the first object image records a target object, and the first camera is a camera with a fixed focus distance.
[0319] The acquisition unit 1010 is further used to acquire distance information of the target object, where the distance information is used to indicate the distance between the target object and the first camera.
[0320] The processing unit 1020 is used to determine a target image based on the distance information and the first object image through a deblurring model, and the deblurring model can be used to perform deblurring processing on any image based on the distance information of an object in the image.
[0321] In some other embodiments, the image processing apparatus 1000 may be used to perform Figure 5 Training method for the deblurring model shown.
[0322] The acquisition unit 1010 is used to acquire multiple groups of first training data, each group of first training data includes a first training sample image of a first training object, a first reference image and training distance information, wherein the training distance information represents the distance between the first training object and a first training camera that captures the first reference image, the camera focus distance corresponding to the first training sample image is equal to the focus distance of the first camera, and the first reference image is an image captured after the first training camera focuses on the first training object.
[0323] The processing unit 1020 is configured to perform a deblurring process on the first training object image based on the training distance information by using the initial deblurring model to obtain a third training object image.
[0324] The processing unit 1020 is further configured to adjust parameters of the initial deblurring model according to the difference between the third training object image and the first reference image.
[0325] It should be noted that the above-mentioned image processing device 1000 is embodied in the form of a functional unit. The term "unit" here can be implemented in the form of software and / or hardware, and is not specifically limited to this.
[0326] For example, a "unit" may be a software program, a hardware circuit, or a combination of the two that implements the above functions. The hardware circuit may include an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a dedicated processor, or a group processor, etc.) and a memory for executing one or more software or firmware programs, a combined logic circuit, and / or other suitable components that support the described functions.
[0327] Therefore, the units of each example described in the embodiments of the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present application.
[0328] The present application also provides a chip, the chip comprising a data interface and one or more processors. When the one or more processors execute instructions, the one or more processors read instructions stored in a memory through the data interface to implement the image processing method and / or deblurring model training method described in the above method embodiment.
[0329] The one or more processors may be general-purpose processors or special-purpose processors. For example, the one or more processors may be a central processing unit (CPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices such as discrete gates, transistor logic devices, or discrete hardware components.
[0330] The chip may be used as a component of a terminal device or other electronic devices. For example, the chip may be located in the electronic device 100 .
[0331] The processor and the memory may be provided separately or integrated together. For example, the processor and the memory may be integrated on a system on chip (SOC) of a terminal device. That is, the chip may also include a memory.
[0332] A program may be stored in the memory, and the program may be executed by the processor to generate instructions, so that the processor executes the image processing method and / or the deblurring model training method described in the above method embodiment according to the instructions.
[0333] Optionally, data may be stored in the memory. Optionally, the processor may read data stored in the memory, and the data may be stored at the same storage address as the program, or may be stored at a different storage address than the program.
[0334] Exemplarily, the memory can be used to store relevant programs of the image processing method provided in the embodiments of the present application, and the processor can be used to call relevant programs of the graphics processing method stored in the memory to implement the image processing method of the embodiments of the present application.
[0335] For example, a first object image is acquired through a first camera of the electronic device, wherein the first object image records a target object, and the first camera is a camera with a fixed focus distance; distance information of the target object is acquired, wherein the distance information is used to represent the distance between the target object and the first camera; and a target image is determined based on the distance information and the first object image through a deblurring model, wherein the deblurring model can be used to deblur any image based on the distance information of an object in the image.
[0336] Exemplarily, the memory can be used to store relevant programs of the deblurring model training method provided in the embodiments of the present application, and the processor can be used to call relevant programs of the deblurring model training method stored in the memory to implement the deblurring model training method of the embodiments of the present application.
[0337] For example, multiple groups of first training data are obtained, each group of first training data includes a first training sample image of a first training object, a first reference image and training distance information, wherein the training distance information represents the distance between the first training object and a first training camera that captures the first reference image, the camera focus distance corresponding to the first training sample image is equal to the focus distance of the first camera, and the first reference image is an image captured after the first training camera focuses on the first training object; through an initial deblurring model, based on the training distance information, the first training object image is deblurred to obtain a third training object image; and according to the difference between the third training object image and the first reference image, the parameters of the initial deblurring model are adjusted.
[0338] The chip can be arranged in an electronic device.
[0339] The present application also provides a computer program product, which, when executed by a processor, implements the image processing method and / or deblurring model training method described in any method embodiment of the present application.
[0340] The computer program product may be stored in a memory, for example, a program, which is finally converted into an executable target file that can be executed by a processor after preprocessing, compiling, assembling and linking.
[0341] The present application also provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a computer, the image processing method and / or the deblurring model training method described in any method embodiment of the present application is implemented. The computer program can be a high-level language program or an executable target program.
[0342] The computer-readable storage medium is, for example, a memory. The memory may be a volatile memory or a nonvolatile memory, or the memory may include both a volatile memory and a nonvolatile memory. Among them, the nonvolatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0343] In the description of this application, the terms "first", "second", etc. are used only for descriptive purposes and should not be understood as indicating or implying relative importance, or a specific order or sequence. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0344] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0345] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0346] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0347] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic; for example, the division of the units is only a logical function division, and there may be other division methods in actual implementation; for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0348] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0349] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0350] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. An image processing method, characterized in that: Applied to electronic equipment, the method comprises: Acquire a first object image through a first camera of the electronic device, wherein the first object image records a target object, and the first camera is a camera with a fixed focus distance; Acquire distance information of the target object, where the distance information is used to indicate the distance between the target object and the first camera; Processing the distance information and the first object image by a deblurring model in the electronic device to obtain a second object image, wherein the definition of the second object image is higher than that of the first object image; The deblurring model is obtained by training an initial deblurring model based on multiple groups of first training data, wherein each group of first training data in the multiple groups of first training data includes a first training sample image of a first training object, a first reference image and training distance information, wherein the training distance information represents the distance between the first training object and a first training camera that captures the first reference image, the focus distance corresponding to the first training sample image is equal to the focus distance of the first camera, and the first reference image is an image captured after the first training camera focuses on the first training object. The initial deblurring model is used to process the first training sample image and the training distance information of each first training object to obtain a third training object image of the first training object, and the difference between the third training object image of each first training object and the first reference image of the first training object is used to adjust the parameters of the initial deblurring model to obtain the deblurring model.
2. The method according to claim 1, characterized in that The electronic device comprises a plurality of cameras, each of the plurality of cameras is a camera with a fixed focus distance, and the focus distances of the plurality of cameras are different; Before the distance information and the first object image are processed by the deblurring model in the electronic device to obtain the second object image, the method further includes: According to a first corresponding relationship, a candidate model corresponding to the first camera is determined, the deblurring model is the candidate model corresponding to the first camera, and the first corresponding relationship represents a corresponding relationship between a plurality of candidate models and the plurality of cameras.
3. The method according to claim 1, characterized in that The target object belongs to a target object category, and before the distance information and the first object image are processed by the deblurring model in the electronic device to obtain the second object image, the method further includes: According to the second corresponding relationship, a candidate model corresponding to the target object category is determined, the defuzzification model is the candidate model corresponding to the target object category, and the second corresponding relationship represents a corresponding relationship between multiple candidate models and multiple object categories.
4. The method according to claim 3, characterized in that: The target object category of the target object is a face.
5. The method according to any one of claims 1 to 4, characterized in that The distance information is an object depth image of the target object, and the object depth image is obtained through a second camera of the electronic device.
6. The method according to claim 5, characterized in that The acquiring a first object image by using a first camera of the electronic device includes: Recognize the first image captured by the first camera to determine an object area, where the object area is an area where the target object is located in the first image; In the first image, an image located in the object area is acquired to obtain the first object image; The acquiring the distance information of the target object includes: determining a depth image of the object located in the object area in a depth image corresponding to the first image, wherein the depth image is acquired by collecting the second camera.
7. The method according to claim 6, characterized in that The method further comprises: The second object image is fused with an image of the first image outside the object area to obtain a target image.
8. The method according to any one of claims 1 to 4, 6 and 7, characterized in that: The first training sample image is obtained by blurring the first reference image using a target blur kernel, wherein the target blur kernel is a blur kernel determined from a plurality of blur kernels corresponding to the focus distance of the first camera and corresponding to a target distance range to which the training distance information belongs, and different distance ranges correspond to different blur kernels.
9. The method according to claim 8, characterized in that Any blur kernel among the multiple blur kernels is obtained by training an initial blur kernel based on multiple groups of second training data, each group of second training data among the multiple groups of second training data includes a second training object image and a second reference image of a second training object, the second training object image and the second reference image are respectively obtained by a second training camera and a third training camera that are equal to the distance from the second training object, the training distance between the third training camera and the second training object belongs to the distance range corresponding to any blur kernel, when acquiring the second training object image, the focus distance of the second training camera is equal to the focus distance of the first camera, and the first reference image is an image acquired after the third training camera focuses on the second training object.
10. The method according to any one of claims 1-4, 6, 7, and 9, characterized in that: The processing of the distance information and the first object image by the deblurring model in the electronic device to obtain the second object image includes: when a difference between the distance represented by the distance information and the focus distance of the first camera is greater than or equal to a first difference threshold, processing the distance information and the first object image by the deblurring model to obtain the second object image; Among them, the training distance information in the multiple groups of first training data includes first training distance information and second training distance information, the difference between the distance represented by the first training distance information and the focusing distance of the first camera is greater than or equal to the first difference threshold, and the difference between the distance represented by the second training distance information and the focusing distance of the first camera is less than the first difference threshold.
11. An electronic device, characterized in that: The electronic device comprises a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the electronic device executes the method according to any one of claims 1 to 10.
12. A chip, characterized in that: The method comprises a processor and a data interface, wherein the processor reads instructions stored in a memory through the data interface to implement the method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is used to implement the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and computer readable storage medium
CN115375567A