Image processing method and related device
By shaking the multi-frame images and fusion of images after eye-closing detection sorting, the problem of a large number of people with eye-closing in portrait photography is solved, and the image quality is improved.
Patent Information
- Application Number
- CN202311724650.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-14
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-12-14
AI Technical Summary
In portrait photography scenes, when there are many portraits to be taken, the number of people with eyes closed is prone to occur in the images taken by electronic devices, which cannot meet the needs of users.
An image processing method is proposed, by acquiring multi-frame images, sorting images using jitter processing algorithm and open-eye detection algorithm, and fusing images according to the sorting results to reduce the probability of eye closing in the image.
It effectively reduces the probability of the number of people with eyes closed in the image, improves the quality of the shot, and meets the users' needs for portrait photography.
Smart Images

Figure CN120201291A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of terminals, and in particular, to an image processing method and related devices. Background Art
[0002] An electronic device can support the function of taking portrait photos. Currently, in a portrait photo-taking scenario, when there are many portraits to be taken, there is a problem that a relatively large number of people have their eyes closed in the images taken by the electronic device, which cannot meet the user's needs. Summary of the Invention
[0003] Embodiments of this application provide an image processing method and related devices, which are applied to the technical field of terminals and are beneficial to reducing the probability of portraits having their eyes closed in the captured images.
[0004] In a first aspect, an embodiment of this application proposes an image processing method, which includes: in response to a photographing operation, obtaining N frames of images, where N is an integer greater than 1; obtaining an image sequence after sorting the N frames of images; among them, the N frames of images in the image sequence are sorted according to the fusion weights when the N frames of images are fused, and the sorting of the N frames of images in the image sequence is related to a first sorting result and a second sorting result. The first sorting result is obtained by sorting the N frames of images using a first algorithm for jitter processing, and the second sorting result is obtained by sorting the N frames of images using a second algorithm for eye opening and closing detection; performing image fusion based on the image sequence. In this way, it is beneficial to reduce the probability of having eyes closed in the image.
[0005] In a possible implementation manner, the second sorting result is obtained by the following method: cropping the N frames of images to obtain the cropped N frames of images, where the cropped N frames of images include the eye regions; using the second algorithm to identify the cropped N frames of images to obtain the eye opening and closing results of the N frames of images, and the eye opening and closing results are used to represent the number of open eyes and / or the number of closed eyes in the image, or are used to represent the probability of having open eyes in the image; sorting the N frames of images according to the eye opening and closing results of the N frames of images to obtain the second sorting result. In this way, since the resolution of the cropped N frames of images is less than the resolution of the N frames of images, after cropping the N frames of images and then performing eye opening and closing detection, it is beneficial to save power consumption.
[0006] In a possible implementation manner, the region included in the cropped N frames of images is the face region. The face region may include the eye regions, so the region included in the cropped N frames of images is the face region, which is beneficial to subsequent implementation of eye opening and closing detection.
[0007] In a possible implementation, before obtaining N frames of images in response to a photographing operation, the method further includes: in response to an operation of turning on the camera, obtaining a first image collected by the camera; shrinking the first image to obtain a second image, where the resolution of the second image is less than that of the first image; performing face recognition on the second image to obtain the position information of the face area in the second image; storing the first image in a first queue and storing the position information of the face area in the second image in a first buffer; obtaining N frames of images in response to a photographing operation, including: in response to a photographing operation, obtaining N frames of images from the first queue; in response to a photographing operation, the method further includes: in response to a photographing operation, obtaining the position information of the face area in a target image from the first buffer, where the timestamp of the target image is the same as that of the N frames of images; converting the position information of the face area in the target image to obtain the position information of the face area in the N frames of images; cropping the N frames of images to obtain the cropped N frames of images, including: cropping the N frames of images based on the position information of the face area in the N frames of images to obtain the cropped N frames of images. In this way, since the resolution of the target image is less than that of the N frames of images, performing face recognition on the target image helps reduce power consumption.
[0008] In a possible implementation, the first image is an original image, the second image is a preview image, or the resolution of the second image is less than that of the preview image.
[0009] If the resolution of the second image is less than that of the preview image. In this way, performing face recognition on the second image helps reduce power consumption.
[0010] In a possible implementation, using a second algorithm to recognize the cropped N frames of images to obtain the open / closed eye results of the N frames of images, including: denoising the cropped N frames of images to obtain a denoised image; using a second algorithm to recognize the denoised image to obtain the open / closed eye results of the N frames of images. In this way, performing open / closed eye detection on the denoised image helps improve the accuracy of open / closed eye detection.
[0011] In a possible implementation, the weight of the first sorting result is the first weight, and the weight of the second sorting result is the second weight, and the second weight is greater than the first weight. The second weight being greater than the first weight can indicate that the importance degree of the open / closed eye detection result is higher. In this way, it helps reduce the probability of the fused image showing closed eyes.
[0012] In a possible implementation, the weight of the second sorting result is positively correlated with the number of face areas included in the N frames of images.
[0013] The weight of the second sorting result is variable. The more the number of face regions included in N frames of images, the greater the weight of the second sorting result can be. In this way, the weight of the second sorting result can change according to the change in the number of face regions in N frames of images, with stronger flexibility. At the same time, it is more conducive to reducing the probability of the fused image having closed eyes.
[0014] In a second aspect, embodiments of the present application provide an image processing device, which may be an electronic device, or a chip or a chip system within the electronic device. When the image processing device is an electronic device, the processing unit may be a processor. The image processing device may further include a storage unit, and the storage unit may be a memory. The storage unit is used to store instructions, and the processing unit executes the instructions stored in the storage unit to enable the electronic device to implement an image processing method described in the first aspect or any one of the possible implementation manners of the first aspect. When the image processing device is a chip or a chip system within the electronic device, the processing unit may be a processor. The processing unit executes the instructions stored in the storage unit to enable the electronic device to implement an image processing method described in the first aspect or any one of the possible implementation manners of the first aspect. The storage unit may be a storage unit within the chip (such as a register, a cache, etc.), or a storage unit outside the chip within the electronic device (such as a read-only memory, a random access memory, etc.).
[0015] In a third aspect, embodiments of the present application provide an electronic device, including a processor and a memory. The memory is used to store code instructions, and the processor is used to run the code instructions to execute the method described in the first aspect or any one of the possible implementation manners of the first aspect.
[0016] In a fourth aspect, embodiments of the present application provide a computer-readable storage medium, in which a computer program or instructions are stored. When the computer program or instructions are run on a computer, the computer is enabled to execute the method described in the first aspect or any one of the possible implementation manners of the first aspect.
[0017] In a fifth aspect, embodiments of the present application provide a computer program product including a computer program. When the computer program is run on a computer, the computer is enabled to execute the method described in the first aspect or any one of the possible implementation manners of the first aspect.
[0018] Sixth aspect, the present application provides a chip or a chip system, the chip or the chip system includes at least one processor and a communication interface, the communication interface and the at least one processor are interconnected by a line, and the at least one processor is configured to run a computer program or instruction to execute the method described in the first aspect or any possible implementation manner of the first aspect. Wherein, the communication interface in the chip may be an input / output interface, a pin or a circuit, etc.
[0019] In a possible implementation, the chip or the chip system described above in the present application further includes at least one memory, and instructions are stored in the at least one memory. The memory may be a storage unit inside the chip, for example, a register, a cache, etc., or it may be a storage unit of the chip (for example, a read-only memory, a random access memory, etc.).
[0020] It should be understood that the second aspect to the sixth aspect of the present application correspond to the technical solutions of the first aspect of the present application, and the beneficial effects obtained by each aspect and the corresponding feasible implementation manners are similar, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 A schematic diagram of a portrait photographing scenario provided by an embodiment of the present application;
[0022] Figure 2 A schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application;
[0023] Figure 3 A schematic diagram of the software architecture of an electronic device provided by an embodiment of the present application;
[0024] Figure 4 A schematic block diagram of an image processing method provided by an embodiment of the present application;
[0025] Figure 5 A schematic flowchart of an image processing method provided by an embodiment of the present application;
[0026] Figure 6 A schematic block diagram of a chip provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] For the convenience of clearly describing the technical solutions of the embodiments of the present application, the following description is first made:
[0028] In the embodiments of the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and effects. For example, the first algorithm and the second algorithm are only used to distinguish different algorithms, and do not limit their sequence. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and "first", "second", etc. do not necessarily mean different.
[0029] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific manner.
[0030] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (s) or plural items (s). For example, at least one (item) of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0031] Currently, in the portrait photography scenario, when there are many portraits to be photographed, there is a problem that a large number of people have their eyes closed in the images captured by the electronic device, which cannot meet the user's needs.
[0032] Exemplarily, Figure 1 shows a schematic diagram of a portrait photography scenario. As Figure 1 shown, in response to the user's operation of opening the camera application, the electronic device can display Figure 1 interface a in. As Figure 1 shown in interface a in, the interface includes an image captured by the camera, a photographing control 101, and a control 102 for viewing the captured image. Among them, the image captured by the camera includes user A and user B, and both user A and user B have their eyes open.
[0033] In response to the user's operation of triggering the photographing control 101, the electronic device captures an image through the camera and can display Figure 1 interface b in. As Figure 1As shown in interface b in FIG. 1 , the interface includes thumbnails 103 of images captured by the camera. The user can trigger the control 102 for viewing the captured images to view the images captured by the camera. In response to the user triggering the control 102 for viewing the captured images, the electronic device can display Figure 1 In the c interface. Figure 1 As shown in interface c, the interface includes an image captured by a camera. In the image, user A is in a closed-eye state and user B is in an open-eye state, which does not meet the needs of user A.
[0034] In some implementations, the electronic device can obtain multiple frames of original images in response to a photo taking operation. The electronic device can select a reference frame image and a suboptimal frame image from the multiple frames of original images based on the jitter amount data, use the reference frame image as the main frame, and the suboptimal frame image as the supplementary frame, and fuse the reference frame image and the suboptimal frame image to obtain a final image, so as to ensure the clarity of the photo.
[0035] In the portrait photography scene, when there are many portraits to be photographed, since people tend to blink, the probability that the multiple frames of original images obtained include images of people with their eyes closed is high. If the reference frame image is an image of a person with eyes closed, the final image is likely to be an image of a person with eyes closed, resulting in the final image not being able to meet the needs of users.
[0036] For example, in the above Figure 1 In the example shown, in response to the user triggering the photo control 101, the electronic device can capture multiple frames of original images through the camera, and select a reference frame image from the multiple frames of original images based on the jitter data. If user A is in a closed-eye state in this reference frame image, user A is likely to be in a closed-eye state in the final image, resulting in failure to meet the user's needs.
[0037] In view of this, an embodiment of the present application provides an image processing method and related devices. In response to a photo-taking operation, multiple frames of original images can be acquired. The electronic device can select a reference frame image and a suboptimal frame image from the multiple frames of original images based on jitter data and eye opening and closing data. The probability that the final image obtained through the reference frame image and the suboptimal frame image includes a person's closed eyes is relatively small, which is beneficial to reducing the probability of closed eyes in the image.
[0038] The electronic devices in the embodiments of this application may include handheld devices with cameras and in-vehicle devices with cameras, etc. For example, some electronic devices are: mobile phones, tablet computers, handheld computers, laptop computers, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grid, wireless terminals in transportation safety, wireless terminals in smart city, wireless terminals in smart home, cellular phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), handheld devices with wireless communication functions, computing devices or other processing devices connected to a wireless modem, in-vehicle devices, wearable devices, terminal devices in a 5G network, or terminal devices in a future evolved public land mobile network (PLMN), etc. The embodiments of this application are not limited thereto.
[0039] The electronic devices in the embodiments of this application may also be referred to as: terminal devices, user equipment (UE), mobile stations (MS), mobile terminals (MT), access terminals, user units, user stations, mobile stations, mobile handsets, remote stations, remote terminals, mobile devices, user terminals, terminals, wireless communication devices, user agents, or user devices, etc.
[0040] To facilitate the understanding of the embodiments of this application, the following introduces the hardware structure of the electronic devices provided in the embodiments of this application.
[0041] Figure 2 The schematic diagram showing the hardware structure of an electronic device provided in the embodiments of this application is as follows Figure 2As shown, the electronic device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, and a display screen 194, etc.
[0042] Optionally, the above sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0043] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device may include more or fewer components than shown, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0044] The camera 193 may include one or more cameras and may be used to capture images. The images captured by the camera 193 may be stored in the internal memory 121. In the embodiments of the present application, the camera 193 may be used to capture images in a portrait photography scenario.
[0045] The software system of the electronic device may adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. The layered architecture may adopt the Android system, or the iOS system, or other operating systems, and the embodiments of the present application do not limit this. Here, taking the Android system with a layered architecture as an example, the software architecture of the electronic device provided in the embodiments of the present application will be exemplarily described.
[0046] Figure 3 It is a schematic diagram of the software architecture of an electronic device provided in the embodiments of the present application. As Figure 3As shown, the layered architecture can divide the software system of an electronic device into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system can be divided into five layers, from top to bottom, namely the applications layer, the application framework layer, the hardware abstraction layer (HAL), the kernel layer, and the hardware layer.
[0047] The applications layer can include a series of application packages. The applications layer runs applications by calling the application programming interfaces (APIs) provided by the application framework layer. As Figure 3 shown, the application packages can include applications such as the camera and the gallery.
[0048] The application framework layer provides APIs and programming frameworks for the applications in the applications layer. The application framework layer includes some predefined functions. As Figure 3 shown, the application framework layer can include a camera access interface and a view system. Among them, the camera access interface can be used to provide application programming interfaces and programming frameworks for camera applications. The view system includes visual controls, for example, it can include controls for taking pictures (such as the picture-taking control 101 shown in interface a above Figure 1 and controls for viewing the captured images (such as the control 102 for viewing the captured images shown in interface a above Figure 1 ).
[0049] As Figure 3 shown, the HAL layer can include a camera hardware abstraction layer and a camera algorithm library. Among them, the camera hardware abstraction layer can provide virtual hardware for camera devices. The camera algorithm library can include the running code and data for implementing the image processing method provided by the embodiments of the present application.
[0050] The kernel layer is the layer between hardware and software. As Figure 3 shown, the kernel layer can include one or more of the following: a camera device driver, a digital signal processor driver, and an image processor driver, etc. Among them, the camera device driver is used to drive the sensor of the camera to collect images. The digital signal processor driver is used to drive the digital signal processor to process images. The image processor driver is used to drive the graphics processor to process images.
[0051] The hardware layer can include hardware such as a camera and a display screen.
[0052] It should be understood that in some embodiments, layers that implement the same function may be referred to by other names, or a layer that can implement the functions of multiple layers may be regarded as one layer, or a layer that can implement the functions of multiple layers may be divided into multiple layers. The embodiments of the present application do not limit this.
[0053] The following combines the above Figure 3 shown software structure to specifically describe the image processing method in the embodiments of the present application:
[0054] In response to the user's operation of opening the camera application, such as the operation of clicking on the camera application icon, the camera application calls the camera access interface of the application framework layer to start the camera application, and then sends an instruction to start the camera through the camera hardware abstraction layer. The camera hardware abstraction layer sends this instruction to the camera device driver in the kernel layer. The camera device driver can start the corresponding camera to capture images.
[0055] The image processor driver can drive the graphics signal processor to reduce the preview image captured by the camera in real time to a tiny thumbnail. The tiny thumbnail can be transmitted to the camera algorithm library in real time through the image processor driver. The original image captured by the camera in real time can be stored in the zero shutter lag (ZSL) queue in real time. The camera algorithm library executes the image processing method provided by the embodiments of the present application to perform face detection on the tiny thumbnail, obtain the position information of the face frame, and store the position information of the face frame in the buffer. In response to the shooting operation, multiple frames of original images are obtained from the ZSL queue, a reference frame image and a sub-optimal frame image are selected from the multiple frames of original images based on the jitter amount data and the open / closed eye data, and the reference frame image and the sub-optimal frame image are fused to obtain the final image. The camera algorithm library can transmit the final image to the camera hardware abstraction layer. The camera hardware abstraction layer can transmit it to the display screen for display. Among them, the position information of the face frame is used to obtain the face area from multiple frames of images and perform open / closed eye detection on the face area to obtain the open / closed eye data. The resolution of the tiny thumbnail is less than that of the original image. In this way, performing face detection on the tiny thumbnail is beneficial to saving power compared to directly performing face detection on the original image.
[0056] To better understand the image processing method provided by the embodiments of the present application. The following will combine Figure 4 for a detailed description.
[0057] Figure 4 shows a schematic block diagram of an image processing method provided by the embodiments of the present application. As Figure 4 shown, the method includes the following steps:
[0058] S401. When the camera is in the open state, obtain an online real-time tiny thumbnail.
[0059] When the camera is in the open state, the electronic device can obtain the original image captured in real time by the camera, process the original image to obtain a preview image, and perform a reduction process on the preview image to obtain an online real-time tiny thumbnail. The resolution of the preview image is less than that of the original image, and the resolution of the tiny thumbnail is less than that of the preview image.
[0060] In some examples, the resolution of the original image can be 4096*3072, the resolution of the preview image can be 1920*1080, and the resolution of the tiny thumbnail can be 640*480.
[0061] Optionally, the electronic device can perform a reduction process on the preview image through hardware such as the Figure 2 graphics signal processor shown above to obtain an online real-time tiny thumbnail. In this way, compared with software processing, it is beneficial to improve the processing efficiency.
[0062] As Figure 4 shown, the online real-time tiny thumbnail can include small Figure 1 , small Figure 2 , small Figure 3 , and small Figure 4 images, etc.
[0063] It can be understood that each tiny thumbnail in the online real-time tiny thumbnail corresponds to a timestamp, and this timestamp is used to represent the moment when this tiny thumbnail is obtained.
[0064] S402. Input the online real-time tiny thumbnail into the face detection module to obtain the position information of the face frame.
[0065] The online real-time tiny thumbnail can include one face or multiple faces, and the embodiments of the present application do not limit this. If the online real-time tiny thumbnail includes one face, the face detection module can output the position information of one face frame. If the online real-time tiny thumbnail includes multiple faces, the face detection module can output the position information of multiple face frames.
[0066] As Figure 4 shown, in some examples, the face detection module can output the position information of 5 face frames, and the position information of these 5 face frames includes the position information of face frame 1, the position information of face frame 2, the position information of face frame 3, the position information of face frame 4, and the position information of face frame 5.
[0067] The position information of the face frame output by the electronic device can be stored in a buffer.
[0068] In the embodiments of the present application, since the resolution of the tiny thumbnail is less than that of the preview image or the original image, performing face detection on the tiny thumbnail helps to save power consumption.
[0069] S403. When the camera is in the open state, obtain the original image collected by the camera and store the original image in the ZSL queue.
[0070] The ZSL queue can be used to store a fixed number of images, and these images can be updated over time. For example, Figure 4 as shown, in some examples, the ZSL queue includes Raw Image 1 (Raw1), Raw Image 2 (Raw2), Raw Image 3 (Raw3), and Raw Image 4 (Raw4), etc.
[0071] S403 and S401 are executed in parallel. In some implementations, S401 and S403 are implemented through different threads.
[0072] It can be understood that the above S401 to S403 can start to be executed after the camera application is launched.
[0073] It can be understood that each frame of the original image collected by the camera corresponds to a timestamp, and this timestamp is used to represent the moment when this frame of the original image is obtained.
[0074] S404. In response to the operation of taking a photo, select multiple frames of the original image from the ZSL queue through the frame selection module.
[0075] The number of multiple frames of the original image can be less than or equal to the number stored in the ZSL queue. The number of multiple frames of the original image can be preset.
[0076] For example, Figure 4 as shown, the number of multiple frames of the original image can be 3, and RawA, RawB, and RawC are selected from images such as Raw1, Raw2, Raw3, and Raw4. Among them, RawA can be Raw1, RawB can be Raw2, and RawC can be Raw3. Or, RawA can be Raw1, RawB can be Raw2, and RawC can be Raw4. Here, they are not listed one by one.
[0077] The embodiments of the present application provide multiple ways to select multiple frames of the original image from the ZSL queue.
[0078] In a possible implementation, the electronic device starts from the moment of taking a photo and obtains multiple frames of the original image after the moment of taking a photo. In this way, it is suitable for the scenario of photographing a static portrait.
[0079] In another possible implementation, the electronic device takes the photo-taking moment as the midpoint, obtains a part of the image before the photo-taking moment and another part of the image after the photo-taking moment, and obtains multiple frames of original images. In this way, it is suitable for shooting moving portrait scenarios.
[0080] S405. According to the timestamps of each frame of the original images among the multiple frames of original images, the corresponding tiny thumbnail can be obtained from the online real-time tiny thumbnails, realizing the matching of the tiny thumbnails with the original images, and further obtaining the position information of the face frame corresponding to the tiny thumbnails.
[0081] S401 and S403 can be implemented in different threads, and there may be time delays. Matching the tiny thumbnails with the original images through timestamps is beneficial to obtaining more accurate position information of the face frame.
[0082] S406. Convert the position information of the face frame to obtain the converted position information of the face frame, and the converted position information of the face frame is used to represent the position information of the face in the original image.
[0083] The resolutions of the original images and the tiny thumbnails are different. If the position information of the face frame in the original image is to be obtained, the position information of the face frame of the tiny thumbnails needs to be converted.
[0084] S407. Input the multiple frames of original images, the jitter amount of each frame of the original images among the multiple frames of original images, and the converted position information of the face frame corresponding to each frame of the original images among the multiple frames of original images into the frame selection algorithm module to obtain the output of the frame selection algorithm module. The frame selection algorithm module is used to sort the multiple frames of original images and output the optimal frame sequence.
[0085] As Figure 4 shown, the frame selection algorithm module may include a jitter amount frame selection sub-module and an eyes-open / closed frame selection sub-module in the portrait scenario. The jitter amount frame selection sub-module can sort the multiple frames of original images based on the jitter amount of each frame of the original images to obtain the first frame sequence. In some examples, in the first frame sequence, the original image sorted earlier is better than the original image sorted later. Or, the original image sorted later is better than the original image sorted earlier. Or, the original image sorted in the middle is better than the original images sorted earlier and later. The embodiments of the present application do not make any limitations in this regard.
[0086] Among them, the jitter amount may include gyroscope (GYRO) anti-shake and / or optical image stabilization (OIS). As Figure 4 shown, the jitter amount may include GYRO and OIS.
[0087] The open / closed eye frame selection sub-module in the portrait scenario can crop each frame of the original image based on the position information of the converted face frame corresponding to each frame of the original image to obtain the image after face cropping; and perform noise reduction (such as using spatial domain filtering or transform domain filtering, etc.) on the image after face cropping to obtain the image after noise reduction; then perform open / closed eye detection on the image after noise reduction to obtain the open / closed eye detection result. In this way, since the number of pixels in the original image is more than that in the tiny image, using the original image for open / closed eye detection has higher accuracy compared to using the tiny image for open / closed eye detection. In addition, performing open / closed eye detection after face cropping is beneficial to improving the processing efficiency and saving power consumption compared to directly using the original image. Moreover, performing open / closed eye detection on the image after noise reduction is beneficial to further improving the recognition accuracy. The frame selection algorithm module can sort multiple frames of the original image based on the open / closed eye detection results of each frame of the original image to obtain the second frame sequence.
[0088] In some examples, in the second frame sequence, the original image ranked earlier is better than the original image ranked later, or rather, the number of open eyes in the original image ranked earlier is more than the number of open eyes in the original image ranked later (which can be applicable to multiple scenarios), or the probability of having open eyes in the original image ranked earlier is greater than the probability of having open eyes in the original image ranked later (which can be applicable to the scenario of one person).
[0089] In some other examples, in the second frame sequence, the original image ranked later is better than the original image ranked earlier, or rather, the number of open eyes in the original image ranked later is more than the number of open eyes in the original image ranked earlier (which can be applicable to multiple scenarios), or the probability of having open eyes in the original image ranked later is greater than the probability of having open eyes in the original image ranked earlier (which can be applicable to the scenario of one person).
[0090] The frame selection algorithm can comprehensively evaluate the first frame sequence and the second frame sequence and output the optimal frame sequence.
[0091] Exemplarily, each frame of the original image in the first frame sequence can correspond to a first score, each frame of the original image in the second frame sequence can correspond to a second score, the first frame sequence can correspond to a first weight, and the second frame sequence can correspond to a second weight. Based on the first score, the second score, the first weight, and the second weight, the score of each frame of the original image can be calculated, and sorting is performed based on this score to obtain the optimal frame sequence. The original image ranked earlier in the optimal frame sequence is better than the original image ranked later.
[0092] It can be understood that the embodiments of the present application focus on portraits without closed eyes in the captured images. Then, in the embodiments of the present application, the weight of the second frame sequence can be greater than the weight of the first frame sequence. In other scenarios, such as night scenes, the probability of portraits appearing is small, and the weight of the second frame sequence can be less than the weight of the first frame sequence. Or rather, the weights of the first frame sequence and the second frame sequence can vary according to the scenario.
[0093] S408. Use the first frame of the optimal frame sequence as the reference frame, and use the other frames in the optimal frame sequence as sub-optimal frames, and fuse the reference frame image and the sub-optimal frame images to obtain the final image.
[0094] The fusion of the reference frame and the sub-optimal frame images involves frame fusion technology, and the embodiments of the present application do not limit this.
[0095] The image processing method provided by the embodiments of the present application can determine the reference frame and sub-optimal frames by comprehensively considering the jitter amount and the open / closed eye detection result in multiple original images, and can reduce the probability of the portrait having closed eyes in the captured image on the basis of anti-shake.
[0096] As described above in combination with Figure 4 the method provided by the embodiments of the present application, the method provided by the embodiments of the present application will be introduced below in combination with the step flow.
[0097] Exemplarily, Figure 5 is a schematic flowchart of an image processing method provided by an embodiment of the present application. As Figure 5 shown, the method may include the following steps:
[0098] S501. In response to a photographing operation, obtain N frames of images, where N is an integer greater than 1.
[0099] The N frames of images may refer to the above-mentioned multiple frames of images. The N frames of images may be used to represent the original images collected by the camera.
[0100] S502. Obtain an image sequence after sorting the N frames of images; among them, the N frames of images in the image sequence are sorted according to the fusion weights when the N frames of images are fused, and the sorting of the N frames of images in the image sequence is related to the first sorting result and the second sorting result. The first sorting result is obtained by sorting the N frames of images using the first algorithm for jitter processing, and the second sorting result is obtained by sorting the N frames of images using the second algorithm for open / closed eye detection.
[0101] The fusion weight may be a value calculated based on the first sorting result and the second sorting result. The first sorting result is used to represent the sorting of the N frames of images based on the jitter data, the second sorting result is used to represent the sorting of the N frames of images based on the open / closed eye data, and the fusion weight is related to the jitter data and the open / closed eye data.
[0102] S503. Image fusion is performed based on an image sequence.
[0103] The image with the largest fusion weight in the image sequence can be a reference frame image, and the images other than the reference frame in the image sequence can be sub-optimal frame images. The reference frame image and the sub-optimal frame images are fused to obtain a final image.
[0104] The image processing method provided by the embodiments of the present application sorts N frames of images based on jitter data and eye opening and closing data, which is beneficial to reducing the probability of a person's eyes being closed in the captured images on the basis of anti-shake.
[0105] Optionally, the second sorting result is obtained in the following manner: The N frames of images are cropped to obtain the cropped N frames of images, and the cropped N frames of images include the eye regions; The second algorithm is used to identify the cropped N frames of images to obtain the eye opening and closing results of the N frames of images. The eye opening and closing results are used to represent the number of open eyes and / or the number of closed eyes in the image, or are used to represent the probability of having open eyes in the image; According to the eye opening and closing results of the N frames of images, the N frames of images are sorted to obtain the second sorting result.
[0106] The cropped N frames of images include the eye regions. The cropped N frames of images can be circular regions, elliptical regions, square regions, rectangular regions, etc. that include the eye regions. The embodiments of the present application do not make limitations in this regard. The resolution of the cropped N frames of images is less than the resolution of the N frames of images, and it only needs to include the eye regions.
[0107] In some examples, the resolution of the eye region can be 30*30, and the resolution of the cropped N frames of images only needs to be greater than 30*30.
[0108] If there are many eyes included in the image, the eye opening and closing results can be used to represent the number of open eyes and / or the number of closed eyes in the image. Exemplarily, if there are 5 people in the image, the eye opening and closing results can include 3 open eyes and 2 closed eyes.
[0109] If there are few eyes included in the image, the eye opening and closing results can be used to represent the probability of having open eyes in the image. Exemplarily, if there is 1 person in the image, the eye opening and closing results can include an open eye probability of 80%.
[0110] In the second sorting result, the images sorted in the front can be superior to the images sorted in the back, or the images sorted in the back can be superior to the images sorted in the front, or the images sorted in the middle can be superior to the images sorted in the front and superior to the images sorted in the back. The embodiments of the present application do not make limitations in this regard.
[0111] In this way, since the resolution of the cropped N frames of images is less than the resolution of the N frames of images, performing eye opening and closing detection after cropping the N frames of images is beneficial to saving power consumption.
[0112] Optionally, the area included in the cropped N-frame images is the face area. The face area may include the eye area. Therefore, the area included in the cropped N-frame images is the face area, which is beneficial to subsequent implementation of eye opening and closing detection.
[0113] Cropping the N-frame images to obtain that the area included in the cropped N-frame images is the face area can have multiple implementations.
[0114] In one possible implementation, face recognition is performed on the N-frame images to obtain the position information of the face area, and the N-frame images are cropped based on the position information of the face to obtain the cropped N-frame images. In this way, the implementation is simple.
[0115] In another possible implementation, face recognition is performed on other images related to the N-frame images to obtain the position information of the face area, where the resolution of the other images is less than the resolution of the N-frame images. The position information of the face area is converted to obtain the position information of the face area in the N-frame images, and the N-frame images are cropped based on the position information of the face in the N-frame images to obtain the cropped N-frame images.
[0116] In this way, since the resolution of the other images is less than the resolution of the N-frame images, performing face recognition based on the resolution of the other images is beneficial to reducing power consumption.
[0117] Optionally, before obtaining the N-frame images in response to the photographing operation, the method further includes: in response to the operation of turning on the camera, obtaining a first image collected by the camera; reducing the resolution of the first image to obtain a second image, where the resolution of the second image is less than the resolution of the first image; performing face recognition on the second image to obtain the position information of the face area in the second image; storing the first image in a first queue and storing the position information of the face area in the second image in a first buffer; in response to the photographing operation, obtaining the N-frame images, including: in response to the photographing operation, obtaining the N-frame images from the first queue; in response to the photographing operation, the method further includes: in response to the photographing operation, obtaining the position information of the face area in the target image from the first buffer, where the time stamp of the target image is the same as the time stamp of the N-frame images; converting the position information of the face area in the target image to obtain the position information of the face area in the N-frame images; cropping the N-frame images to obtain the cropped N-frame images, including: cropping the N-frame images based on the position information of the face area in the N-frame images to obtain the cropped N-frame images.
[0118] The first image is used to represent the image captured by the camera in real time. The electronic device can store the first image in the first queue in real time. The number of images that the first queue can store can be infinite or finite, and the embodiments of the present application do not limit this. If the number of images that can be stored in the first queue is finite, the images in the first queue can be updated over time. In this way, it is beneficial to save memory and at the same time can store the images required for use. The first queue can refer to the above-mentioned ZSL queue.
[0119] The resolution of the second image is less than the resolution of the first image. If the first image is the original image, the second image can be the preview image or the above-mentioned tiny thumbnail.
[0120] In response to the operation of turning on the camera, storing the first image in the first queue, and shrinking the first image to obtain the second image, and performing face recognition on the second image, etc., can be executed in parallel and are executed in parallel by different threads.
[0121] For example, the processing of the first image can refer to the above-mentioned Figure 4 shown offline photo-taking process, and the processing of the second image can refer to the above-mentioned online preview process.
[0122] It can be understood that when the camera is in the on state, the first image and the second image can be continuously generated, and the timestamps of the first image and the second image are the same.
[0123] If the timestamp of the target image is the same as the timestamp of the N-frame image, it can be explained that other than the resolution, the target image and the N-frame image are the same. Therefore, the position information of the face area in the target image can be converted to obtain the position information of the face area in the N-frame image, so as to crop the N-frame image based on the position information of the face area in the N-frame image to obtain the cropped N-frame image.
[0124] This method can refer to the above S401 to S406.
[0125] In this way, since the resolution of the target image is less than the resolution of the N-frame image, performing face recognition on the target image is beneficial to reducing power consumption.
[0126] Optionally, the first image is the original image, the second image is the preview image, or the resolution of the second image is less than the resolution of the preview image.
[0127] If the resolution of the second image is less than the resolution of the preview image. In one example, the second image can be the above-mentioned Figure 4 shown tiny thumbnail. In this way, performing face recognition on the second image is beneficial to reducing power consumption.
[0128] Optionally, a second algorithm is used to identify the N cropped images to obtain the open / closed eye results of the N images, including: denoising the N cropped images to obtain denoised images; using the second algorithm to identify the denoised images to obtain the open / closed eye results of the N images.
[0129] In this way, performing open / closed eye detection on the denoised images is beneficial to improving the accuracy of open / closed eye detection.
[0130] Optionally, the weight of the first sorting result is the first weight, and the weight of the second sorting result is the second weight, and the second weight is greater than the first weight. That the second weight is greater than the first weight can indicate that the importance degree of the open / closed eye detection result is higher. In this way, it is beneficial to reduce the probability that the fused image shows closed eyes.
[0131] Exemplarily, in a portrait photography scenario, the probability of eyes existing in the image is relatively high. That the second weight is greater than the first weight is beneficial to reducing the probability that the fused image shows closed eyes.
[0132] Optionally, the weight of the second sorting result is positively correlated with the number of face regions included in the N images.
[0133] The weight of the second sorting result is variable. The more the number of face regions included in each of the N images, the greater the weight of the second sorting result can be. It can be understood that the number of face regions included in each of the N images is the same.
[0134] Exemplarily, if there are 5 face regions included in each of the N images, the weight of the second sorting result can be 90%. If there are 3 face regions included in each of the N images, the weight of the second sorting result can be 80%.
[0135] In this way, the weight of the second sorting result can change according to the change in the number of face regions in the N images, with stronger flexibility. At the same time, it is more beneficial to reduce the probability that the fused image shows closed eyes.
[0136] It should be noted that the module names involved in the embodiments of the present application can all be defined as other names, as long as the functions of each module can be realized, and no specific restrictions are imposed on the module names.
[0137] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present application are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for the user to choose to authorize or refuse.
[0138] The image processing method of the embodiments of the present application has been described above. Next, the apparatus for executing the above method provided by the embodiments of the present application will be described. Those skilled in the art can understand that the method and the apparatus can be combined and cited with each other. The related apparatus provided by the embodiments of the present application can execute the steps in the above image processing method.
[0139] Figure 6 It is a schematic structural diagram of a chip provided by an embodiment of the present application. As Figure 6 shown, the chip 60 includes one or more (including two) processors 601, a communication line 602, a communication interface 603, and a memory 604.
[0140] In some embodiments, the memory 604 stores the following elements: executable modules or data structures, or subsets thereof, or extended sets thereof.
[0141] The above-described image processing method described in the embodiments of the present application can be applied to the processor 601 or implemented by the processor 601. The processor 601 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above image processing method can be completed by the integrated logic circuit in the hardware of the processor 601 or by instructions in software form. The above-mentioned processor 601 may be a general-purpose processor (for example, a microprocessor or a conventional processor), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates, transistor logic devices, or discrete hardware components. The processor 601 can implement or execute various processing-related methods, steps, and logic block diagrams disclosed in the embodiments of the present application.
[0142] The steps of the image processing method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. Among them, the software module may be located in a mature storage medium in the art such as a random access memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable read-only memory (EEPROM). This storage medium is located in the memory 604, and the processor 601 reads the information in the memory 604 and combines its hardware to complete the steps of the above method.
[0143] Communication can be carried out between the processor 601, the memory 604, and the communication interface 603 through the communication line 602.
[0144] In the above embodiment, the instructions stored in the memory for the processor to execute can be implemented in the form of a computer program product. Among them, the computer program product can be pre-written in the memory or downloaded and installed in the memory in the form of software.
[0145] The image processing method provided by the embodiments of this application can be applied to an electronic device with communication functions. The electronic device includes a terminal device. For the specific device form of the terminal device, reference can be made to the above relevant description, which will not be elaborated here.
[0146] The embodiments of this application provide a terminal device, which includes: a processor and a memory; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, so that the terminal device executes the above method.
[0147] The embodiments of this application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the above method is implemented. The methods described in the above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. If implemented in software, the functions can be stored as one or more instructions or codes on a computer-readable medium or transmitted on a computer-readable medium. The computer-readable medium can include a computer storage medium and a communication medium, and can also include any medium that can transfer a computer program from one place to another. The storage medium can be any target medium accessible by a computer.
[0148] In one possible implementation, the computer-readable medium may include RAM, ROM, compact disc read-only memory (CD-ROM), or other optical disc storage, magnetic disk storage, or any other medium targeted to carry or store the required program code in the form of instructions or data structures and accessible by a computer. Moreover, any connection is properly termed a computer-readable medium. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. As used herein, disk and optical disc include optical disc, laser disc, optical disc, Digital Versatile Disc (DVD), floppy disk, and Blu-ray disc, where disks typically reproduce data magnetically, while optical discs utilize lasers to optically reproduce data. Combinations of the above should also be included within the scope of computer-readable media.
[0149] An embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is run, it causes the computer to execute the above method.
[0150] Embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processing unit of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable devices to generate a machine, such that the instructions executed by the processing unit of the computer or other programmable data processing device generate means for implementing the functions specified in one Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0151] The above specific embodiments further elaborate on the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present invention should be included within the protection scope of the present invention.
Claims
1. An image processing method, characterized in that, Including: In response to a photographing operation, obtain N frames of images, where N is an integer greater than 1; Obtain an image sequence after sorting the N frames of images; wherein, the N frames of images in the image sequence are sorted according to the fusion weights when the N frames of images are subjected to image fusion, and the sorting of the N frames of images in the image sequence is related to a first sorting result and a second sorting result, the first sorting result is obtained by sorting the N frames of images using a first algorithm for jitter processing, and the second sorting result is obtained by sorting the N frames of images using a second algorithm for eye opening and closing detection; Perform image fusion based on the image sequence.
2. The method according to claim 1, wherein The second sorting result is obtained through the following method: Crop the N frames of images to obtain the cropped N frames of images, and the cropped N frames of images include eye regions; Use the second algorithm to identify the cropped N frames of images to obtain the eye opening and closing results of the N frames of images, and the eye opening and closing results are used to represent the number of open eyes and / or the number of closed eyes in the image, or are used to represent the probability of having open eyes in the image; Sort the N frames of images according to the eye opening and closing results of the N frames of images to obtain the second sorting result.
3. The method according to claim 2, wherein The region included in the cropped N frames of images is the face region.
4. The method according to claim 3, wherein Before the step of, in response to a photographing operation, obtaining N frames of images, the method further includes: In response to an operation of turning on the camera, obtain a first image collected by the camera; Reduce the size of the first image to obtain a second image, and the resolution of the second image is less than the resolution of the first image; Perform face recognition on the second image to obtain the position information of the face region in the second image; Store the first image in a first queue, and store the position information of the face region in the second image in a first buffer; The step of, in response to a photographing operation, obtaining N frames of images includes: In response to the photographing operation, obtain the N frames of images from the first queue; In response to the photographing operation, the method further includes: In response to the photographing operation, obtain the position information of the face region in the target image from the first buffer, and the time stamp of the target image is the same as the time stamp of the N frames of images; Convert the position information of the face region in the target image to obtain the position information of the face region in the N frames of images; The step of cropping the N frames of images to obtain the cropped N frames of images includes: Crop the N frames of images based on the position information of the face region in the N frames of images to obtain the cropped N frames of images.
5. The method according to claim 4, wherein The first image is the original image, the second image is the preview image, or the resolution of the second image is less than the resolution of the preview image.
6. The method according to any one of claims 2 to 5, characterized in that, The step of using the second algorithm to identify the cropped N frames of images to obtain the eye opening and closing results of the N frames of images includes: Denoise the cropped N frames of images to obtain a denoised image; Use the second algorithm to identify the denoised image to obtain the eye opening and closing results of the N frames of images.
7. The method according to any one of claims 1 to 6, characterized in that, The weight of the first sorting result is the first weight, and the weight of the second sorting result is the second weight, and the second weight is greater than the first weight.
8. The method according to any one of claims 1 to 7, characterized in that The weight of the second sorting result is positively correlated with the number of face regions included in the N frames of images.
9. An electronic device, characterized in that, Comprising: A processor and a memory; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the electronic device executes the method according to any one of claims 1-8.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1-8.
11. A chip system, characterized in that, Comprising at least one processor and a communication interface, the communication interface and the at least one processor are interconnected by a line, and the at least one processor is used to run a computer program or instructions to execute the method according to any one of claims 1-8.
12. A computer program product, characterized in that, Comprising a computer program, when the computer program is run, it causes a computer to execute the method according to any one of claims 1-8.
Citation Information
Patent Citations
Method for determining quality of iris image based on machine learning
CN102567744A
Image processing method and apparatus, mobile terminal and computer readable storage medium
CN107734253A
Image de-noising method and device, electronic device and computer readable storage medium
CN109348088A
An anti-eye-closing photographing method
CN109740472A
Face image storage method, device and equipment, computer medium and program product
CN114944004A