An image processing method, an electronic device, and a storage medium

By directly using the object detection results of the detection stream image after the photo-taking function to process the display effect, the problem of long user waiting time is solved, and the photo-taking speed and user experience are improved.

CN120282012BActive Publication Date: 2026-03-20HONOR DEVICE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

When users take photos with electronic devices, the time it takes from triggering the camera function to displaying the thumbnail is relatively long, which affects the user experience.

Method used

After the photo-taking function is triggered, the display effect is processed by directly using the detection results of specific objects in the detection stream image with lower resolution, instead of detecting the higher resolution photo-taking stream image and using the detection results of specific objects in the image detection path for fast display.

Benefits of technology

It reduces the time from taking a photo to displaying the thumbnail, improving the user's photo-taking speed and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120282012B_ABST
    Figure CN120282012B_ABST
Patent Text Reader

Abstract

The application provides an image processing method, an electronic device and a storage medium, and relates to the technical field of image processing. In the method, in response to a camera application being invoked, an image detection channel of the camera application is first started, which is used for detecting a specific object in a detection stream image to obtain a specific object detection result; in response to a photographing function of the camera application being triggered, a photographing channel of the camera application is then started, which is used for processing a photographing stream image to obtain a photographing picture; the resolution of the detection stream image is lower than that of the photographing stream image; the photographing channel can use the specific object detection result to perform display effect processing on the specific object in the photographing stream image to generate the photographing picture. In this way, the image detection channel can quickly obtain the specific object detection result by detecting the detection stream image with low resolution, so that the photographing channel can be reused, the speed of obtaining the photographing picture can be improved, and the photographing speed experienced by a user can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, and particularly relates to an image processing method, an electronic device and a storage medium. BACKGROUND

[0002] With the wide popularity of electronic devices such as mobile phones and tablet computers and the rapid development of photographing technology, more and more users choose to use electronic devices to record wonderful moments in life, so people widely use the photographing function of electronic devices in daily life.

[0003] Taking a mobile phone as an example, the mobile phone has a photo instant display function. A user triggers the photographing function of the mobile phone to take a photo, and the mobile phone instantaneously displays the photo in the form of a thumbnail in a specific area (for example, the lower left corner) of the mobile phone screen. The user can confirm that the mobile phone has completed photographing by seeing the thumbnail. Then the user triggers the thumbnail, and the mobile phone can display the photo taken by the mobile phone for browsing. However, the time consumed from the user triggering the photographing function of the electronic device to the electronic device displaying the thumbnail is relatively long, and the user feels that the photographing completion time is relatively long, which easily affects the user experience. SUMMARY

[0004] In order to solve the above problems, the present application provides an image processing method, an electronic device and a storage medium. The purpose is to reduce the time consumed from the user triggering the photographing function of the electronic device to the electronic device displaying the thumbnail, improve the photographing speed, and further improve the user experience.

[0005] In a first aspect, the present application provides an image processing method. The method can be applied to an electronic device, such as a mobile phone, a tablet computer, a notebook computer, or other devices including a camera application. In the method, the camera application is invoked, for example, a user clicks an icon of the camera application to start the camera application. In response to the invocation, the electronic device can start an image detection pass of the camera application. The image detection pass can detect a specific object in a detection stream image to obtain a specific object detection result. The specific object can be a face, a human body, or other objects in the image. Then, when a photographing function of the camera application is triggered, for example, a user clicks a photographing control of the camera application to trigger the photographing function, or the user issues a gesture instruction or a voice instruction to trigger the photographing function, the electronic device can start a photographing pass of the camera application in response to the triggering of the photographing function. The photographing pass is used to process a photographing stream image to obtain a photographing picture. The detection stream image and the photographing stream image are both generated based on a raw image obtained by the camera application. For example, the raw image is captured by a camera of the electronic device based on hardware associated with the photographing function of the camera application. The resolution of the detection stream image is lower than that of the photographing stream image. Subsequently, the photographing pass can use the specific object detection result obtained by the image detection pass to perform display effect processing on the specific object in the photographing stream image to generate the photographing picture. The specific object detection result can be sent to the photographing pass after being obtained by the image detection pass, or the photographing pass can obtain the specific object detection result from the image detection pass.

[0006] In this way, after the photographing function of the camera application is triggered, the photographing pass does not need to detect the specific object based on the photographing stream image with a higher resolution. Instead, the photographing pass can directly use the specific object detection result obtained by the image detection pass. Because the resolution of the detection stream image is lower than that of the photographing stream image, the image detection pass can detect faster, and the photographing pass can obtain the specific object detection result faster, generate the photographing picture earlier, and thus improve the user's perception of the photographing speed.

[0007] In a possible implementation, after the electronic device obtains the raw image, the image detection pass can reduce the resolution of the raw image to obtain a detection stream image with a lower resolution. In this way, the detection speed of the image detection pass for detecting the specific object in the detection stream image is improved, and the specific object detection result is obtained faster.

[0008] In a possible implementation, the first vertex of the detection stream image is the origin of the first coordinate system, and the second vertex of the photograph stream image is the origin of the second coordinate system; the position of the first vertex in the detection stream image is equivalent to the position of the second vertex in the photograph stream image, for example, the top-left vertex of the detection stream image is the origin of the first coordinate system, and the top-left vertex of the photograph stream image is the origin of the second coordinate system. The specific object detection result can include first coordinate data of the specific object in the detection stream image in the first coordinate system, for example, the first coordinate data can be coordinate data of a face or a face feature point in the first coordinate system. The photograph channel can convert the first coordinate data included in the specific object detection result into second coordinate data in the second coordinate system, and then perform display effect processing on the specific object in the photograph stream image by using the converted specific object detection result to generate a photograph.

[0009] In this way, considering that the positions of the specific object in the detection stream image and the photograph stream image are deviated due to different coordinate systems, the converted second coordinate data obtained by the photograph channel is corresponding to the position of the specific object in the photograph stream image, and therefore the accuracy of the display effect processing can be improved, and a photograph with higher quality can be obtained.

[0010] In a possible implementation, the image processing method can further include: based on the second coordinate data, the photograph channel can filter out data that does not belong to the photograph stream image from the converted specific object detection result; and then, the photograph channel can perform display effect processing on the specific object in the photograph stream image by using the converted and filtered specific object detection result to generate a photograph. In this way, considering that the number of specific objects included in the detection stream image and the photograph stream image can be deviated due to different field angles, the specific object detection result is obtained based on detection of the detection stream image, and when the field angle of the detection stream image is greater than that of the photograph stream image, there can be multiple specific objects in the detection stream image, but one or some of the multiple specific objects can not be completely in the photograph stream image, and therefore the specific object detection result needs to be filtered out. The converted and filtered specific object detection result only includes detection results of specific objects that exist in the photograph stream image, and therefore the time for display effect processing can be avoided, the photograph can be generated more quickly, and the user's experience of the photographing speed can be further improved.

[0011] In a possible implementation, the detection flow image can include a plurality of specific objects, and the converted specific object detection result can include a plurality of groups of detection results corresponding to the plurality of specific objects one by one. The photographing passageway can first filter out the first coordinate points that are not in the photographing flow image from the second coordinate data. For example, when the upper left corner of the photographing flow image is the origin of the second coordinate system, and the coordinate values of the coordinate points on the photographing flow image are all non-negative, the photographing passageway can filter out the coordinate points with negative coordinate values from the coordinate points included in the second coordinate data as the first coordinate points. Then, the photographing passageway can determine the target detection result to which the first coordinate points belong from the plurality of groups of detection results. Subsequently, the photographing passageway can filter out the target detection result from the converted specific object detection result, that is, from the plurality of groups of detection results. In this way, the specific objects included in the converted and filtered specific object detection result are all on the photographing flow image, which can avoid wasting time on detecting specific objects that are not complete or not on the photographing flow image, and further improve the user's experience of the photographing speed.

[0012] In a possible implementation, the photographing passageway can first determine the field angles of the detection flow image and the photographing flow image, and when the two field angles are determined to be the same, the photographing passageway can calculate the square root of the ratio of the resolution of the photographing flow image to the resolution of the detection flow image as a first multiple, and then multiply the first multiple by the coordinate values of the first coordinate data to obtain the second coordinate data. In this way, when the field angles of the detection flow image and the photographing flow image are the same, it indicates that the contents included in the detection flow image and the photographing flow image are the same and there is no deviation, and therefore the coordinate conversion can be performed only based on the difference between the resolutions of the two, which can ensure the accuracy of the converted second coordinate data.

[0013] In a possible implementation, the third vertex of the original image is the origin of the third coordinate system, and the position of the third vertex on the original image is the same as the position of the second vertex on the photographing flow image. For example, the left upper corner vertex of the original image is the origin of the third coordinate system, and the left upper corner vertex of the photographing flow image is the origin of the second coordinate system. The photographing passageway can first determine the field angles of the detection flow image and the photographing flow image, and when the two field angles are determined to be different, the photographing passageway can first convert the first coordinate data into third coordinate data in the third coordinate system, and then convert the third coordinate data into the second coordinate data. In this way, when the field angles of the detection flow image and the photographing flow image are different, it indicates that the contents included in the detection flow image and the photographing flow image are different, and the first coordinate data cannot be directly converted into the second coordinate data. However, the detection flow image and the photographing flow image are both obtained based on the original image, and therefore the conversion can be performed by using the third coordinate system in which the original image is located, so as to ensure the accuracy of the converted second coordinate data.

[0014] In a possible implementation, the photographing passage can first acquire first vertex coordinate data and a second multiple, where the first vertex coordinate data is data of a vertex of the detection stream image in the third coordinate system; the second multiple is a square root of a ratio of a resolution of the original image to a resolution of the detection stream image; and the photographing passage can convert the first coordinate data into third coordinate data based on the first vertex coordinate data and the second multiple. In this way, the conversion accuracy of the coordinate data is ensured by taking into account the difference between the resolutions of the original image and the detection stream image and the difference between the field angles.

[0015] In a possible implementation, the first vertex coordinate data includes second vertex coordinate data of a first vertex of the detection stream image in the third coordinate system, which can include, for example, coordinate values of a top-left vertex of the detection stream image in the third coordinate system; the photographing passage can first multiply the coordinate values of the first coordinate data by the second multiple to obtain fourth coordinate data, that is, the resolution conversion is first performed on the first coordinate data to obtain the fourth coordinate data, and then add the coordinate values of the fourth coordinate data to the coordinate values of the second vertex coordinate data to obtain the third coordinate data, that is, the field angle conversion is further performed on the fourth coordinate data to obtain the third coordinate data. In this way, when the detection stream image is obtained by first performing the field angle conversion and then performing the resolution conversion on the original image, the photographing passage can first convert the coordinate points based on the difference between the resolutions and then convert the coordinate points based on the difference between the field angles, so as to restore the specific object in the detection stream image to the original image, and ensure the conversion accuracy of the coordinate data.

[0016] In a possible implementation, the first vertex coordinate data includes third vertex coordinate data of a first vertex of the detection stream image in the third coordinate system, which can include, for example, coordinate values of a top-left vertex of the detection stream image in the third coordinate system; the photographing passage can first add the coordinate values of the first coordinate data to the coordinate values of the third vertex coordinate data to obtain fifth coordinate data, that is, the field angle conversion is first performed on the first coordinate data to obtain the fifth coordinate data, and then multiply the coordinate values of the fifth coordinate data by the second multiple to obtain the third coordinate data, that is, the resolution conversion is further performed on the fifth coordinate data to obtain the third coordinate data. In this way, when the detection stream image is obtained by first performing the resolution conversion and then performing the field angle conversion on the original image, the photographing passage can first convert the coordinate points based on the difference between the field angles and then convert the coordinate points based on the difference between the resolutions, so as to restore the specific object in the detection stream image to the original image, and ensure the conversion accuracy of the coordinate data.

[0017] In a possible implementation, when converting the third coordinate data into the second coordinate data in the photographing pass, fourth vertex coordinate data and a third multiple can be acquired first, where the fourth vertex coordinate data is data of a vertex of the photographing stream image in the third coordinate system, for example, coordinate points of each vertex of the photographing stream image in the third coordinate system, and the third multiple is a square root of a ratio of a resolution of the photographing stream image to a resolution of the original image, that is, a multiple of magnification or reduction when converting the original image into the photographing stream image; then the photographing pass can convert the third coordinate data into the second coordinate data based on the fourth vertex coordinate data and the third multiple. In this way, the conversion accuracy of the coordinate data is ensured by taking into account the difference in resolution and the difference in field of view between the original image and the photographing stream image.

[0018] In a possible implementation, the fourth vertex coordinate data includes fifth vertex coordinate data of a second vertex of the photographing stream image in the third coordinate system, for example, coordinate values of the top-left corner vertex of the photographing stream image in the third coordinate system. The photographing pass can multiply the coordinate values of the third coordinate data by the third multiple to obtain sixth coordinate data, that is, the resolution conversion is performed on the third coordinate data to obtain the sixth coordinate data; then the photographing pass can subtract the coordinate values of the fifth vertex coordinate data from the coordinate values of the sixth coordinate data to obtain the second coordinate data, that is, the field of view conversion is performed on the sixth coordinate data to obtain the second coordinate data. In this way, the photographing stream image can be obtained by performing the resolution conversion on the original image first and then performing the field of view conversion, and the conversion accuracy of the coordinate data can be ensured.

[0019] In a possible implementation, the fourth vertex coordinate data includes sixth vertex coordinate data of a second vertex of the photographing stream image in the third coordinate system, for example, coordinate values of the top-left corner vertex of the photographing stream image in the third coordinate system. The photographing pass can subtract the coordinate values of the sixth vertex coordinate data from the coordinate values of the third coordinate data to obtain seventh coordinate data, that is, the field of view conversion is performed on the third coordinate data to obtain the seventh coordinate data; then the photographing pass can multiply the coordinate values of the seventh coordinate data by the third multiple to obtain the second coordinate data, that is, the resolution conversion is performed on the seventh coordinate data to obtain the second coordinate data. In this way, the photographing stream image can be obtained by performing the field of view conversion on the original image first and then performing the resolution conversion, and the conversion accuracy of the coordinate data can be ensured.

[0020] In a possible implementation, the photographing stream image can be obtained by reducing the field of view of the original image in the photographing pass. For example, when the user adjusts the focal length of the camera application from 1.0x to 2.0x, it indicates that the user needs to reduce the content contained in the photographed picture, and therefore the field of view needs to be reduced. In this way, the photographed picture that meets the user's demand can be obtained.

[0021] In a possible implementation, when the photographing pass reduces the field of view of the original image to obtain the photographing stream image, the photographing pass can first acquire the zoom ratio of the camera application, then determine the second coordinate point of the original image based on the zoom ratio, that is, the coordinate point required when the original image is cropped, and then crop the original image based on the second coordinate point to obtain the photographing stream image. In this way, the photographing picture that meets the user's demand can be obtained.

[0022] In a possible implementation, the detection stream image and the photographing stream image can both be generated based on a first original image obtained by the camera application, that is, both are generated based on the same original image. For example, the first original image can be captured by the camera before the photographing function of the camera application is triggered, or can be captured by the camera after the photographing function of the camera application is triggered.

[0023] In a possible implementation, the photographing stream image can be generated based on a first original image obtained by the camera application, and the detection stream image can be generated based on a second original image obtained by the camera application, that is, the two are generated based on different original images, and the position of the specific object in the first original image is the same as the position of the specific object in the second original image.

[0024] In a possible implementation, the first original image can be obtained by performing image fusion processing on a plurality of original images including the second original image. For example, when the camera application is in an HDR shooting mode, the photographing picture displayed to the user can be obtained based on a plurality of original images.

[0025] In a possible implementation, the specific object includes a face, and the specific object detection result is a face detection result. For example, the face detection result can include coordinate data of a face frame, coordinate data of feature points in the face, face attributes, face angles, and the like. The photographing pass can use the face detection result to perform retouching processing on the face in the photographing stream image, for example, large-eye, skin beautifying, and the like, to generate a photographing picture that can be displayed to the user.

[0026] In a second aspect, the present application provides an electronic device, which includes a memory and a processor; the memory stores computer program code, and the computer program code includes computer instructions; the one or more processors invoke the computer instructions to enable the electronic device to execute the image processing method of the first aspect.

[0027] In a third aspect, the present application provides a computer readable storage medium, which stores a computer program. When the computer program is executed by a processor, the image processing method of the first aspect is implemented.

[0028] From the above technical solution, the present application has the following beneficial effects:

[0029] After the camera application is invoked, the electronic device can start the image detection channel of the camera application, which is used to detect specific objects in the detection stream image to obtain specific object detection results. After the camera application is triggered, the electronic device can start the camera application photographing channel, which is used to process the photographing stream image to obtain a photographing picture. The detection stream image and the photographing stream image are both generated based on the original image obtained by the camera application, and the resolution of the detection stream image is lower than that of the photographing stream image. Therefore, compared with the photographing channel, the image detection channel obtains specific object detection results faster. The photographing channel can directly reuse the specific object detection results to display the specific objects in the photographing stream image and generate the photographing picture faster. In this way, the image processing method provided by the present application can improve the speed of obtaining the photographing picture, and then present the photographing picture to the user faster, reduce the time the user feels the completion of the photographing, and improve the user experience. BRIEF DESCRIPTION OF DRAWINGS

[0030] Figure 1 A schematic diagram of a photographing process in a related technology is provided for an embodiment of the present application.

[0031] Figure 2 A schematic diagram of a camera running process in a related technology is provided for an embodiment of the present application.

[0032] Figure 3 A schematic diagram of a face detection module is provided for an embodiment of the present application.

[0033] Figure 4 A signaling interaction diagram of an image processing method is provided for an embodiment of the present application.

[0034] Figure 5 A schematic diagram of reusing first face data is provided for an embodiment of the present application.

[0035] Figure 6a A schematic diagram of adjusting the focal length is provided for an embodiment of the present application.

[0036] Figure 6b A schematic diagram of adjusting the photo ratio is provided for an embodiment of the present application.

[0037] Figure 7 A schematic diagram of mapping first face data is provided for an embodiment of the present application.

[0038] Figure 8 A schematic diagram of mapping second face data is provided for an embodiment of the present application.

[0039] Figure 9A schematic diagram of invalid face data provided for an embodiment of the present application.

[0040] Figure 10 A structural schematic diagram of an electronic device provided for an embodiment of the present application. DETAILED DESCRIPTION

[0041] For the sake of clear and concise description of the following embodiments, first, the glossary involved in the embodiments of the present application is explained. It should be understood that the explanation is for clearer understanding of the embodiments of the present application, and does not necessarily constitute a limitation on the embodiments of the present application.

[0042] Image sensor: used for sensing light, when light shines on the image sensor, the image sensor converts the optical signal into an electrical signal. In the embodiments of the present application, the camera of the electronic device includes an image sensor.

[0043] Raw image: also known as RAW image, is an image obtained by converting the electrical signal obtained by the image sensor through the image signal processor. In some embodiments, the camera of the electronic device includes an image signal processor.

[0044] Field of view (FOV): used to indicate the maximum angle range that can be shot by the camera. The object to be photographed is within this angle range, and the object to be photographed will be captured by the camera. The object to be photographed is outside this angle range, and the object to be photographed will not be captured by the camera. The larger the field of view of the camera, the larger the shooting range; and the smaller the field of view of the camera, the smaller the shooting range.

[0045] Focal length: refers to the focal length range of the camera, also known as zoom ratio. In some embodiments, the focal length of the electronic device includes 0.5x-30x, x represents the fixed focal length of the electronic device.

[0046] Resolution: refers to the number of pixels in the vertical and horizontal directions of the image, which determines the degree of detail of the image and is used to represent the sharpness of the image. The higher the resolution, the more pixels it contains, and the clearer the image. In some embodiments, the resolution is 400x300, which means that the image includes 400 pixels in the vertical direction and 300 pixels in the horizontal direction.

[0047] Detection stream image: a low-resolution image obtained by image processing the raw image obtained by the camera based on the first image parameter of the detection stream image. In some embodiments, the electronic device can perform perception-based detection such as subject detection, face detection, and body detection on the detection stream image to identify the object to be photographed, and then determine the shooting scene.

[0048] The subject detection refers to detecting an object included in the detection stream image, and the object includes, for example, a human, an animal, a plant, and the like. The face detection refers to detecting a human face included in the detection stream image, including detecting a position of the human face in the image, positions of feature points such as eyes, a nose, and a mouth of the human face in the image, a human face attribute, a human face angle, and the like. The human body detection refers to detecting a human body included in the detection stream image, including detecting a position of a limb and a trunk in the image, and the like.

[0049] The photograph stream image is an image obtained by performing image processing on an original image obtained by a camera based on a second image parameter of the photograph stream image. In some embodiments, the electronic device can optimize the photograph stream image to obtain a higher-quality image to display to the user.

[0050] The technical advantages of the image processing method, the electronic device, and the storage medium provided in the present application are described below in comparison with related technologies. For the convenience of understanding, an example scenario is described. In the example scenario, the electronic device is a mobile phone 100.

[0051] In related technologies, the time consumed from triggering the photograph function by the user to obtaining a photograph that can be presented to the user by the mobile phone 100 can be referred to as shot2review time. Taking a human portrait photograph as an example, the user can use the photograph function (i.e., the photographing function) of the mobile phone 100 to take a human portrait photograph including a human face.

[0052] As shown in Figure 1 , the mobile phone 100 displays a camera preview interface 11 of a camera application, the camera preview interface includes a photograph control 12 and a thumbnail display area 13, and after the user holds the mobile phone 100 to aim at a human face to be photographed, the user can click the photograph control 12 to trigger the photograph function of the mobile phone 100. The mobile phone 100 displays the obtained photograph in the form of a thumbnail in the thumbnail display area 13.

[0053] As shown in Figure 1 , Figure 2 , and Figure 3 , during the running of the camera application, the image sensor continuously works to obtain multiple RAW images, and a detection channel (which can also be referred to as an image detection channel) and a preview channel also continuously work for each RAW image. After the user clicks the photograph control 12 (i.e., triggers the photograph function), the photograph channel starts to work, the mobile phone 100 first performs related processing on a RAW image to obtain a photograph stream image, then performs related processing such as denoising on the photograph stream image by using an algorithm a of the photograph channel, and then processes the human face by using a human face module.

[0054] The face module includes a face detection module and a face processing module. The mobile phone 100 first detects the face information of the face included in the photograph stream image by the face detection module, such as face position detection, face attribute detection, and face feature point detection. Then, the mobile phone 100 processes the face included in the photograph stream image by the face processing module based on the face information, such as adding a beautifying effect (big eyes, slim face, and skin smoothing) to the face based on the face information. Finally, the mobile phone 100 displays the photographed portrait photo in the form of a thumbnail in the thumbnail display area 13.

[0055] However, in the related art, the mobile phone 100 needs to complete the face detection when taking a photo, and may also need to complete other detections such as subject detection and human body detection. The detection may take 100-200 ms. Subsequently, the mobile phone 100 needs to process the face, human body, and subject based on the detected information to obtain the final photo, that is, the shot2review is long. Therefore, the time from the user clicking the photograph control 12 to the thumbnail display area 13 displaying the thumbnail is also long, which makes the user feel that the photographing speed is slow, which affects the user experience.

[0056] Therefore, to solve the above problems, the embodiment of the present application provides an image processing method, an electronic device, and a storage medium. After the photographing function of the camera application is triggered, the photographing path no longer performs face detection on the photograph stream image, but directly reuses the face information obtained by detecting the detection stream image by the detection path. The resolution of the detection stream image is lower than that of the photograph stream image, so the speed of face detection can be accelerated, the speed of obtaining face information by the photographing path can be improved, the face can be processed earlier, the speed of obtaining a photo (which can also be referred to as a photographing picture) can be improved, and the photographing speed perceived by the user can be improved, thereby improving the user experience.

[0057] Next, still taking the mobile phone as an example, the user uses the photographing function of the mobile phone to take a portrait photo including a face, and combining the above Figure 4 and Figure 5 The image processing method provided by the embodiment of the present application is described in detail.

[0058] Taking the operating system of the mobile phone as an example, the Android system can adopt a layered architecture, and each layer has a clear role and division of labor. The layers communicate with each other through a software interface. In some embodiments, the system is divided into five layers, from top to bottom, the application program layer, the application program framework layer, the hardware abstraction layer (HAL), the driver layer, and the hardware layer. It should be noted that the mobile phone can also run the iOS operating system, and the present application does not limit this.

[0059] As shown in Figure 4 application layer includes a camera application (hereinafter referred to as a camera), the application framework layer includes a camera access interface, the hardware abstraction layer includes a camera HAL module, the driver layer includes a camera driver, and the hardware layer includes a camera.

[0060] The image processing method includes the following steps:

[0061] S401: The camera application receives a start operation of the camera application by a user.

[0062] The application layer can include a series of application packages. In the embodiment of the present application, the application package can include the camera application.

[0063] The start operation refers to a triggering operation of the user to start the camera application of the mobile phone.

[0064] Exemplarily, the start operation includes a click operation of the user on the application icon of the camera application included in the main interface of the mobile phone. Exemplarily, the start operation includes a click operation of the user on the application icon of the camera application of the lock screen interface of the mobile phone, etc. Exemplarily, the start operation includes an operation of the user to issue a sound of “opening the camera”, “selfie”, etc. to trigger the intelligent voice service of the mobile phone. The present application does not make any limitation in this regard.

[0065] S402: The camera application sends a start request of the camera application to the camera access interface.

[0066] The application framework layer provides application programming interfaces and programming frameworks for the application programs of the application layer. The application framework layer includes some pre-defined functions. In the embodiment of the present application, the application framework layer can include the camera access interface for providing application programming interfaces and programming frameworks for the camera application.

[0067] In response to the operation of the user to start the camera application, the camera application calls the camera access interface of the application framework layer to start the camera application.

[0068] S403: The camera access interface sends a call request to the camera HAL module.

[0069] The HAL layer is an interface layer between the application framework layer and the driver layer. In the embodiment of the present application, the hardware abstraction layer can include the camera HAL module for providing running codes and data for implementing the image processing method provided in the embodiment of the present application.

[0070] The camera access interface calls the camera HAL module to implement the image processing method provided in the embodiment of the present application.

[0071] S404: The camera HAL module sends a start request of the camera to the camera driver.

[0072] The driver layer is the layer between hardware and software. It includes drivers for various hardware components. In embodiments of this application, the driver layer may include a camera driver, which drives the image sensor of the camera to convert sensed light signals into electrical signals and drives an image signal processor to preprocess the electrical signals to obtain the original image. In some embodiments, the image signal processor may be located within the camera.

[0073] S405: Camera driver starts the camera.

[0074] After receiving the camera's startup request, the camera driver starts the image sensor and image signal processor.

[0075] S406: The camera captures a second raw image.

[0076] The camera's image sensor can detect light signals. The image sensor converts the detected light signals into electrical signals, and the image signal processor preprocesses the electrical signals to obtain a second original image.

[0077] In some embodiments, the image parameters of the second original image are obtained based on the parameters of the camera. For example, the camera parameters include field of view, pixels, image ratio, etc. For instance, if the camera has 12 million pixels, the second original image acquired by the camera includes 12 million pixels. If the camera's image ratio is 4:3, then the ratio of the number of pixels in the vertical direction to the number of pixels in the horizontal direction of the second original image is 4:3, indicating that the resolution of the second original image obtained by the camera is 4000×3000.

[0078] S407: The camera sends a second raw image to the camera HAL module via the camera driver.

[0079] After receiving the second raw image, the camera HAL module processes the second raw image based on the running code and data of the included image processing methods.

[0080] S408: The camera HAL module performs image processing on the second raw image based on the first image parameters to obtain the detection stream image.

[0081] like Figure 5 As shown, after the second raw image is acquired based on the image sensor, the detection path and the preview path begin to operate. It should be noted that the image signal processor is not shown in the figure, but the raw image is obtained based on the image sensor and the image signal processor.

[0082] In some embodiments, the camera HAL module includes a face detection module, a human body detection module, an object detection module, and the like for detecting the passageway. The camera HAL module includes an algorithm 1, an algorithm 2, and the like for the preview passageway. Exemplarily, the algorithm 1 and the algorithm 2 can be a denoising algorithm, a sharpening algorithm, and the like, which are not limited in the present application.

[0083] In some embodiments, the camera HAL module includes a first image parameter for detecting the stream image and an image parameter for previewing the stream image. The camera HAL module performs image processing on the second raw image based on the first image parameter to obtain the detecting stream image, and performs image processing on the second raw image based on the image parameter for previewing the stream image to obtain the preview stream image, which can be seen from Figure 5 .

[0084] It should be understood that the detecting stream image is mainly used for perception type detection to determine the object to be photographed and thus determine the shooting scene, which is not presented to the user, and thus the resolution of the detecting stream image is low, i.e., the resolution of the first image parameter is low. It should be noted that the first image parameter is pre-stored in the camera HAL module, which can be fixed or can be adaptively changed according to the image parameter of the second raw image, the external environment, and the like, which are not limited in the present application.

[0085] In some embodiments, the first image parameter includes a resolution of 40x30 (i.e., the number of pixels) and a field of view of 65°. The resolution of the second raw image is 4000x3000, and the field of view is 130°. Then the mobile phone can crop (i.e., field of view conversion) the second raw image based on the field of view of 65° to obtain an image with a number of pixels of 2000x1500, and then perform resolution conversion to obtain a detecting stream image with a number of pixels of 40x30.

[0086] Exemplarily, the field of view conversion can be to crop the periphery of the second raw image with the center point of the second raw image as the reference. Or it can be to crop the right and lower edges of the second raw image with the top left vertex of the second raw image as the reference, which are not limited in the present application.

[0087] Exemplarily, the resolution conversion can be to convert the size of the image at a constant ratio, such as to increase the resolution by enlarging the size of the image at a constant ratio, or to decrease the resolution by reducing the size of the image at a constant ratio. The present application is not limited in this regard.

[0088] Based on the above example, the center point of the second raw image can be taken as the center point of the cropped image, 1000 pixels are cropped from the top and bottom of the second raw image in the vertical direction, and 750 pixels are cropped from the left and right of the second raw image in the horizontal direction. Then the cropped image is reduced at a constant ratio of 50 times to obtain a detecting stream image with a field of view of 65° and a resolution of 40x30.

[0089] It should be understood that, in the case of the image scale being invariant, the resolution conversion is to convert the definition of the image, which changes the number of pixels included in the image, but does not change the content contained in the image. The field of view conversion changes the shooting range of the camera, which also changes the number of pixels included in the image, but increases or reduces the content contained in the image, and does not change the definition of the image. In some embodiments, adjusting the focal length and the image scale both change the field of view of the image. As shown in the following embodiments Figure 6a and Figure 6b It should be understood that, in the case of the image scale being invariant, the resolution conversion is to convert the definition of the image, which changes the number of pixels included in the image, but does not change the content contained in the image. The field of view conversion changes the shooting range of the camera, which also changes the number of pixels included in the image, but increases or reduces the content contained in the image, and does not change the definition of the image. In some embodiments, adjusting the focal length and the image scale both change the field of view of the image. As shown in the following embodiments

[0090] It should be understood that, in the case of the image scale being invariant, the resolution conversion is to convert the definition of the image, which changes the number of pixels included in the image, but does not change the content contained in the image. The field of view conversion changes the shooting range of the camera, which also changes the number of pixels included in the image, but increases or reduces the content contained in the image, and does not change the definition of the image. In some embodiments, adjusting the focal length and the image scale both change the field of view of the image. As shown in the following embodiments

[0091] In addition, it should be noted that the field of view included in the first image parameter can be the same as the field of view corresponding to the second original image, and only the resolution conversion is needed.

[0092] S409: The camera HAL module detects the detection stream image based on a face detection algorithm to obtain first face data.

[0093] As shown in Figure 5 The detection channel includes a face detection module, a human body detection module, and an object detection module, etc. After obtaining the detection stream image, the camera HAL module can detect the detection stream image based on each detection module included in the detection channel.

[0094] In some embodiments, each detection module includes a detection algorithm. For example, the camera HAL module can detect the detection stream image based on the face detection algorithm included in the face detection module to obtain the first face data (also referred to as a specific object detection result). For example, the face detection algorithm can include a face position detection algorithm, a face attribute detection algorithm, and a face feature point detection algorithm. Correspondingly, the first face data can include coordinate data corresponding to the face position, coordinate data corresponding to the face feature point, and face angle data, etc.

[0095] It should be understood that the resolution of the detection stream image is low. Therefore, when the face detection algorithm is used to detect the face, the processing speed is fast, and the first face data can be quickly obtained.

[0096] In some embodiments, the first face data can include one face-related face data, or can include multiple face-related face data. For example, the first face data includes face data 1, face data 2 and face data 3, which are one face-related face data respectively.

[0097] In some embodiments, after the camera is successfully started, the camera continues to work to obtain a plurality of second raw images, and the detection channel also continues to work. The camera HAL module processes each second raw image based on the first image parameter to obtain a plurality of detection stream images, and detects each detection stream image based on the face detection algorithm to obtain the first face data of each detection stream image. That is, during the running of the camera application, the camera and the detection channel continue to work, and after the camera obtains a second raw image, the S407-S409 are executed once.

[0098] S410: The camera application receives a trigger operation of a photographing function of the camera application.

[0099] After the user determines that the camera of the mobile phone is aligned with the object to be photographed through the preview image displayed on the screen of the mobile phone, the user can trigger the photographing function of the camera application.

[0100] In some embodiments, as shown in FIG. 12, the trigger operation of the photographing function can be a click operation of the user on the photographing control 12. Figure 1

[0101] In some embodiments, the trigger operation of the photographing function can include a setting operation of the user on the smiley face snapshot function of the camera application, and an operation of changing the face in the preview image into a smiley face, so as to trigger the smiley face snapshot function of the camera application. The present application does not limit this.

[0102] S411: The camera application sends a photographing request and second image parameters to the camera HAL module through the camera access interface.

[0103] The second image parameters are image parameters of the photographing stream image.

[0104] In some embodiments, the second image parameters can include image parameters such as focal length, resolution and photo size (also referred to as image size). For example, the second image parameters of the photographing stream image can include a focal length of 2.0x, a photo size of 4:3, a resolution of 3072x4096, and the like.

[0105] In some embodiments, the second image parameters can be obtained based on the default settings of the camera application.

[0106] In some embodiments, the second image parameters can also be obtained based on the configuration operation of the user, which is not limited by the present application.

[0107] ​As shown in Figure 6a and 6b , the mobile phone 100 displays a camera preview interface 610. The camera preview interface 610 includes a photographing control 611, a focus section selection control 612, and a setting control 613, etc. The photographing control 611 is used to implement the function of taking a photo, the focus section selection control 612 is used to implement the function of selecting a focus section, and the setting control 613 is used to implement the function of setting a camera parameter. The camera preview interface 610 further includes a preview photo display area 614 and a thumbnail display area 615. In some embodiments, before triggering the photographing control 612, i.e., before performing S410, the user can set an image parameter (which can also be referred to as a photo parameter), such as setting a focus section and a photo size, etc. The camera application receives the user's setting operation on the image parameter, and generates a second image parameter.

[0108] As shown in Figure 6a , the user swipes the focus section selection control 612 included in the camera preview interface 610 to the right. In response to the user's swiping operation to the right, the mobile phone 100 adjusts the focus section from 1.0x to 2.0x, the focus section becomes larger, the photographing range of the camera lens becomes smaller, and the content displayed by the camera preview interface 420 is reduced.

[0109] As shown in Figure 6b , the user clicks the setting control 613 included in the camera preview interface 610. In response to the user's clicking operation, the mobile phone 100 displays a camera parameter setting interface 620. The camera parameter setting interface 620 includes a photo size setting control 621 and other controls. The photo size setting control 621 is used to set the size of a photo. The parameter setting interface 620 further includes a return control 622, which is used to return to the camera preview interface 610.

[0110] The user triggers the photo size setting control 621 to select 4:3, and the mobile phone 100 can set it to 4:3. Subsequently, the user triggers the return control 622, and the mobile phone 100 returns to the camera preview interface 610. The photo size of the preview photo displayed in the preview photo display area 614 is adjusted from 1:1 to 4:3, and the content displayed by the preview photo also changes.

[0111] It should be noted that the above setting of the focus section and the photo size is only an example. The user can set any image parameter, or can not set an image parameter. In this case, the mobile phone 100 uses the default image parameter of the mobile phone 100 to take a photo, i.e., the camera application generates a second image parameter based on the default setting. This application does not limit this.

[0112] As shown in Figure 6a and Figure 6bAs shown, changes in the photo ratio and focal length of the second image parameters will affect the content displayed in the preview photo. Therefore, adjusting the photo ratio or focal length will cause a change in the field of view.

[0113] S412: The camera HAL module performs image processing on the first original image based on the second image parameters to obtain the captured image stream.

[0114] The camera HAL module adjusts the image parameters of the first original image based on the second image parameters to obtain the captured image stream.

[0115] like Figure 5 As shown, when a user triggers the camera application's photo-taking function, that is, after the camera HAL module receives the photo-taking request and the second image parameters, the photo-taking path begins to work. In some embodiments, the camera HAL module includes multiple algorithms such as algorithm a, algorithm b, and algorithm c for the photo-taking path.

[0116] Based on the above example, before the user triggers the camera application's photo-taking function, the camera HAL module can obtain multiple detection stream images and the first face data of each detection stream image.

[0117] In some embodiments, the mobile phone's shooting modes include single-frame shooting mode and multi-frame fusion shooting mode. Single-frame shooting mode means that the photo displayed to the user is obtained by the camera's HAL module processing a single RAW image. Multi-frame fusion shooting mode means that the photo displayed to the user is obtained by the camera's HAL module first fusing multiple RAW images to obtain a single RAW image, and then processing the fused RAW image.

[0118] In one possible implementation, the phone's camera mode is a single-frame shooting mode, where the first original image is captured by the camera, meaning that the first original image and the second original image are the same original image.

[0119] For example, the first original image can be captured by the camera before the user triggers the photo-taking function. Before the user triggers the photo-taking function, the camera HAL module can obtain the first face data.

[0120] For example, the first original image may be captured by the camera after the user triggers the photo-taking function. Then, after the user triggers the photo-taking function, S406-S409 are executed again to obtain the first face data.

[0121] In one possible implementation, the phone's shooting mode is a multi-frame fusion shooting mode (which can be called HDR). The first original image is obtained by fusing multiple original images, including the second original image, which means that the first original image is different from the second original image.

[0122] Exemplarily, the first raw image is obtained by image fusion processing of five continuous raw images, one of which is the second raw image. The number of raw images is not limited in the present application.

[0123] Taking five raw images including RAW1-RAW5 as an example, RAW1-RAW3 can be obtained by the camera before the user triggers the photographing function, and RAW4-RAW5 can be obtained by the camera after the user triggers the photographing function. The present application does not limit this.

[0124] In some embodiments, the second image parameter includes a resolution of 4096x3072 (i.e. the number of pixels) and a field of view angle of 65° (obtained based on the focal length and the photo size). The resolution of the first raw image is 4000x3000, and the field of view angle is 130°. Then the mobile phone can crop the first raw image based on the field of view angle of 65° to obtain an image with a pixel number of 2000x1500, and then scale up the cropped image by 2.048 times to obtain a photographing stream image with a field of view angle of 65° and a resolution of 4096x3072.

[0125] In addition, it should be noted that the field of view angle and the resolution included in the second image parameter can be the same as the field of view angle and the resolution corresponding to the first raw image, and in this case, the resolution conversion and the field of view angle conversion of the first raw image can be omitted.

[0126] For ease of understanding, in the embodiments of the present application, one vertex of the image is taken as the origin of the coordinate system, the upper boundary of the image is taken as the horizontal axis of the coordinate system, the left boundary of the image is taken as the vertical axis of the coordinate system, the right direction of the origin is taken as the positive direction of the horizontal axis of the coordinate system, and the downward direction of the origin is taken as the positive direction of the vertical axis of the coordinate system. For example, the vertex of the image can be the top-left corner vertex.

[0127] In some embodiments, the detection stream image corresponds to a first coordinate system, the first vertex of which can be the origin; the photographing stream image corresponds to a second coordinate system, the second vertex of which can be the origin; and the raw image corresponds to a third coordinate system, the third vertex of which can be the origin. It should be noted that the positions of the vertices in the corresponding images are the same, for example, all being the top-left corner vertex.

[0128] It can be understood that the coordinate value of a coordinate point of an image will change with the change of the origin of the coordinate system, and will also change with the enlargement or reduction of the image. The field of view angle conversion of the image can cause the change of the origin of the coordinate system. The resolution conversion of the image can cause the enlargement or reduction of the image. Therefore, the field of view angle conversion and the resolution conversion of the image can both cause the change of the coordinate value of the coordinate point on the image.

[0129] In some embodiments, when performing face processing on the face included in the image, the face position and the face feature point position need to be determined based on the coordinate data of the face first, and then the skin color is adjusted and the eyes are enlarged.

[0130] The first face data can include face data related to coordinate points, such as coordinate data of face positions and coordinate data of face feature points, which change with the change of the field of view angle and the resolution. The first face data can also include face data unrelated to coordinate points, such as face angle data, or face attribute data such as gender, age, posture, and expression, which do not change with the change of the field of view angle and the resolution. In the embodiments of the present application, the face data related to coordinate points can be referred to as first coordinate data, which represents the coordinate data of the face detected in the flow image in its corresponding coordinate system.

[0131] For example, the coordinate data of the face position includes the coordinate points of the four vertices of the face frame (such as a rectangle), and the rectangle frame contains a complete face. For example, the coordinate data of the face position includes the coordinate points of the vertices corresponding to multiple face frames, and each face frame contains a complete face.

[0132] For example, the coordinate data of the face feature points includes multiple coordinate points used to represent the shape and position of the feature points. For example, the coordinate points used to represent the shape and position of the left eye include: a left outer corner point, which is a point located on the outer side of the left eye and used to mark the outer corner position of the eye; a left inner corner point, which is a point located on the inner side of the left eye and used to mark the inner corner position of the eye; and the right eye is the same. For example, the coordinate points used to represent the shape of the nose include: a left side point of the nose, which is a point located on the left side of the nose and used to mark the shape of the nose; and a right side point of the nose, which is a point located on the right side of the nose. The coordinate points used to represent the position of the nose include: a tip point of the nose, which is a point located at the tip of the nose. The coordinate data of the face feature points can also include coordinate data of other feature points such as the mouth, the eyebrow, and the chin, which are not described here.

[0133] Based on the above examples, it can be known that the flow image is obtained by performing field of view angle conversion and resolution conversion on the second original image, and the photographing flow image is obtained by performing field of view angle conversion and resolution conversion on the first original image.

[0134] It should be understood that when the first original image is different from the second original image, the first original image is obtained by fusing multiple original images, and the multiple original images are obtained at a very close time. One of the multiple original images is the second original image, and therefore the change of the coordinate data corresponding to the first original image and the second original image can be ignored.

[0135] Exemplarily, one first face data can be randomly selected from the plurality of first face data of the plurality of detection stream images corresponding to the plurality of RAW images as the first face data corresponding to the first RAW image.

[0136] Exemplarily, the second RAW image can be taken as a reference image, and image fusion processing can be performed on the remaining RAW images, and the first face data corresponding to the second RAW image can be taken as the first face data corresponding to the first RAW image.

[0137] Therefore, in order to enable the camera HAL module to determine the coordinate point related to the face on the photograph stream image based on the first face data, that is, to obtain the coordinate point corresponding to the coordinate point of the first face data on the photograph stream image, the first face data can be first mapped to the first RAW image to obtain the second face data of the first RAW image, and then the second face data is mapped to the photograph stream image to obtain the third face data (also referred to as the converted specific object detection result). In addition, in some embodiments, the field of view angle of the detection stream image and the field of view angle of the photograph stream image are the same, but the resolutions of the two are different. In this case, the field of view angle of the detection stream image and the content contained in the photograph stream image are the same, and the difference lies in the size of the image. Therefore, the first face data can be directly mapped to the photograph stream image to obtain the third face data based on only the difference in resolution.

[0138] It should be noted that in the mapping process, the first face data including the face data related to the coordinate point (which can be referred to as first coordinate data) is first mapped to obtain third coordinate data, and the third coordinate data is used to replace the first coordinate data in the first face data to obtain the second face data; the second face data including the face data related to the coordinate point (that is, the third coordinate data) is secondly mapped to obtain second coordinate data, and the second coordinate data is used to replace the third coordinate data in the second face data to obtain the third face data.

[0139] Exemplarily, the coordinate value of the coordinate point included in the first face data can be multiplied by the square root of the ratio of the resolution of the photograph stream image to the resolution of the detection stream image as a multiple C (which can be referred to as a first multiple) to obtain the third face data.

[0140] S413: The camera HAL module maps the first face data of the detection stream image to the first RAW image to obtain the second face data.

[0141] Based on the above example, the detection stream image is obtained by the camera HAL module performing image processing on the second original image, and is obtained by first performing field of view conversion and then performing resolution conversion. Based on the above example, the second original image can be equivalent to the first original image. Therefore, the camera HAL module can map the first face data to the first original image, and it can be understood that the camera HAL module restores the detection stream image to the first original image to determine the change of the coordinate point of the first face data on the detection stream image to the coordinate point of the first original image. Therefore, the camera HAL module needs to first perform resolution conversion and then perform field of view conversion on the first face data.

[0142] It should be noted that the detection stream image is obtained by the camera HAL module performing resolution conversion and then performing field of view conversion on the second original image. Therefore, in order to restore, the camera HAL module can also first perform field of view conversion and then perform resolution conversion on the first face data.

[0143] In some embodiments, the camera HAL module can store the corresponding first image processing information when S408 is performed. For example, when the second original image is cropped based on the field of view included in the first image parameter, the camera HAL module stores the coordinate point of each vertex of the detection stream image on the second original image, and the data can be referred to as first vertex coordinate data. For example, the camera HAL module stores the coordinate point 1 (which can be referred to as second vertex coordinate data or third vertex coordinate data) of the top-left vertex of the detection stream image obtained by cropping on the second original image. When the resolution conversion is performed on the cropped image based on the resolution included in the first image parameter, the camera HAL module stores the scaling factor A (which can be referred to as second scaling factor) of the image, that is, stores the ratio of the resolution of the cropped original image to the resolution of the first image parameter.

[0144] In some embodiments, the mapping process of the first face data can include the following steps 1-3:

[0145] Step 1: The camera HAL module obtains the scaling factor A and the coordinate point 1 included in the stored first image processing information, and obtains the coordinate point 2 included in the first face data.

[0146] The coordinate point 2 includes a plurality of coordinate points corresponding to the face position and the feature point position included in the first face data.

[0147] It should be noted that the coordinate point 1 can also be the coordinate point of another vertex of the detection stream image on the second original image, which is not limited in the present application.

[0148] Step 2: The camera HAL module multiplies the coordinate value of the coordinate point 2 by the scaling factor A to obtain the coordinate point 3.

[0149] The detection stream image is magnified proportionally based on factor A, and coordinate point 3 is the coordinate point of multiple coordinate points included in the first face data in the magnified detection stream image.

[0150] Step 3: The camera HAL module adds the coordinate value of coordinate point 3 to the coordinate value of coordinate point 1 to obtain coordinate point 4, and uses coordinate point 4 to replace coordinate point 2 of the first face data, and uses the replaced first face data as the second face data.

[0151] Coordinate point 4 is the coordinate point of the first original image, which includes multiple coordinate points of the first face data.

[0152] For example, the value of the multiple A is u, and the coordinate point 1 is (x1, y1). For example... Figure 7 As shown, the first face data in the detection stream image includes face data 1. The top left corner of the detection stream image is the origin (0, 0). Taking the coordinates (x2, y2) of the top left corner of the face bounding box of face data 1 as an example (i.e., coordinate point 2), the camera HAL module needs to multiply the coordinate value of (x2, y2) with u to obtain coordinate point 3 (ux2, uy2). Then, the coordinate values ​​of coordinate point 1 and coordinate point 3 are added to obtain coordinate point 4 (x1+ux2, y1+uy2). The coordinates of other vertices of the face bounding box and the multiple coordinate points corresponding to the face feature points are handled similarly, and will not be elaborated here. Subsequently, the camera HAL module replaces coordinate point 1 of the first face data with coordinate point 4, and uses the face data including coordinate point 4 as the second face data.

[0153] S414: The camera HAL module maps the second face data to the image stream to obtain the third face data.

[0154] Based on the above example, the image stream is obtained by the camera HAL module processing the first original image. For example, the camera HAL module first performs field-of-view transformation on the first original image, and then performs resolution transformation to obtain the image stream. Therefore, mapping the second face data to the image stream can be understood as the camera HAL module performing image processing on the first original image to obtain the image stream, in order to determine the changes in the coordinates of the second face data points on the first original image to the coordinates of the points on the image stream.

[0155] In some embodiments, the camera HAL module can store the corresponding second image processing information when performing S412. Illustratively, when the first raw image is cropped based on the field of view angle included in the second image parameter, the coordinate points of the cropping can be referred to as second coordinate points. With the top-left vertex of the first raw image as the origin, the camera HAL module stores the coordinate point 5 (which can be referred to as the fifth vertex coordinate data, or the sixth vertex coordinate data) of the top-left vertex of the photographed stream image obtained by cropping. When the cropped image is converted in resolution based on the resolution included in the second image parameter, the camera HAL module stores the multiple B (which can be referred to as the third multiple) of the image in equal proportion. The coordinate data of the vertex of the photographed stream image in the first raw image can be referred to as the fourth vertex coordinate data.

[0156] In some embodiments, the square root of the ratio of the resolution of the photographed stream image to the resolution of the second raw image can be used as the multiple B. The multiple B is greater than 1, indicating that the resolution of the photographed stream image is larger, and the image is enlarged in equal proportion; the multiple B is less than 1, indicating that the resolution of the photographed stream image is smaller, and the image is reduced in equal proportion.

[0157] In some embodiments, the mapping process of the second face data can include the following steps 4-6:

[0158] Step 4: The camera HAL module obtains the multiple B and the coordinate point 5 included in the stored second image processing information, and obtains the coordinate point 4 included in the second face data.

[0159] It should be noted that the coordinate point 5 can also be the coordinate point of another vertex of the photographed stream image in the first raw image, which is not limited in the present application.

[0160] Step 5: The camera HAL module subtracts the coordinate value of the coordinate point 5 from the coordinate value of the coordinate point 4 to obtain the coordinate point 6.

[0161] The coordinate point 6 is the coordinate point of the multiple coordinate points included in the second face data in the cropped first raw image.

[0162] Step 6: The camera HAL module multiplies the coordinate value of the coordinate point 6 by the multiple B to obtain the coordinate point 7, replaces the coordinate point 4 of the second face data with the coordinate point 7, and uses the replaced second face data as the third face data.

[0163] The coordinate point 7 is the coordinate point of the multiple coordinate points included in the second face data in the photographed stream image.

[0164] Illustratively, the value of the multiple B is v, and the coordinate point 5 is (x3, y3). As Figure 8As shown, the second face data of the first original image includes face data 1. The top left corner of the first original image is the origin (0, 0). Taking the coordinates (x4, y4) of the top left corner of the face bounding box of face data 1 as an example (i.e., coordinate point 4), the camera HAL module needs to subtract the coordinates of coordinate point 5 from the coordinates of coordinate point 4 to obtain coordinate point 6 (x4-x3, y4-y3). Then, multiply (x4-x3, y4-y3) by the multiple Bv to obtain coordinate point 7 (vx4-vx3, vy4-vy3). The coordinates of other vertices of the face bounding box and the multiple coordinates corresponding to the face feature points are handled similarly, and will not be elaborated here. In this way, coordinate point 7 is used to replace coordinate point 4 of the second face data, and the face data including coordinate point 7 is used as the third face data.

[0165] Furthermore, the field of view included in the first image parameter and the second image parameter may be different, resulting in differences between the content contained in the detection stream image and the content contained in the captured stream image.

[0166] For example, when the field of view of the detection stream image is larger than that of the image capture stream image, the detection stream image contains more content than the image capture stream image. Therefore, a complete face in the detection stream image may be a face that does not exist or is incomplete in the image capture stream image, meaning that invalid faces exist in the image capture stream image. In this case, invalid faces in the third-party face data need to be filtered out.

[0167] like Figure 9 As shown, the first face data in the detection stream image includes face data 1, face data 2, and face data 3. In the captured stream image, the face indicated by face data 2 is a face that does not exist in the captured stream image, and the face indicated by face data 3 is an incomplete face in the captured stream image. Therefore, face data 2 and face data 3 are invalid face data in the captured stream image and need to be removed. The captured stream image only includes face data 1.

[0168] In some embodiments, a situation may arise where a face is detected in the detection stream image, but the captured image does not contain a complete face. In this case, all third-party face data is invalid and needs to be removed. The capture path then does not need to perform face-related processing. That is, when the camera preview interface of the camera application does not contain a complete face, the electronic device's capture path does not need to execute face-related algorithms after the user presses the capture control. For example, it does not need to execute... Figure 5 Algorithms b and c are shown.

[0169] In some embodiments, the detected stream image can include multiple faces, the first face data can include multiple face data (also referred to as multiple sets of detection results) corresponding to the multiple faces, and the converted third face data also includes multiple face data (converted) corresponding to the multiple faces. Assuming that each face data includes multiple coordinate points 7, for a face data, if there is a coordinate point 7 not on the photographed stream image, it means that the face data corresponding to the coordinate point 7 is invalid face data. For example, the multiple coordinate points 7 of the face data 2 described above are all not on the photographed stream image, and a part of the coordinate points 7 of the face data 3 described above (for example, the coordinate point of the top right corner of the face frame) are on the photographed stream image, and a part of the coordinate points 7 (for example, the coordinate point of the top left corner of the face frame) are not on the photographed stream image, both of which are invalid face data. In the embodiments of the present application, the coordinate point 7 not on the photographed stream image can also be referred to as a first coordinate point.

[0170] In some embodiments, whether a coordinate point 7 is on the photographed stream image can be determined by the coordinate value of the coordinate point.

[0171] For example, the coordinate value of a coordinate point 7 of the third face data is ≥ 0, indicating that the coordinate point 7 is on the photographed stream image, and the coordinate value of a coordinate point 7 is < 0, indicating that the coordinate point 7 is not on the photographed stream image. For each face data included in the third face data, if there is a coordinate point 7 with a coordinate value < 0, the face data is determined to be invalid face data. Then, the invalid face data of the third face data is removed, and the remaining face data is taken as fourth face data (also referred to as converted and screened specific object detection results).

[0172] In this way, the fourth face data is prevented from including invalid face data, which can cause the mobile phone to perform face processing on a face not on the photographed stream image, thereby avoiding increasing the shot2review time.

[0173] S415: The camera HAL module performs face processing on the photographed stream image based on the fourth face data to obtain a photographed picture.

[0174] The photographed picture is an image displayed to the user on the screen of the mobile phone. In some embodiments, the fourth face data includes face attributes, face positions, feature point positions, and face angles, and can also be referred to as converted and screened specific object detection results. Based on the fourth face data, the camera HAL module can perform face processing on the face on the photographed stream image by using a face processing algorithm to obtain an image displayed to the user.

[0175] In some embodiments, as Figure 5As shown, the image capture process includes algorithm a, algorithm b, and algorithm c. Algorithm a is an algorithm unrelated to face processing algorithms, such as noise reduction algorithms and brightness adjustment algorithms. Algorithms b and c are algorithms related to face processing algorithms, such as beautification algorithms like skin tone adjustment algorithms and facial feature point adjustment algorithms.

[0176] like Figure 5 As shown, the image processing method provided in this application embodiment, after the user triggers the camera application's photo-taking function, converts and filters the first face data obtained from the detected stream image, that is, reuses the first face data. The first face data is obtained by the camera HAL module detecting the detected stream image. The detected stream image has a small resolution and a fast processing speed, and the conversion speed of the first face data is also very fast, usually completed in about 1-2ms. Therefore, compared with the prior art, it can improve the acquisition speed of the fourth face data, thereby greatly saving shot2review time.

[0177] S416: The camera HAL module sends captured images to the camera application through the camera access interface.

[0178] S417: The camera app displays thumbnails of captured images.

[0179] like Figure 1 As shown, thumbnails of captured images are displayed in thumbnail display area 13. When a user triggers the capture of an image thumbnail, the phone can display the captured image.

[0180] It should be noted that the above embodiments take the reuse of the first face data as an example. This application does not limit this, and other data such as subject data and human body data obtained by the detection path to detect the detection stream image can also be reused.

[0181] Furthermore, in some embodiments, when the user triggers the recording function in video recording mode, the camera captures a series of RAW images. For each RAW image, the phone can reuse the first face data from the detection stream image corresponding to that RAW image to perform face processing on the corresponding capture stream image, resulting in a video frame displayed to the user. During the shooting process, the resolution of the detection stream image is lower than that of the capture stream image; therefore, the face detection processing speed for the detection stream image is faster, reducing the time from when the user clicks the shooting control to when the thumbnail of the captured video is displayed to the user, thereby improving the user experience.

[0182] Next, we will introduce the components of the electronic device.

[0183] It should be noted that the electronic device is a mobile phone in the above embodiments, which is only an example. In some embodiments, the electronic device can be a terminal device such as a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. The specific form of the electronic device is not specially limited in the present application, and the electronic device can only have a photographing function.

[0184] Exemplarily, Figure 10 A structural schematic diagram of the electronic device 1000 is shown.

[0185] The electronic device 1000 can include a processor 1010, an internal memory 1020, a camera 1030, and a display screen 1040.

[0186] It can be understood that the structure shown in the embodiment is not a specific limitation on the electronic device. In other embodiments, the electronic device can include more or fewer components than shown, or combine certain components, or split certain components, or different component arrangements. The components shown can be implemented in hardware, software, or a combination of software and hardware.

[0187] The processor 1010 can include one or more processing units, for example: the processor 1010 can include an application processor (AP), a graphics processing unit (GPU), an image signal processor (ISP), a digital signal processor (DSP), etc. Different processing units can be independent devices or integrated into one or more processors.

[0188] The processor 1010 can also be provided with a memory for storing instructions and data.

[0189] The internal memory 1020 can be used to store computer executable program codes, and the executable program codes include instructions. The processor 1010 executes various functional applications and data processing of the electronic device 1000 by running the instructions stored in the internal memory 1020.

[0190] In some embodiments, the internal memory 1020 stores instructions for performing the image processing method. The processor 1010 can implement the image processing method provided by the embodiments of the present application by executing the instructions stored in the internal memory 1020.

[0191] The electronic device 1000 implements the display function through the image processor, the display screen 1040, and the application processor, etc. The image processor is a microprocessor for image processing, connected to the display screen 1040 and the application processor. The image processor is used to perform mathematical and geometric calculations for graphics rendering. The processor 1010 can include one or more image processors that execute program instructions to generate or change display information. The display screen 1040 is used to display images, videos, etc.

[0192] In some embodiments, the display screen 1040 is used to display the running interface of the camera application of the electronic device 1000, such as the camera preview interface, the camera parameter setting interface, etc.

[0193] The electronic device 1000 can implement the shooting function through the ISP, the camera 1030, the video codec, the image processor, the display screen 1040, and the application processor, etc.

[0194] In some embodiments, the shooting function includes the photographing function and the video recording function of the camera application of the electronic device 1000.

[0195] The ISP is used to process the data fed back by the camera 1030. For example, when taking a picture, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, and the light signal is converted into an electrical signal. The camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also optimize the algorithm of the noise, brightness, and skin color of the image. The ISP can also optimize the exposure, color temperature, and other parameters of the shooting scene.

[0196] In some embodiments, the ISP can be arranged in the camera 1030. The camera 1030 is used to capture still images or videos. Objects generate optical images through lenses and project them onto photosensitive elements. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into a standard RGB, YUV, etc. format image signal. In some embodiments, the electronic device 1000 can include 1 or N cameras 1030, N being a positive integer greater than 1.

[0197] In some embodiments, the photosensitive element such as a CCD or CMOS is the image sensor mentioned in the embodiments of the present application, which converts the optical signal into an electrical signal. The image sensor transmits the electrical signal to the ISP to obtain a RAW image.

[0198] The embodiments of the present application also provide a computer readable storage medium storing a computer program, and the computer program can implement one or more steps in any of the image processing methods when executed by a computer.

[0199] The computer readable storage medium can be a non-transitory computer readable storage medium, for example, the non-transitory computer readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0200] The embodiments of the present application also provide a computer program product containing instructions. The computer program product can implement one or more steps in any of the image processing methods when executed by a computer.

[0201] The electronic device, the computer readable storage medium, and the computer program product provided by the embodiments of the present application are used to execute the corresponding image processing method provided above, and thus the beneficial effects achieved by the electronic device, the computer readable storage medium, and the computer program product can refer to the beneficial effects of the corresponding image processing method provided above, which will not be described here again.

[0202] The terms “first”, “second”, and “third” and the like in the specification of the present application, the claims, and the accompanying drawings are used to distinguish different objects, and are not used to limit a specific sequence.

[0203] In the embodiments of the present application, the words “exemplary” or “for example” are used to mean serving as an example, instance, or illustration. Any embodiment or design presented as “exemplary” or “for example” in the embodiments of the present application should not be interpreted as being more preferred or advantageous than other embodiments or designs. In fact, a word “exemplary” or “for example” is used to represent one of the embodiments of the present application in a specific manner.

[0204] The above-described embodiments are merely used to illustrate the technical solutions of the present application, but not for limiting the present application; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An image processing method, characterized in that, include: In response to the camera application being invoked, the image detection path of the camera application is started; the image detection path is used to detect specific objects in the detection stream image to obtain specific object detection results; In response to the camera application's photo-taking function being triggered, the camera application's photo-taking path is initiated; The image capture path is used to process the image capture stream to obtain the captured image; both the detection stream image and the image capture stream image are generated based on the original image obtained by the camera application; The resolution of the detection stream image is lower than the resolution of the captured stream image; The image capture path uses the specific object detection results to perform display effect processing on specific objects in the image capture stream to generate a captured image. The first vertex of the detection stream image is the origin of the first coordinate system; the second vertex of the image capture stream image is the origin of the second coordinate system; the position of the first vertex in the detection stream image is equivalent to the position of the second vertex in the image capture stream image; the specific object detection result includes the first coordinate data of the specific object in the detection stream image in the first coordinate system; the image capture path uses the specific object detection result to perform display effect processing on the specific object in the image capture stream image to generate an image capture, including: The image capture path converts the first coordinate data included in the specific object detection result into second coordinate data in the second coordinate system; The image capture path uses the converted specific object detection results to perform display effect processing on specific objects in the image capture stream to generate a captured image. The method further includes: The image capture path, based on the second coordinate data, filters out data that does not belong to the image capture stream from the transformed specific object detection results; The image capture path utilizes the converted object detection results to perform display effect processing on specific objects in the image capture stream, generating a captured image, including: The image capture path uses the transformed and filtered specific object detection results to perform display effect processing on specific objects in the image capture stream, generating a captured image.

2. The method according to claim 1, characterized in that, The detected stream image is obtained through the following steps: The image detection path reduces the resolution of the original image to obtain the detection stream image.

3. The method according to claim 1, characterized in that, The detection stream image includes multiple specific objects, and the converted specific object detection results include multiple sets of detection results corresponding one-to-one with the multiple specific objects; the image capture path, based on the second coordinate data, filters out data that does not belong to the image capture stream image from the converted specific object detection results, including: The image capture path filters out first coordinate points that are not in the image capture stream from the second coordinate data; The imaging path determines the target detection result to which the first coordinate point belongs from the multiple sets of detection results; The imaging path filters out the target detection results from the multiple sets of detection results.

4. The method according to claim 1, characterized in that, The image capture path converts the first coordinate data included in the specific object detection result into second coordinate data in the second coordinate system, including: The imaging path determines that the field of view of the detection stream image is the same as the field of view of the captured stream image; The image capture path calculates the square root of the ratio of the resolution of the captured image stream to the resolution of the detected image stream as the first multiple; The imaging path multiplies the first multiple by the coordinate value of the first coordinate data to obtain the second coordinate data.

5. The method according to claim 1, characterized in that, The third vertex of the original image is the origin of the third coordinate system; the position of the third vertex in the original image is the same as the position of the second vertex in the captured image stream. The image capture path converts the first coordinate data included in the specific object detection result into second coordinate data in the second coordinate system, including: The imaging path determines that the field of view of the detection stream image is different from the field of view of the captured stream image; The imaging path converts the first coordinate data into third coordinate data in the third coordinate system; The imaging path converts the third coordinate data into the second coordinate data.

6. The method according to claim 5, characterized in that, The imaging path converts the first coordinate data into third coordinate data in the third coordinate system, including: The image capture path acquires first vertex coordinate data and a second multiple; the first vertex coordinate data is the data of the vertices of the detection stream image in the third coordinate system; the second multiple is the square root of the ratio of the resolution of the original image to the resolution of the detection stream image; The imaging path converts the first coordinate data into the third coordinate data based on the first vertex coordinate data and the second multiple.

7. The method according to claim 6, characterized in that, The first vertex coordinate data includes the second vertex coordinate data of the first vertex of the detected stream image in the third coordinate system; the imaging path converts the first coordinate data into the third coordinate data based on the first vertex coordinate data and the second multiple, including: The imaging path uses the coordinate values ​​of the first coordinate data and the second multiple to obtain the fourth coordinate data; The imaging path obtains the third coordinate data by adding the coordinate values ​​of the fourth coordinate data to the coordinate values ​​of the second vertex coordinate data.

8. The method according to claim 7, characterized in that, The first vertex coordinate data includes the third vertex coordinate data of the first vertex of the detected stream image in the third coordinate system; the imaging path converts the first coordinate data into the third coordinate data based on the first vertex coordinate data and the second multiple, including: The imaging path uses the coordinate values ​​of the first coordinate data and the coordinate values ​​of the third vertex coordinate data to obtain the fifth coordinate data; The imaging path uses the coordinate value of the fifth coordinate data to multiply by the second multiple to obtain the third coordinate data.

9. The method according to claim 5, characterized in that, The imaging path converts the third coordinate data into the second coordinate data, including: The image capture path acquires the coordinate data of the fourth vertex and the third multiple; the coordinate data of the fourth vertex is the data of the vertex of the image capture stream in the third coordinate system; the third multiple is the square root of the ratio of the resolution of the image capture stream to the resolution of the original image; The imaging path converts the third coordinate data into the second coordinate data based on the fourth vertex coordinate data and the third multiple.

10. The method according to claim 9, characterized in that, The fourth vertex coordinate data includes the fifth vertex coordinate data of the second vertex of the captured image in the third coordinate system; the capturing path converts the third coordinate data into the second coordinate data based on the fourth vertex coordinate data and the third multiple, including: The imaging path uses the coordinate value of the third coordinate data and the third multiple to obtain the sixth coordinate data; The imaging path obtains the second coordinate data by subtracting the coordinate value of the fifth vertex coordinate data from the coordinate value of the sixth coordinate data.

11. The method according to claim 9, characterized in that, The fourth vertex coordinate data includes the sixth vertex coordinate data of the second vertex of the captured image in the third coordinate system; the capturing path converts the third coordinate data into the second coordinate data based on the fourth vertex coordinate data and the third multiple, including: The imaging path uses the coordinate value of the third coordinate data to subtract the coordinate value of the sixth vertex coordinate data to obtain the seventh coordinate data; The imaging path obtains the second coordinate data by multiplying the coordinate value of the seventh coordinate data by the third multiple.

12. The method according to any one of claims 1-11, characterized in that, The captured images are obtained through the following steps: The imaging path reduces the field of view of the original image to obtain the imaging stream image.

13. The method according to claim 12, characterized in that, The image capture path reduces the field of view of the original image to obtain the image capture stream image, including: The image capture path obtains the zoom level of the camera application; The imaging path determines the second coordinate point of the original image based on the zoom ratio; The image capture path obtains the image capture stream by cropping the original image based on the second coordinate point.

14. The method according to any one of claims 1-11, characterized in that, Both the detection stream image and the image capture stream image are generated based on the first original image obtained from the camera application.

15. The method according to any one of claims 1-11, characterized in that, The captured image stream is generated based on a first original image obtained by the camera application, and the detected image stream is generated based on a second original image obtained by the camera application; the position of a specific object in the first original image is the same as the position of a specific object in the second original image.

16. The method according to claim 15, characterized in that, The first original image was obtained through the following steps: The first original image is obtained by performing image fusion processing on multiple original images, including the second original image.

17. The method according to any one of claims 1-11, characterized in that, The specific object includes a human face; the image capture path uses the detection results of the specific object to perform display effect processing on the specific object in the image capture stream, generating a captured image, including: The image capture path uses the specific object detection results to retouch the faces in the image capture stream and generate a captured image.

18. An electronic device, characterized in that, Including memory and processor; The memory is coupled to the processor and is used to store computer program code, the computer program code including computer instructions, wherein one or more of the processors invoke the computer instructions to cause the electronic device to perform the image processing method as described in any one of claims 1-17.

19. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the image processing method as described in any one of claims 1-17.

Citation Information

Patent Citations

  • Internet-based face beautifying system

    CN107392110A