Image processing method, electronic equipment and storage medium
By using image detection paths in electronic devices to detect specific objects on low-resolution detection stream images, the problem of too long photography time is solved, and faster photography speed and better user experience is achieved.
Patent Information
- Application Number
- CN202311867816.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-12-29
AI Technical Summary
When users take photos with electronic devices, it takes a long time from triggering the photo function to displaying thumbnails, which affects the user experience.
The image detection path is used to detect specific objects on the detection stream image, generate specific object detection results, and directly use this result to process the photographic stream image after the photography function is triggered, reducing dependence on high-resolution images and improving detection speed.
It shortens the time from taking photos to displaying thumbnails, and improves the user's photo speed and user experience.
Smart Images

Figure CN120282012A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and particularly to an image processing method, an electronic device, and a storage medium. Background Art
[0002] With the wide popularity of electronic devices such as mobile phones and tablet computers and the rapid development of photographing technologies, more and more users choose to use electronic devices to record wonderful moments in life. Therefore, people widely use the photographing function of electronic devices to take pictures in daily life.
[0003] Taking a mobile phone as an example, the mobile phone has an instant photo display function. When a user triggers the photographing function of the mobile phone to take a picture, the mobile phone immediately displays the picture in the form of a thumbnail in a specific area of the mobile phone screen (such as the lower left corner, etc.). The user can confirm that the mobile phone has finished taking a picture by seeing the thumbnail. Subsequently, when the user triggers the thumbnail, the mobile phone can display the picture taken by the mobile phone for browsing. However, the time consumed from when the user triggers the photographing function of the electronic device to when the electronic device displays the thumbnail is relatively long, and the user feels that the time for finishing taking a picture is relatively long, which is likely to affect the user experience. Summary of the Invention
[0004] To solve the above problems, this application provides an image processing method, an electronic device, and a storage medium. The aim is to reduce the time consumed from when the user triggers the photographing function of the electronic device to when the electronic device displays the thumbnail, improve the photographing speed, and thus improve the user experience.
[0005] In a first aspect, the present application provides an image processing method. Exemplarily, this method can be applied to an electronic device, which can be a device including a camera application such as a mobile phone, a tablet computer, a laptop computer, etc. In this method, the camera application is called. For example, the user clicks on the icon of the camera application to start the camera application. In response to being called, the electronic device can activate the image detection path of the camera application, which can detect a specific object in the detection stream image to obtain a specific object detection result. Exemplarily, the specific object can be a face, a human body, or other objects in the image. Then, when the photographing function of the camera application is triggered, for example, the user clicks on the photographing control of the camera application to trigger the photographing function, or the user issues a gesture instruction or a voice instruction to trigger the photographing function. In response to the photographing function being triggered, the electronic device can activate the photographing path of the camera application, which is used to process the photographing stream image to obtain a photographed picture; wherein, both the above-mentioned detection stream image and the above-mentioned photographing stream image are generated based on the original image obtained by the camera application. For example, based on the hardware associated with the photographing function of the camera application, the original image captured by the camera of the electronic device, and the resolution of the detection stream image is lower than that of the photographing stream image; subsequently, the photographing path can use the specific object detection result obtained by the detection path to perform display effect processing on the specific object in the photographing stream image to generate a photographed picture. Exemplarily, the specific object detection result can be sent to the photographing path after being obtained by the image detection path, or can be obtained by the photographing path from the image detection path.
[0006] In this way, after the photographing function of the camera application is triggered, the photographing path does not need to detect the specific object based on the higher-resolution photographing stream image, but can directly use the specific object detection result obtained by the image detection path. Because compared with the photographing stream image, the resolution of the detection stream image is lower, the image detection path can detect faster, the photographing path can obtain the specific object detection result faster, and can generate the photographed picture earlier, thereby improving the photographing speed felt by the user.
[0007] In a possible implementation manner, after the electronic device obtains the original image, the image detection path can reduce its resolution to obtain a detection stream image with a lower resolution. In this way, the detection speed of the detection path for the specific object in the detection stream image is improved, and the specific object detection result can be obtained faster.
[0008] In a possible implementation, the first vertex of the detection stream image is the origin of the first coordinate system; the second vertex of the captured stream image is the origin of the second coordinate system; the position of the first vertex in the detection stream image is the same as the position of the second vertex in the captured stream image. For example, the upper left corner vertex of the detection stream image is the origin of the first coordinate system, and at the same time, the upper left corner vertex of the captured stream image is the origin of the second coordinate system. The above specific object detection result may include the first coordinate data of the specific object in the detection stream image in the first coordinate system. Exemplarily, the first coordinate data may be the coordinate data of a human face, facial feature points, etc. in the first coordinate system. The capture path may first convert the first coordinate data included in the specific object detection result into the second coordinate data in the second coordinate system; then, the capture path may use the converted specific object detection result to perform display effect processing on the specific object in the captured stream image to generate a captured picture.
[0009] In this way, considering that the positions of the specific object in the detection stream image and the captured stream image deviate due to the different corresponding coordinate systems, the converted second coordinate data obtained after the capture path performs the conversion corresponds to the position of the specific object in the captured stream image. Therefore, the accuracy of the display effect processing can be improved, and a captured picture with higher quality can be obtained.
[0010] In a possible implementation, the image processing method may further include: based on the second coordinate data, the capture path may filter out the data that does not belong to the captured stream image from the converted specific object detection result; then, the capture path may use the converted and filtered specific object detection result to perform display effect processing on the specific object in the captured stream image to generate a captured picture. In this way, considering that the field of view angles corresponding to the detection stream image and the captured stream image are different, which may lead to a deviation in the number of specific objects included in the two, and the specific object detection result is obtained by detecting the detection stream image. When the field of view angle of the detection stream image is larger than that of the captured stream image, there may be multiple specific objects in the detection stream image, but one or some of the multiple specific objects may not be completely in the captured stream image. Therefore, it is necessary to filter out the specific object detection result. The converted and filtered specific object detection result only includes the detection results corresponding to the specific objects existing in the captured stream image. Therefore, it can avoid increasing the time for display effect processing, generate the captured picture more quickly, and further improve the captured speed felt by the user.
[0011] In a possible implementation, the detected stream image may include multiple specific objects, and the converted specific object detection results may include multiple sets of detection results corresponding one-to-one to the multiple specific objects. The photographing path may first screen out the first coordinate points that are not in the photographing stream image from the second coordinate data. For example, if the upper left corner of the photographing stream image is the origin of the second coordinate system and the coordinate values of the coordinate points on the photographing stream image are all non-negative, then the coordinate points with negative coordinate values can be screened out from the coordinate points included in the second coordinate data as the first coordinate points. Then, the photographing path may determine the target detection result to which the first coordinate point belongs from the multiple sets of detection results. Subsequently, the photographing path may screen out the target detection result from the converted specific object detection results, that is, the multiple sets of detection results. In this way, the specific objects included in the converted and screened specific object detection results are all on the photographing stream image, which can avoid wasting time on detecting incomplete or non-existent specific objects on the photographing stream image and further improve the photographing speed felt by the user.
[0012] In a possible implementation, the photographing path may first determine the field of view angles of the detected stream image and the photographing stream image. When it is determined that the two are the same, the photographing path may calculate the square root of the ratio of the resolution of the photographing stream image to the resolution of the detected stream image as the first multiple. Subsequently, the photographing path multiplies the first multiple by the coordinate values of the first coordinate data to obtain the second coordinate data. In this way, when the field of view angles of the detected stream image and the photographing stream image are the same, it indicates that the contents included in the detected stream image and the photographing stream image are the same and there is no deviation. Then, the coordinate conversion can be performed only based on the difference in their resolutions, which can ensure the accuracy of the converted second coordinate data.
[0013] In a possible implementation, the third vertex of the original image is the origin of the third coordinate system; the position of the third vertex in the original image is the same as the position of the second vertex in the photographing stream image. Exemplarily, the upper left corner vertex of the original image is the origin of the third coordinate system, and at the same time, the upper left corner vertex of the photographing stream image is the origin of the second coordinate system. The photographing path may first determine the field of view angles of the detected stream image and the photographing stream image. When it is determined that the two are different, the photographing path may first convert the first coordinate data into the third coordinate data in the third coordinate system; then the photographing path converts the third coordinate data into the second coordinate data. In this way, when the field of view angles of the detected stream image and the photographing stream image are different, it indicates that there is a deviation in the contents included in the detected stream image and the photographing stream image, and the first coordinate data cannot be directly converted into the second coordinate data. Since both the detected stream image and the photographing stream image are obtained based on the original image, the third coordinate system where the original image is located can be used for conversion, thereby ensuring the accuracy of the converted second coordinate data.
[0014] In a possible implementation, the photographing path may first obtain first vertex coordinate data and a second multiple, where the first vertex coordinate data is the data of the vertices of the detected stream image in a third coordinate system; the second multiple is the square root of the ratio of the resolution of the original image to the resolution of the detected stream image; the photographing path may convert the first coordinate data into third coordinate data based on the first vertex coordinate data and the second multiple. In this way, considering the differences in resolution and field of view angle between the original image and the detected stream image, the accuracy of coordinate data conversion is ensured.
[0015] In a possible implementation, the above first vertex coordinate data includes second vertex coordinate data of the first vertex of the detected stream image in the third coordinate system. Exemplarily, it may include the coordinate values of the upper left corner vertex of the detected stream image in the third coordinate system; the photographing path may first multiply the coordinate values of the first coordinate data by the second multiple to obtain fourth coordinate data, that is, first perform resolution conversion on the first coordinate data to obtain fourth coordinate data, and then add the coordinate values of the fourth coordinate data to the coordinate values of the second vertex coordinate data to obtain third coordinate data, that is, then perform field of view angle conversion on the fourth coordinate data to obtain third coordinate data. In this way, when the detected stream image is obtained by first performing field of view angle conversion on the original image and then performing resolution conversion, the photographing path can first perform conversion on the coordinate points based on different resolutions, and then perform conversion based on different field of view angles, so as to restore a specific object in the detected stream image to the original image and ensure the accuracy of coordinate data conversion.
[0016] In a possible implementation, the above first vertex coordinate data includes third vertex coordinate data of the first vertex of the detected stream image in the third coordinate system. Exemplarily, it may include the coordinate values of the upper left corner vertex of the detected stream image in the third coordinate system; the photographing path may first add the coordinate values of the first coordinate data to the coordinate values of the third vertex coordinate data to obtain fifth coordinate data, that is, first perform field of view angle conversion on the first coordinate data to obtain fifth coordinate data, and then multiply the coordinate values of the fifth coordinate data by the second multiple to obtain third coordinate data, that is, then perform resolution conversion on the fifth coordinate data to obtain third coordinate data. In this way, when the detected stream image is obtained by first performing resolution conversion on the original image and then performing field of view angle conversion, the photographing path can first perform conversion on the coordinate points based on different field of view angles, and then perform conversion based on different resolutions, so as to restore a specific object in the detected stream image to the original image and ensure the accuracy of coordinate data conversion.
[0017] In a possible implementation, when converting the third coordinate data into the second coordinate data in the photographing path, the fourth vertex coordinate data and the third multiple can be obtained first. The fourth vertex coordinate data is the data of the vertices of the photographing stream image in the third coordinate system. For example, the coordinate points corresponding to the respective vertices of the photographing stream image in the third coordinate system. The third multiple is the square root of the ratio of the resolution of the photographing stream image to the resolution of the original image, that is, the multiple by which the original image needs to be enlarged or reduced to be converted into the photographing stream image. Subsequently, the photographing path can convert the third coordinate data into the second coordinate data based on the fourth vertex coordinate data and the third multiple. In this way, considering the differences in resolution and field of view angle between the original image and the photographing stream image, the accuracy of coordinate data conversion is ensured.
[0018] In a possible implementation, the above-mentioned fourth vertex coordinate data includes the fifth vertex coordinate data of the second vertex of the photographing stream image in the third coordinate system. Exemplarily, it may include the coordinate values of the upper left vertex of the photographing stream image in the third coordinate system. The photographing path can multiply the coordinate values of the third coordinate data by the third multiple to obtain the sixth coordinate data, that is, first perform resolution conversion on the third coordinate data to obtain the sixth coordinate data. Subsequently, subtract the coordinate values of the fifth vertex coordinate data from the coordinate values of the sixth coordinate data to obtain the second coordinate data, that is, perform field of view angle conversion on the sixth coordinate data to obtain the second coordinate data. In this way, the photographing stream image can be obtained by first performing resolution conversion on the original image and then performing field of view angle conversion, which can ensure the accuracy of coordinate data conversion.
[0019] In a possible implementation, the above-mentioned fourth vertex coordinate data includes the sixth vertex coordinate data of the second vertex of the photographing stream image in the third coordinate system. Exemplarily, it may include the coordinate values of the upper left vertex of the photographing stream image in the third coordinate system. The photographing path can first subtract the coordinate values of the sixth vertex coordinate data from the coordinate values of the third coordinate data to obtain the seventh coordinate data, that is, first perform field of view angle conversion on the third coordinate data to obtain the seventh coordinate data. Subsequently, multiply the coordinate values of the seventh coordinate data by the third multiple to obtain the second coordinate data, that is, perform resolution conversion on the seventh coordinate data to obtain the second coordinate data. In this way, the photographing stream image can be obtained by first performing field of view angle conversion on the original image and then performing resolution conversion, which can ensure the accuracy of coordinate data conversion.
[0020] In a possible implementation, the above-mentioned photographing stream image can be obtained by the photographing path reducing the field of view angle of the original image. Exemplarily, when the user adjusts the focal length of the camera application from 1.0x to 2.0x, it indicates that the user needs to reduce the content included in the photographed picture, so the field of view angle needs to be reduced. In this way, a photographed picture that meets the user's needs can be obtained.
[0021] In a possible implementation, when the photographing path reduces the field of view angle of the original image to obtain a photographing stream image, the photographing path may first obtain the zoom ratio of the camera application, and then, based on the zoom ratio, determine the second coordinate point of the original image, that is, the coordinate point required for cropping the original image. Subsequently, the original image is cropped based on the second coordinate point to obtain a photographing stream image. In this way, a photographed picture that meets the user's needs can be obtained.
[0022] In a possible implementation, the above-mentioned detection stream image and the above-mentioned photographing stream image may both be generated based on a first original image obtained by a camera application, that is, both are generated based on the same original image. Exemplarily, the first original image may be captured by the camera before the photographing function of the camera application is triggered, or may be captured by the camera after the photographing function of the camera application is triggered.
[0023] In a possible implementation, the above-mentioned photographing stream image may be generated based on a first original image obtained by a camera application, and the above-mentioned detection stream image may be generated based on a second original image obtained by a camera application, that is, both are generated based on different original images. Among them, the position of a specific object in the first original image is the same as the position of the specific object in the second original image.
[0024] In a possible implementation, the above-mentioned first original image may be obtained by performing image fusion processing on multiple original images including the second original image. Exemplarily, when the camera application is in the HDR shooting mode, the photographed picture presented to the user may be obtained based on multiple original images.
[0025] In a possible implementation, the above-mentioned specific object includes a human face, and the specific object detection result is also the human face detection result. Exemplarily, the human face detection result may include coordinate data of a human face frame, coordinate data of feature points in the human face, human face attributes, human face angles, etc. Then, the photographing path may use the human face detection result to perform retouching processing on the human face in the photographing stream image, such as beauty processing such as enlarging eyes and skin beautification, to generate a photographed picture that can be presented to the user.
[0026] In a second aspect, the present application provides an electronic device, which includes a memory and a processor; the memory stores computer program code, and the computer program code includes computer instructions; one or more processors call the computer instructions to enable the electronic device to execute the image processing method in the first aspect above.
[0027] In a third aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the image processing method in the first aspect above is implemented.
[0028] As can be seen from the above technical solutions, the present application has the following beneficial effects:
[0029] After the camera application is called, the electronic device can activate the image detection path of the camera application, which is used to detect specific objects in the detection stream image to obtain specific object detection results; after the photographing function of the camera application is triggered, the electronic device can activate the photographing path of the camera application, which is used to process the photographing stream image to obtain a photographed picture; both the detection stream image and the photographing stream image are generated based on the original image obtained by the camera application, and the resolution of the detection stream image is lower than that of the photographing stream image. Therefore, compared with the photographing path, the image detection path can obtain the specific object detection results faster. The photographing path can directly reuse the specific object detection results to process the display effect of the specific object in the photographing stream image and generate the photographed picture faster. Thus, the image processing method provided by the present application can improve the acquisition speed of the photographed picture, and then present the photographed picture to the user faster, reduce the photographing completion time felt by the user, and improve its user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 Schematic diagram of a photographing process in a related technology provided by an embodiment of the present application;
[0031] Figure 2 Schematic diagram of a camera operation process in a related technology provided by an embodiment of the present application;
[0032] Figure 3 Schematic diagram of a face detection module provided by an embodiment of the present application;
[0033] Figure 4 Signaling interaction diagram of an image processing method provided by an embodiment of the present application;
[0034] Figure 5 Schematic diagram of reusing first face data provided by an embodiment of the present application;
[0035] Figure 6a Schematic diagram of adjusting the focal length provided by an embodiment of the present application;
[0036] Figure 6b Schematic diagram of adjusting the photo ratio provided by an embodiment of the present application;
[0037] Figure 7 Schematic diagram of mapping first face data provided by an embodiment of the present application;
[0038] Figure 8 Schematic diagram of mapping second face data provided by an embodiment of the present application;
[0039] Figure 9Schematic diagram of invalid face data provided by an embodiment of this application;
[0040] Figure 10 Schematic structural diagram of an electronic device provided by an embodiment of this application. Detailed implementation manners
[0041] For the sake of clear and concise description of the following embodiments, first, the vocabulary involved in the embodiments of this application will be described. It should be understood that this description is for a clearer understanding of the embodiments of this application and does not necessarily constitute a limitation on the embodiments of this application.
[0042] Image sensor: Used to sense light. When light shines on the image sensor, the image sensor converts the optical signal into an electrical signal. In the embodiments of this application, the camera of the electronic device includes an image sensor.
[0043] Original image: It can be called a RAW image, which is an electrical signal converted by the image sensor and then a pre-processed image obtained by the image signal processor. In some embodiments, the camera of the electronic device includes an image signal processor.
[0044] Field of view (FOV): Used to indicate the maximum angular range that the camera can capture. If the object to be photographed is within this angular range, the object to be photographed will be captured by the camera. If the object to be photographed is outside this angular range, the object to be photographed will not be captured by the camera. The larger the field of view of the camera, the larger the photographing range; the smaller the field of view of the camera, the smaller the photographing range.
[0045] Focal length range: Refers to the focal length range of the camera, which can also be called the zoom ratio. In some embodiments, the focal length range of the electronic device is 0.5x - 30x, where x represents the fixed focal length of the electronic device.
[0046] Resolution: Refers to the number of pixels in the vertical and horizontal directions of the image, which determines the fineness of the image details and is used to represent the clarity of the image. The higher the resolution, the more pixels it contains, and the clearer the image. In some embodiments, the resolution is 400×300, indicating that the image includes 400 pixels vertically and 300 pixels horizontally.
[0047] Detected stream image: An image with a lower resolution obtained by performing image processing on the original image obtained through the camera based on the first image parameter of the detected stream image. In some embodiments, the electronic device can perform perception detections such as object detection, face detection, and human body detection on the detected stream image to identify the object to be photographed and then determine the photographing scene.
[0048] Among them, subject detection refers to detecting the objects included in the detection stream image. Exemplarily, the objects include people, animals, plants, etc. Face detection refers to detecting the faces included in the detection stream image, including detecting the position of the face in the image, the positions of feature points such as the eyes, nose, and mouth of the face in the image, face attributes, face angles, etc. Human body detection refers to detecting the human bodies included in the detection stream image, including detecting the positions of the limbs and torso in the image, etc.
[0049] Photo stream image: An image obtained by performing image processing on the original image obtained by the camera based on the second image parameter of the photo stream image. In some embodiments, the electronic device can optimize the photo stream image to obtain a higher-quality image for display to the user.
[0050] Next, in combination with related technologies, the technical advantages of an image processing method, an electronic device, and a storage medium provided by this application will be compared and described. For the convenience of understanding, it will be described through an example scenario. In this example scenario, the electronic device is mobile phone 100.
[0051] In related technologies, the time consumed from when the user triggers the camera function until mobile phone 100 captures a photo that can be presented to the user can be referred to as the shot2review time. Taking the example of the user taking a portrait photo, the user can use the camera function of mobile phone 100 (i.e., the function of taking a photo) to take a portrait photo including a face.
[0052] As Figure 1 shown, mobile phone 100 displays the camera preview interface 11 of the camera application. The camera preview interface includes a shooting control 12 and a thumbnail display area 13. After the user holds mobile phone 100 and aims it at the face to be photographed, the user can click the shooting control 12 to trigger the camera function of mobile phone 100. Mobile phone 100 displays the captured photo in the form of a thumbnail in the thumbnail display area 13.
[0053] Combined with Figure 1 、 Figure 2 and Figure 3 shown, during the operation of the camera application, the image sensor continuously works and multiple RAW images are collected. Moreover, the detection path (also referred to as the image detection path) and the preview path also continuously work for each RAW image. After the user clicks the shooting control 12 (i.e., triggers the camera function), the shooting path starts to work. Mobile phone 100 will first perform relevant processing on a RAW image to obtain a photo stream image, then perform relevant processing such as denoising on the photo stream image through algorithm a of the shooting path, and subsequently perform face processing through the face module.
[0054] The face module includes a face detection module and a face processing module. First, the mobile phone 100 performs various detections on the face included in the captured image stream through the face detection module, such as face position detection, face attribute detection, and face feature point detection, to obtain face information. Subsequently, based on the face information, the mobile phone 100 performs face processing on the face included in the captured image stream through algorithms b and c included in the face processing module. For example, relevant processing such as adding beauty effects (big eyes, face slimming, skin smoothing) to the face based on the face information. Finally, the mobile phone 100 displays the captured portrait photo in the thumbnail display area 13 in the form of a thumbnail.
[0055] However, in the related art, when the mobile phone 100 takes a photo, it needs to complete the above-mentioned face detection, and may also need to complete other detections such as subject detection and human body detection. The detection may take 100ms - 200ms. Subsequently, the mobile phone 100 needs to perform relevant processing on the face, human body, subject, etc. according to the detected information to obtain the final photo, that is, the shot2review is relatively long. Therefore, the time from when the user clicks the camera control 12 to when the thumbnail is displayed in the thumbnail display area 13 is also relatively long, resulting in a slow photo-taking speed felt by the user, which is likely to affect the user experience.
[0056] Therefore, to solve the above problems, the embodiments of the present application provide an image processing method, an electronic device, and a storage medium. After the camera application's photo-taking function is triggered, the photo-taking path no longer performs face detection on the captured image stream, but directly reuses the face information obtained by detecting the detection image stream through the detection path. The resolution of the detection image stream is lower than that of the captured image stream. Therefore, the speed of face detection can be accelerated, the speed of obtaining face information by the photo-taking path can be improved, the face can be processed earlier, the speed of obtaining the photo (also called the captured picture) can be increased, thereby improving the photo-taking speed felt by the user and enhancing its user experience.
[0057] Next, still taking the electronic device as a mobile phone and the user using the mobile phone's photo-taking function to capture a portrait photo including a face as an example, in combination with Figure 4 and Figure 5 The image processing method provided by the embodiments of the present application will be described in detail.
[0058] Taking the Android system as an example for the operating system of the mobile phone. The Android system can adopt a layered architecture, and each layer has clear roles and divisions. The layers communicate with each other through software interfaces. In some embodiments, the system is divided into five layers, from top to bottom, namely the application layer, the application framework layer, the hardware abstraction layer (HAL (Hardware Abstraction Layer)), the driver layer, and the hardware layer. It should be noted that the mobile phone can also run the iOS operating system, and the present application does not make any limitations in this regard.
[0059] As Figure 4 shown, the application layer includes a camera application (hereinafter simply referred to as the camera), the application framework layer includes a camera access interface, the hardware abstraction layer includes a camera HAL module, the driver layer includes a camera driver, and the hardware layer includes a camera.
[0060] The image processing method includes the following steps:
[0061] S401: The camera application receives a start operation of the user on the camera application.
[0062] The application layer may include a series of application packages. In the embodiment of the present application, the application package may include a camera application.
[0063] The start operation refers to a trigger operation by which the user starts the camera application of the mobile phone.
[0064] Exemplarily, the start operation includes a click operation on the application icon of the camera application included in the main interface of the mobile phone by the user. Also exemplarily, the start operation includes a click operation on the application icon of the camera application on the lock screen interface of the mobile phone by the user, etc. Still exemplarily, the start operation includes an operation in which the user issues voices such as "open the camera", "take a selfie", etc., to trigger the intelligent voice service of the mobile phone. The present application does not make a limitation thereto.
[0065] S402: The camera application sends a start request of the camera application to the camera access interface.
[0066] The application framework layer provides application programming interfaces and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions. In the embodiment of the present application, the application framework layer may include a camera access interface for providing application programming interfaces and programming frameworks for the camera application.
[0067] In response to the operation of the user starting the camera application, the camera application calls the camera access interface in the application framework layer to start the camera application.
[0068] S403: The camera access interface sends a call request to the camera HAL module.
[0069] The HAL layer is an interface layer located between the application framework layer and the driver layer. In the embodiment of the present application, the hardware abstraction layer may include a camera HAL module for providing the running code and data for implementing the image processing method provided in the embodiment of the present application.
[0070] The camera access interface calls the camera HAL module to implement the image processing method provided in the embodiment of the present application.
[0071] S404: The camera HAL module sends a start request of the camera to the camera driver.
[0072] The driver layer is the layer between hardware and software. The driver layer includes drivers for various hardware. In the embodiments of the present application, the driver layer may include a camera driver, etc., which is used to drive the image sensor of the camera to convert the sensed optical signal into an electrical signal, and drive the image signal processor to preprocess the electrical signal to obtain a raw image. In some embodiments, the image signal processor may be disposed in the camera.
[0073] S405: The camera driver starts the camera.
[0074] After receiving the start request of the camera, the camera driver starts the image sensor and the image signal processor.
[0075] S406: The camera captures a second raw image.
[0076] The image sensor of the camera can sense the optical signal. The image sensor converts the sensed optical signal into an electrical signal, and the image signal processor preprocesses the electrical signal to obtain a second raw image.
[0077] In some embodiments, the image parameters of the second raw image are obtained based on the parameters of the camera. Exemplarily, the parameters of the camera include the field of view angle, pixels, image ratio, etc. For example, if the number of pixels of the camera is 12 million pixels, then the second raw image captured by the camera includes 12 million pixels. The image ratio of the camera is 4:3, then the ratio of the number of pixels included longitudinally in the second raw image to the number of pixels included horizontally in the second raw image is 4:3, indicating that the resolution of the second raw image obtained by the camera is 4000×3000.
[0078] S407: The camera sends the second raw image to the camera HAL module through the camera driver.
[0079] After receiving the second raw image, the camera HAL module processes the second raw image based on the operating code and data of the image processing method included.
[0080] S408: The camera HAL module processes the second raw image based on the first image parameters to obtain a detection stream image.
[0081] As Figure 5 shown, after the second raw image is captured by the image sensor, the detection path and the preview path start to work. It should be noted that the image signal processor is not shown in the figure, but the raw image is obtained based on the image sensor and the image signal processor.
[0082] In some embodiments, the camera HAL module includes modules such as a face detection module, a human body detection module, and an object detection module for detecting passages. The camera HAL module includes algorithms such as Algorithm 1 and Algorithm 2 for the preview passage. Exemplarily, Algorithm 1 and Algorithm 2 can be denoising algorithms, sharpening algorithms, etc., and the present application does not limit this.
[0083] In some embodiments, the camera HAL module includes the first image parameters of the detection stream image and the image parameters of the preview stream image. The camera HAL module performs image processing on the second raw image based on the first image parameters to obtain the detection stream image, and the camera HAL module performs image processing on the second raw image based on the image parameters of the preview stream image to obtain the preview stream image. Refer to Figure 5 。
[0084] It should be understood that the detection stream image is mainly used for perception detection to determine the object to be photographed and then determine the photographing scene, and it will not be presented to the user. Therefore, the resolution of the detection stream image is relatively low, that is, the resolution of the first image parameters is relatively low. It should be noted that the first image parameters are pre-stored in the camera HAL module, and they can be fixed or adaptively changed according to the image parameters of the second raw image, the external environment, etc. The present application does not limit this.
[0085] In some embodiments, the first image parameters include a resolution of 40×30 (i.e., the number of pixels) and a field of view angle of 65°. The resolution of the second raw image is 4000×3000, and the field of view angle is 130°. Then the mobile phone can crop the second raw image (i.e., perform field of view angle conversion) based on the field of view angle of 65° to obtain an image with a pixel count of 2000×1500, and then perform resolution conversion on it to obtain a detection stream image with a pixel count of 40×30.
[0086] Exemplarily, the field of view angle conversion can be centered on the center point of the second raw image and crop the periphery of the second raw image. It can also be centered on the upper left vertex of the second raw image and crop its right and lower sides. The present application does not limit this.
[0087] Exemplarily, the resolution conversion can be to proportionally convert the size of the image. For example, increase the resolution by proportionally enlarging the size of the image, or decrease the resolution by proportionally reducing the size of the image. The present application does not limit this.
[0088] Based on the above example, the center point of the second raw image can be used as the center point of the cropped image first. Vertically, 1000 pixels are cropped from the top and bottom of the second raw image; horizontally, 750 pixels are cropped from each side. Subsequently, the cropped image is proportionally reduced by 50 times to obtain a detection stream image with a field of view angle of 65° and a resolution of 40×30.
[0089] It should be understood that when the image ratio remains unchanged, the resolution conversion is to convert the clarity of the image, which will change the number of pixels included in the image, but will not change the content included in the image. The field of view conversion will change the photographing range of the camera, which will also change the number of pixels included in the image, but it will increase or decrease the content included in the image and will not change the clarity of the image. In some embodiments, adjusting the focal length and the image ratio will both cause a change in the field of view of the image. As in the following embodiments Figure 6a and Figure 6b As shown, adjusting the focal length and the photo ratio will both change the amount of content displayed in the image (which can also be called a photo).
[0090] It should be noted that the above image processing sequence of first performing field of view conversion on the second original image and then performing resolution conversion is only an example. It is also possible to first perform resolution conversion on the second original image and then perform field of view conversion, and the present application does not make any limitations in this regard.
[0091] In addition, it should be noted that the field of view included in the first image parameter may be the same as the field of view corresponding to the second original image, in which case only resolution conversion is required.
[0092] S409: The camera HAL module detects the detection stream image based on the face detection algorithm to obtain the first face data.
[0093] As Figure 5 shown, the detection path includes a face detection module, a human body detection module, an object detection module, etc. After the camera HAL module obtains the detection stream image, it can detect the detection stream image based on each detection module included in the detection path.
[0094] In some embodiments, each detection module includes a detection algorithm. Exemplarily, the camera HAL module can detect the detection stream image based on the face detection algorithm included in the face detection module to obtain the first face data (which can also be called the specific object detection result). For example, the face detection algorithm can include a face position detection algorithm, a face attribute detection algorithm, and a face feature point detection algorithm. Correspondingly, the first face data can include coordinate data corresponding to the face position, coordinate data corresponding to the face feature points, and face angle data, etc.
[0095] It should be understood that the resolution of the detection stream image is relatively low. Therefore, when using the face detection algorithm to perform face detection on it, the processing speed is relatively fast, and the first face data can be obtained quickly.
[0096] In some embodiments, the first face data may include face data related to one face or may include face data related to multiple faces. Exemplarily, the first face data includes Face Data 1, Face Data 2, and Face Data 3, which are face data related to one face respectively.
[0097] In some embodiments, after the camera is successfully started, the camera keeps working, and a series of second original images are captured. The detection path also keeps working. The camera HAL module processes each second original image based on the first image parameters to obtain a series of detection stream images, and also detects each detection stream image based on the face detection algorithm to obtain the first face data of each detection stream image. That is, during the operation of the camera application, the camera and the detection path keep working. After the camera captures a second original image, S407 - S409 are executed once.
[0098] S410: The camera application receives a trigger operation from the user for the camera application's photographing function.
[0099] After the user determines that the camera of the mobile phone is aligned with the object to be photographed through the preview image displayed on the mobile phone screen, the user can trigger the photographing function of the camera application.
[0100] In some embodiments, as Figure 1 shown, the trigger operation of the photographing function may be the user's click operation on the photographing control 12.
[0101] In some embodiments, the trigger operation of the photographing function may include the user's setting operation for the smiley face capture function of the camera application and the operation of the face in the preview image becoming a smiley face to trigger the smiley face capture function of the camera application. This application does not limit this.
[0102] S411: The camera application sends a photographing request and second image parameters to the camera HAL module through the camera access interface.
[0103] The second image parameters are the image parameters of the photographing stream images.
[0104] In some embodiments, the second image parameters may include image parameters such as focal length, resolution, and photo size (also referred to as image size). Exemplarily, the second image parameters of the photographing stream images may include a focal length of 2.0x, a photo size of 4:3, a resolution of 3072×4096, etc.
[0105] In some embodiments, the second image parameters may be obtained based on the default settings of the camera application.
[0106] In some embodiments, the second image parameters may also be obtained based on the user's configuration operation. This application does not limit this.
[0107] As Figure 6a and 6b shown, the mobile phone 100 displays a camera preview interface 610. The camera preview interface 610 includes controls such as a photo-taking control 611, a focal length selection control 612, and a settings control 613. The photo-taking control 611 is used to implement the function of taking photos, the focal length selection control 612 is used to implement the function of selecting the focal length, and the settings control 613 is used to implement the function of setting camera parameters. The camera preview interface 610 also includes a preview photo display area 614 and a thumbnail display area 615. In some embodiments, before the user triggers the photo-taking control 612, that is, before executing S410, the user can set image parameters (which can also be called photo parameters), such as setting the focal length and photo size, etc. The camera application receives the user's setting operation on the image parameters and generates second image parameters.
[0108] As Figure 6a shown, the user swipes the focal length selection control 612 included in the camera preview interface 610 to the right. In response to the user's right swipe operation, the mobile phone 100 adjusts the focal length from 1.0x to 2.0x. As the focal length increases, the photographing range of the camera decreases, and the content displayed in the camera preview interface 420 decreases.
[0109] As Figure 6b shown, the user clicks the settings control 613 included in the camera preview interface 610. In response to the user's click operation, the mobile phone 100 displays a camera parameter settings interface 620. The camera parameter settings interface 620 includes multiple controls such as a photo ratio settings control 621. The photo ratio settings control 621 is used to set the ratio of the photo. The parameter settings interface 620 also includes a return control 622, which is used to return to the camera preview interface 610.
[0110] The user triggers the photo ratio settings control 621 to select 4:3, and the mobile phone 100 can set it to 4:3. Subsequently, the user triggers the return control 622, and the mobile phone 100 returns to the camera preview interface 610. The photo ratio of the preview photo displayed in the preview photo display area 614 of the mobile phone 100 is adjusted from 1:1 to 4:3, and the content displayed in the preview photo also changes.
[0111] It should be noted that the above settings of the focal length and photo ratio are only examples. The user can set any image parameter or not set the image parameter. In this case, the mobile phone 100 takes photos using the default image parameters of the mobile phone 100, that is, the camera application generates second image parameters based on the default settings. This application does not make any limitations in this regard.
[0112] As Figure 6a and Figure 6bAs shown, changes in the photo ratio and focal length of the second image parameter will both affect the content displayed in the preview photo. Therefore, adjusting the photo ratio or focal length will cause the field of view angle to change.
[0113] S412: The camera HAL module processes the first raw image based on the second image parameter to obtain a captured stream image.
[0114] The camera HAL module adjusts the image parameter of the first raw image based on the second image parameter to obtain a captured stream image.
[0115] As Figure 5 shown, when the user triggers the photo-taking function of the camera application, that is, after the camera HAL module receives a photo-taking request and the second image parameter, the photo-taking path starts to work. In some embodiments, the camera HAL module includes multiple algorithms such as algorithm a, algorithm b, and algorithm c of the photo-taking path.
[0116] Based on the above example, before the user triggers the photo-taking function of the camera application, the camera HAL module can obtain multiple detection stream images and the first face data of each detection stream image.
[0117] In some embodiments, the photo-taking modes of the mobile phone include a single-frame photo-taking mode and a multi-frame fusion photo-taking mode. The single-frame photo-taking mode means that the photo shown to the user is obtained by the camera HAL module processing a single RAW image. The multi-frame fusion photo-taking mode means that the photo shown to the user is obtained by the camera HAL module first fusing multiple RAW images to obtain a single RAW image and then processing the fused single RAW image.
[0118] In a possible implementation, the photo-taking mode of the mobile phone is the single-frame photo-taking mode, and the first raw image is captured by the camera, that is, the first raw image and the second raw image are the same raw image.
[0119] Exemplarily, the first raw image can be captured by the camera before the user triggers the photo-taking function. Before the user triggers the photo-taking function, the camera HAL module can obtain the first face data.
[0120] Exemplarily again, the first raw image can be captured by the camera after the user triggers the photo-taking function. Then, after the user triggers the photo-taking function, S406 - S409 are executed again to obtain the first face data.
[0121] In a possible implementation, the photo-taking mode of the mobile phone is the multi-frame fusion photo-taking mode (which can be called HDR), and the first raw image is obtained by image fusion of multiple raw images including the second raw image, that is, the first raw image is different from the second raw image.
[0122] Exemplarily, the first original image is obtained by the camera HAL module performing image fusion processing on 5 continuously acquired original images, and one of the original images is the second original image. The number of original images in this application is not limited.
[0123] Taking the 5 original images including RAW1 - RAW5 as an example, RAW1 - RAW3 can be acquired by the front camera before the user triggers the camera function, and RAW4 - RAW5 can be acquired by the rear camera after the user triggers the camera function. This application does not limit this.
[0124] In some embodiments, the second image parameter includes a resolution of 4096×3072 (i.e., the number of pixels) and a field of view of 65° (obtained based on the focal length and the photo size). The resolution of the first original image is 4000×3000, and the field of view is 130°. Then the mobile phone can crop the first original image based on the field of view of 65° to obtain an image with a number of pixels of 2000×1500, and then scale up the cropped image by a factor of 2.048 to obtain a captured stream image with a field of view of 65° and a resolution of 4096×3072.
[0125] In addition, it should be noted that the field of view and resolution included in the second image parameter may be the same as those corresponding to the first original image. In this case, no resolution conversion and field of view conversion are required for the first original image.
[0126] For ease of understanding, in the embodiments of this application, one vertex of the image is used as the origin of the coordinate system, the upper boundary of the image is the horizontal axis of the coordinate system, the left boundary of the image is the vertical axis of the coordinate system, the positive direction of the horizontal axis of the coordinate system is to the right of the origin, and the positive direction of the vertical axis of the coordinate system is downward from the origin. For example, the vertex of the image can be the upper left vertex.
[0127] In some embodiments, the detected stream image corresponds to a first coordinate system, and its first vertex can be the origin; the captured stream image corresponds to a second coordinate system, and its second vertex can be the origin; the original image corresponds to a third coordinate system, and its third vertex can be the origin. It should be noted that the positions of the respective vertices in the corresponding images are the same, for example, they are all the upper left vertices.
[0128] It can be understood that the coordinate values of the coordinate points of an image will change with the change of the origin of the coordinate system and will also change with the magnification or reduction of the image. Performing a field of view conversion on the image may cause a change in the origin of the coordinate system. Performing a resolution conversion on the image will cause the image to be magnified or reduced. Therefore, both performing a field of view conversion and a resolution conversion on the image may cause the coordinate values of the coordinate points on the image to change.
[0129] In some embodiments, when performing face processing on the faces included in an image, it is necessary to first determine the face position and the positions of facial feature points based on the coordinate data of the face, and then perform relevant face processing such as adjusting skin color and enlarging eyes.
[0130] The first face data may include face data related to coordinate points, such as coordinate data of the face position and coordinate data of facial feature points, etc. These data will change with the changes in the field of view angle and resolution. The first face data may also include face data unrelated to coordinate points, such as face angle data, or face attribute data such as gender, age, pose, expression, etc. These data will not change due to the changes in the field of view angle and resolution. In the embodiments of the present application, the face data related to coordinate points may be referred to as first coordinate data, representing the coordinate data of the face in the detection stream image in its corresponding coordinate system.
[0131] Exemplarily, the coordinate data of the face position includes the coordinate points of the four vertices of a face frame (such as a rectangle), and there is a complete face within this rectangular frame. Taking the example that the portrait photo obtained by final photographing includes multiple faces, the coordinate data of the face position will include the coordinate points of the vertices corresponding to multiple face frames respectively, and there is a complete face in each face frame.
[0132] Exemplarily, the coordinate data of facial feature points includes multiple coordinate points for representing the shape and position of the feature points. Taking the left eye as an example, the coordinate points representing the shape and position of the eye include: the outer corner point of the left eye, which is a point located on the outer side of the left eye and is used to mark the position of the outer corner of the eye; the inner corner point of the left eye, which is a point located on the inner side of the left eye and is used to mark the position of the inner corner of the eye; the same applies to the right eye. Taking the nose as an example, the coordinate points representing the shape of the nose include: the left side point of the nose, which is a point located on the left side of the nose and is used to mark the shape of the nose; the right side point of the nose, which is a point located on the right side of the nose. The coordinate points representing the position of the nose include: the tip point of the nose, which is a point located at the tip of the nose. The coordinate data of facial feature points may also include the coordinate data of other feature points such as the mouth, eyebrows, chin, etc., which will not be elaborated here.
[0133] Based on the above examples, it can be known that the detection stream image is an image obtained by performing field of view angle conversion and resolution conversion on the second original image, and the photographing stream image is an image obtained by performing field of view angle conversion and resolution conversion on the first original image.
[0134] It should be understood that when the first original image is different from the second original image, the first original image is obtained by fusing multiple original images, and the multiple original images are collected at very close times. One of the multiple original images is the second original image. Therefore, the changes in the coordinate data corresponding to the first original image and the second original image can be ignored.
[0135] Exemplarily, one first face data can be randomly selected from the multiple first face data of the multiple detection stream images corresponding to these multiple RAW images as the first face data corresponding to the first original image.
[0136] Exemplarily again, the second original image can be used as a reference image, and the remaining RAW images can be used to perform image fusion processing on it. Then, the first face data corresponding to the second original image can be used as the first face data corresponding to the first original image.
[0137] Therefore, in order for the camera HAL module to determine the coordinate points related to the face on the capture stream image based on the first face data, that is, in order to obtain the coordinate points on the capture stream image corresponding to the coordinate points of the first face data, the first face data can be first mapped to the first original image to obtain the second face data of the first original image, and then the second face data can be mapped to the capture stream image to obtain the third face data (which can also be called the converted specific object detection result). In addition, in some embodiments, the field of view angle of the detection stream image is the same as that of the capture stream image, but their resolutions are different. In this case, the field of view angles of the detection stream image and the capture stream image contain the same content, and the difference lies in the size of the image. Therefore, the first face data can be directly mapped to the capture stream image based only on the difference in resolution to obtain the third face data.
[0138] It should be noted that during the mapping process, the first mapping is of the face data related to the coordinate points included in the first face data (which can be called the first coordinate data) to obtain the third coordinate data. The first coordinate data in the first face data can be replaced with the third coordinate data to obtain the second face data; the second mapping is of the face data related to the coordinate points included in the second face data (that is, the third coordinate data) to obtain the second coordinate data. The third coordinate data in the second face data can be replaced with the second coordinate data to obtain the third face data.
[0139] Exemplarily, the square root of the ratio of the resolution of the capture stream image to the resolution of the detection stream image can be used as a multiple C (which can be called the first multiple), and the coordinate values of the coordinate points included in the first face data are multiplied by the multiple C to obtain the third face data.
[0140] S413: The camera HAL module maps the first face data of the detection stream image to the first original image to obtain the second face data.
[0141] Based on the above example, the detected stream image is obtained by the camera HAL module performing image processing on the second original image. It is obtained by first performing a field of view angle conversion and then a resolution conversion. And based on the above example, the second original image can be equivalent to the first original image. Therefore, the camera HAL module can map the first face data to the first original image. It can be understood that the camera HAL module restores the detected stream image to the first original image to determine the change in the coordinate points of the first face data on the detected stream image on the first original image. Then, the camera HAL module needs to first perform a resolution conversion on the first face data and then a field of view angle conversion.
[0142] It should be noted that the detected stream image is obtained by the camera HAL module first performing a resolution conversion on the second original image and then a field of view angle conversion. Then, for restoration, the camera HAL module can also first perform a field of view angle conversion on the first face data and then a resolution conversion.
[0143] In some embodiments, when executing S408, the camera HAL module can store the corresponding first image processing information. Exemplarily, when cropping the second original image based on the field of view angle included in the first image parameters, the camera HAL module stores the coordinate points of each vertex of the detected stream image on the second original image, and the data can be referred to as the first vertex coordinate data. Taking the upper left corner vertex of the second original image as an example, the camera HAL module stores the coordinate point 1 of the upper left corner vertex of the cropped detected stream image on the second original image (which can be referred to as the second vertex coordinate data or the third vertex coordinate data). When performing a resolution conversion on the cropped image based on the resolution included in the first image parameters, the camera HAL module stores the reduction multiple A of the image in equal proportion (which can be referred to as the second multiple), that is, stores the ratio of the resolution of the cropped original image to the resolution of the first image parameters.
[0144] In some embodiments, the mapping process of the first face data may include the following steps 1-step 3:
[0145] Step 1: The camera HAL module obtains the multiple A and the coordinate point 1 included in the stored first image processing information, and obtains the coordinate point 2 included in the first face data.
[0146] The coordinate point 2 includes multiple coordinate points corresponding to the face position and the face feature point position included in the first face data respectively.
[0147] It should be noted that the coordinate point 1 can also be the coordinate point of other vertices of the detected stream image on the second original image, and this application does not make a limitation on this.
[0148] Step 2: The camera HAL module multiplies the coordinate value of the coordinate point 2 by the multiple A to obtain the coordinate point 3.
[0149] The detection stream image is enlarged proportionally based on multiple A, and coordinate point 3 is the coordinate point of multiple coordinate points included in the first face data in the enlarged detection stream image.
[0150] Step 3: The camera HAL module adds the coordinate value of coordinate point 3 to the coordinate value of coordinate point 1 to obtain coordinate point 4, and uses coordinate point 4 to replace coordinate point 2 of the first face data, and uses the replaced first face data as the second face data.
[0151] Coordinate point 4 is the coordinate point of multiple coordinate points included in the first face data in the first original image.
[0152] Exemplarily, the value of multiple A is u, and coordinate point 1 is (x1, y1). As Figure 7 shown, the first face data of the detection stream image includes face data 1, the upper left corner of the detection stream image is the origin (0, 0), taking the coordinate point (x2, y2) at the upper left corner of the face frame of face data 1 as an example (that is, coordinate point 2), then the camera HAL module needs to multiply the coordinate value of (x2, y2) by u to obtain coordinate point 3 (ux2, uy2), and then add the coordinate value of coordinate point 1 and the coordinate value of coordinate point 3 to obtain coordinate point 4 (x1 + ux2, y1 + uy2). The coordinate points of other vertices of the face frame and the multiple coordinate points corresponding to the face feature points are the same, which will not be elaborated here. Subsequently, the camera HAL module uses coordinate point 4 to replace coordinate point 1 of the first face data, and uses the face data including coordinate point 4 as the second face data.
[0153] S414: The camera HAL module maps the second face data to the capture stream image to obtain the third face data.
[0154] Based on the above example, the capture stream image is obtained by the camera HAL module performing image processing on the first original image. For example, the camera HAL module first performs field of view angle conversion on the first original image, and then performs resolution conversion to obtain the capture stream image. Therefore, mapping the second face data to the capture stream image can be understood as the camera HAL module performing image processing on the first original image to obtain the capture stream image to determine the coordinate point change of the coordinate point of the second face data on the first original image on the capture stream image.
[0155] In some embodiments, when performing S412, the camera HAL module may store corresponding second image processing information. Exemplarily, when cropping the first original image based on the field of view angle included in the second image parameters, the cropping coordinate points may be referred to as second coordinate points. Taking the upper left vertex of the first original image as the origin, the camera HAL module stores the coordinate points of the upper left vertex of the cropped captured stream image in the first original image (which may be referred to as the fifth vertex coordinate data or the sixth vertex coordinate data). When performing resolution conversion on the cropped image based on the resolution included in the second image parameters, the camera HAL module stores the magnification multiple B (which may be referred to as the third multiple) of the image's proportional change. The coordinate data of the vertices of the captured stream image in the first original image may be referred to as the fourth vertex coordinate data.
[0156] In some embodiments, the square root of the ratio of the resolution of the captured stream image to the resolution of the second original image may be used as the multiple B. When the multiple B is greater than 1, it indicates that the resolution of the captured stream image is larger and the image is enlarged proportionally; when the multiple B is less than 1, it indicates that the resolution of the captured stream image is smaller and the image is reduced proportionally.
[0157] In some embodiments, the mapping process of the second face data may include the following steps 4-step 6:
[0158] Step 4: The camera HAL module obtains the multiple B and the coordinate point 5 included in the stored second image processing information, and obtains the coordinate point 4 included in the second face data.
[0159] It should be noted that the coordinate point 5 may also be the coordinate points of other vertices of the captured stream image in the first original image, and the present application does not make any limitations in this regard.
[0160] Step 5: The camera HAL module subtracts the coordinate value of the coordinate point 5 from the coordinate value of the coordinate point 4 to obtain the coordinate point 6.
[0161] The coordinate point 6 is the coordinate point of the multiple coordinate points included in the second face data in the cropped first original image.
[0162] Step 6: The camera HAL module multiplies the coordinate value of the coordinate point 6 by the multiple B to obtain the coordinate point 7, replaces the coordinate point 4 of the second face data with the coordinate point 7, and uses the replaced second face data as the third face data.
[0163] The coordinate point 7 is the coordinate point of the multiple coordinate points included in the second face data in the captured stream image.
[0164] Exemplarily, the value of the multiple B is v, and the coordinate point 5 is (x3, y3). For example Figure 8As shown in the figure, the second face data of the first original image includes face data 1. The upper left corner of the first original image is the origin (0, 0). Taking the coordinate point (x4, y4) at the upper left corner of the face frame of face data 1 as an example (i.e., coordinate point 4), the camera HAL module needs to subtract the coordinate value of coordinate point 5 from the coordinate value of coordinate point 4 to obtain coordinate point 6 (x4 - x3, y4 - y3), and then multiply (x4 - x3, y4 - y3) by the multiple Bv to obtain coordinate point 7 (vx4 - vx3, vy4 - vy3). The coordinate points of other vertices of the face frame and the multiple coordinate points corresponding to the face feature points are the same and will not be elaborated here. In this way, coordinate point 7 is used to replace coordinate point 4 of the second face data, and the face data including coordinate point 7 is used as the third face data.
[0165] In addition, the field of view angles included in the first image parameter and the second image parameter may be different, resulting in differences in the content included in the detection stream image and the content included in the capture stream image.
[0166] For example, when the field of view angle corresponding to the detection stream image is larger than the field of view angle corresponding to the capture stream image, the content included in the detection stream image is more than the content included in the capture stream image. Therefore, the complete face included in the detection stream image may be a non-existent face or an incomplete face in the capture stream image, that is, there are invalid faces in the capture stream image. Then, it is necessary to screen out the invalid faces existing in the third face data.
[0167] As Figure 9 shown, the first face data of the detection stream image includes face data 1, face data 2, and face data 3. In the capture stream image, the face indicated by face data 2 is a non-existent face in the capture stream image, and the face indicated by face data 3 is an incomplete face in the capture stream image. Therefore, face data 2 and face data 3 are invalid face data of the capture stream image and need to be removed. The capture stream image only includes face data 1.
[0168] In some embodiments, it may also occur that there are faces in the detection stream image but the capture stream image does not contain a complete face. In this case, all the third face data are invalid face data, and then it is necessary to remove all the third face data. The capture path does not need to perform face-related processing. That is, when the camera preview interface of the camera application does not contain a complete face, after the user presses the capture control, the capture path of the electronic device does not need to perform face-related algorithms, such as not needing to perform Figure 5 the algorithms b and c shown in the figure.
[0169] In some embodiments, the detected stream image may include multiple human faces. Then the first human face data may include multiple human face data (also referred to as multiple groups of detection results) corresponding one by one to the multiple human faces. The converted third human face data also includes multiple human face data corresponding one by one to these human faces (after conversion). Assume that each human face data includes multiple coordinate points 7. For a piece of human face data, if there is a coordinate point 7 not on the captured stream image, it indicates that the human face data corresponding to this coordinate point 7 is invalid human face data. For example, all the multiple coordinate points 7 of the above human face data 2 are not on the captured stream image, and a part of the coordinate points 7 of the above human face data 3 (such as the coordinate point of the upper right vertex of the face frame, etc.) are on the captured stream image, while a part of the coordinate points 7 (such as the coordinate point of the upper left vertex of the face frame, etc.) are not on the captured stream image. Both are invalid human face data. In the embodiments of the present application, the coordinate point 7 not on the captured stream image may also be referred to as the first coordinate point.
[0170] In some embodiments, it is possible to determine whether the coordinate point 7 is on the captured stream image through the coordinate values of the coordinate points.
[0171] Exemplarily, if the coordinate value of a coordinate point 7 of the third human face data is ≥0, it indicates that this coordinate point 7 is within the captured stream image; if the coordinate value of a coordinate point 7 is <0, it indicates that this coordinate point 7 is outside the captured stream image. For each piece of human face data included in the third human face data, if there is a coordinate point 7 with a coordinate value <0, then this piece of human face data is determined as invalid human face data. Subsequently, the invalid human face data of the third human face data is removed, and the remaining human face data is used as the fourth human face data (also referred to as the converted and screened specific object detection result).
[0172] In this way, it is avoided that the fourth human face data contains invalid human face data, which may cause the mobile phone to perform face processing on a face that does not exist in the captured stream image, thereby avoiding an increase in the shot2review time.
[0173] S415: The camera HAL module performs face processing on the captured stream image based on the fourth human face data through a face processing algorithm to obtain a captured picture.
[0174] The captured picture is the image displayed on the mobile phone screen. In some embodiments, the fourth human face data includes data such as face attributes, face positions, feature point positions, and face angles, and is also referred to as the converted and screened specific object detection result. Based on the fourth human face data, the camera HAL module can perform face processing on the faces existing in the captured stream image through a face processing algorithm to obtain the image displayed to the user.
[0175] In some embodiments, such as Figure 5As shown in the figure, the photographing path includes algorithm a, algorithm b, and algorithm c. Algorithm a is an algorithm unrelated to the face processing algorithm, such as a denoising algorithm, a brightness adjustment algorithm, etc. Algorithm b and algorithm c are algorithms related to the face processing algorithm, such as beauty algorithms like the face skin color adjustment algorithm, the face feature point adjustment algorithm, etc.
[0176] As Figure 5 shown, in the image processing method provided by the embodiment of the present application, after the user triggers the photographing function of the camera application, based on the first face data obtained after the detection stream image is converted and screened, that is, the first face data is reused. The first face data is obtained by the camera HAL module detecting the detection stream image. The resolution of the detection stream image is small, the processing speed is fast, and the conversion speed of the first face data is also very fast, usually able to be completed in about 1 - 2 ms. Therefore, compared with the prior art, the acquisition speed of the fourth face data can be improved, and thus the shot2review time can be greatly saved.
[0177] S416: The camera HAL module sends the photographed picture to the camera application through the camera access interface.
[0178] S417: The camera application displays the thumbnail of the photographed picture.
[0179] As Figure 1 shown, the thumbnail of the photographed picture is displayed in the thumbnail display area 13. When the user triggers the thumbnail of the photographed picture, the mobile phone can display the photographed picture.
[0180] It should be noted that in the above embodiment, the first face data is reused as an example, and the present application is not limited thereto. Other data such as the main body data and human body data obtained by the detection path detecting the detection stream image can also be reused.
[0181] In addition, in some embodiments, in the video recording mode of the mobile phone, after the user triggers the video recording function of the mobile phone, the camera will collect RAW images one by one. For each RAW image, the mobile phone can reuse the first face data of the detection stream image corresponding to the RAW image to perform face processing on the photographing stream image corresponding to the RAW image to obtain a video frame displayed to the user. During the shooting process, the resolution of the detection stream image is lower than that of the photographing stream image. Therefore, the processing speed of face detection on the detection stream image is faster, which can reduce the time from the user clicking the shooting control to the thumbnail of the shot video being displayed to the user, thereby improving the user experience.
[0182] Next, the composition of the electronic device will be introduced.
[0183] It should be noted that the electronic device in the above embodiments being a mobile phone is only an example. In some embodiments, the electronic device may be a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), or other terminal devices. The present application does not impose any special restrictions on the specific form of the above-mentioned electronic device, as long as it can implement the photographing function.
[0184] Exemplarily, Figure 10 A schematic structural diagram of an electronic device 1000 is shown.
[0185] The electronic device 1000 may include a processor 1010, an internal memory 1020, a camera 1030, and a display screen 1040.
[0186] It can be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device. In other embodiments, the electronic device may include more or fewer components than those shown, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0187] The processor 1010 may include one or more processing units. For example, the processor 1010 may include an application processor (AP), a graphics processing unit (GPU), an image signal processor (ISP), a digital signal processor (DSP), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0188] The processor 1010 may also be provided with a memory for storing instructions and data.
[0189] The internal memory 1020 may be used to store computer-executable program code, and the executable program code includes instructions. The processor 1010 executes various functional applications and data processing of the electronic device 1000 by running the instructions stored in the internal memory 1020.
[0190] In some embodiments, the internal memory 1020 stores instructions for executing an image processing method. The processor 1010 can implement the image processing method provided in the embodiments of the present application by executing the instructions stored in the internal memory 1020.
[0191] The electronic device 1000 realizes a display function through an image processor, a display screen 1040, an application processor, etc. The image processor is a microprocessor for image processing, connected to the display screen 1040 and the application processor. The image processor is used to perform mathematical and geometric calculations for graphics rendering. The processor 1010 may include one or more image processors, which execute program instructions to generate or change display information. The display screen 1040 is used to display images, videos, etc.
[0192] In some embodiments, the display screen 1040 is used to display the operation interface of the camera application of the electronic device 1000, such as a camera preview interface, a camera parameter setting interface, etc.
[0193] The electronic device 1000 can realize a shooting function through an ISP, a camera 1030, a video codec, an image processor, a display screen 1040, an application processor, etc.
[0194] In some embodiments, the shooting function includes the photographing function and the video recording function of the camera application of the electronic device 1000.
[0195] The ISP is used to process the data fed back by the camera 1030. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera sensor. The optical signal is converted into an electrical signal, and the camera sensor transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene.
[0196] In some embodiments, the ISP can be set in the camera 1030. The camera 1030 is used to capture static images or videos. An object generates an optical image through the lens and projects it onto the sensor. The sensor can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The sensor converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV, etc. format. In some embodiments, the electronic device 1000 may include 1 or N cameras 1030, where N is a positive integer greater than 1.
[0197] In some embodiments, photosensitive elements such as CCD or CMOS are the image sensors mentioned in the embodiments of the present application, and electrical signals are obtained after the conversion of optical signals. The image sensor transmits the electrical signals to the ISP for processing to obtain a RAW image.
[0198] The embodiments of the present application also provide a computer-readable storage medium storing a computer program. When the computer program is executed by a computer, one or more steps in any of the above-mentioned image processing methods can be implemented.
[0199] The computer-readable storage medium may be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium may be ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage devices, etc.
[0200] Another embodiment of the present application also provides a computer program product containing instructions. When the computer program product is executed by a computer, one or more steps in any of the above-mentioned image processing methods can be implemented.
[0201] The electronic device, computer-readable storage medium, and computer program product provided in this embodiment are all used to execute the corresponding image processing method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding image processing method provided above, and will not be elaborated here.
[0202] The terms "first", "second", "third", etc. in the specification, claims, and drawings of the present application are used to distinguish different objects, rather than to limit a specific order.
[0203] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific way.
[0204] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An image processing method, characterized in that, Including: In response to the camera application being called, starting an image detection path of the camera application; the image detection path is used to detect a specific object in a detection stream image to obtain a specific object detection result; In response to the shooting function of the camera application being triggered, starting a shooting path of the camera application; The shooting path is used to process a shooting stream image to obtain a shooting picture; both the detection stream image and the shooting stream image are generated based on an original image obtained by the camera application; The resolution of the detection stream image is lower than the resolution of the shooting stream image; The shooting path uses the specific object detection result to perform display effect processing on the specific object in the shooting stream image to generate a shooting picture.
2. The method according to claim 1, wherein The detection stream image is obtained through the following steps: The image detection path reduces the resolution of the original image to obtain the detection stream image.
3. The method according to claim 1, wherein The first vertex of the detection stream image is the origin of the first coordinate system; the second vertex of the shooting stream image is the origin of the second coordinate system; the position of the first vertex in the detection stream image is the same as the position of the second vertex in the shooting stream image; the specific object detection result includes first coordinate data of a specific object in the detection stream image in the first coordinate system, and the shooting path uses the specific object detection result to perform display effect processing on the specific object in the shooting stream image to generate a shooting picture, including: The shooting path converts the first coordinate data included in the specific object detection result into second coordinate data in the second coordinate system; The shooting path uses the converted specific object detection result to perform display effect processing on the specific object in the shooting stream image to generate a shooting picture.
4. The method according to claim 3, characterized in that The method further includes: The shooting path screens out data that does not belong to the shooting stream image from the converted specific object detection result based on the second coordinate data; The shooting path uses the converted specific object detection result to perform display effect processing on the specific object in the shooting stream image to generate a shooting picture, including: The shooting path uses the converted and screened specific object detection result to perform display effect processing on the specific object in the shooting stream image to generate a shooting picture.
5. The method according to claim 4, characterized in that, The detection stream image includes multiple specific objects, and the converted specific object detection result includes multiple groups of detection results corresponding one by one to the multiple specific objects; the shooting path screens out data that does not belong to the shooting stream image from the converted specific object detection result based on the second coordinate data, including: The shooting path screens out first coordinate points that are not in the shooting stream image from the second coordinate data; The shooting path determines the target detection result to which the first coordinate point belongs from the multiple groups of detection results; The shooting path screens out the target detection result from the multiple groups of detection results.
6. The method according to claim 3, wherein The shooting path converts the first coordinate data included in the specific object detection result into second coordinate data in the second coordinate system, including: The shooting path determines that the field of view angle of the detection stream image is the same as the field of view angle of the shooting stream image; The photographing path calculates the square root of the ratio of the resolution of the photographed stream image to the resolution of the detected stream image as the first multiple; The photographing path multiplies the first multiple by the coordinate value of the first coordinate data to obtain the second coordinate data.
7. The method according to claim 3, characterized in that The third vertex of the original image is the origin of the third coordinate system; the position of the third vertex in the original image is the same as the position of the second vertex in the photographed stream image; The photographing path converts the first coordinate data included in the specific object detection result into second coordinate data in the second coordinate system, including: The photographing path determines that the field of view angle of the detected stream image is different from the field of view angle of the photographed stream image; The photographing path converts the first coordinate data into third coordinate data in the third coordinate system; The photographing path converts the third coordinate data into the second coordinate data.
8. The method according to claim 7, wherein The photographing path converts the first coordinate data into third coordinate data in the third coordinate system, including: The photographing path obtains first vertex coordinate data and a second multiple; the first vertex coordinate data is the data of the vertex of the detected stream image in the third coordinate system; the second multiple is the square root of the ratio of the resolution of the original image to the resolution of the detected stream image; The photographing path converts the first coordinate data into the third coordinate data based on the first vertex coordinate data and the second multiple.
9. The method according to claim 8, wherein The first vertex coordinate data includes second vertex coordinate data of the first vertex of the detected stream image in the third coordinate system; The photographing path converts the first coordinate data into the third coordinate data based on the first vertex coordinate data and the second multiple, including: The photographing path multiplies the coordinate value of the first coordinate data by the second multiple to obtain fourth coordinate data; The photographing path adds the coordinate value of the fourth coordinate data to the coordinate value of the second vertex coordinate data to obtain the third coordinate data.
10. The method according to claim 8, characterized in that The first vertex coordinate data includes third vertex coordinate data of the first vertex of the detected stream image in the third coordinate system; the photographing path converts the first coordinate data into the third coordinate data based on the first vertex coordinate data and the second multiple, including: The photographing path adds the coordinate value of the first coordinate data to the coordinate value of the third vertex coordinate data to obtain fifth coordinate data; The photographing path multiplies the coordinate value of the fifth coordinate data by the second multiple to obtain the third coordinate data.
11. The method according to claim 7, wherein The photographing path converts the third coordinate data into the second coordinate data, including: The photographing path obtains fourth vertex coordinate data and a third multiple; the fourth vertex coordinate data is the data of the vertex of the photographed stream image in the third coordinate system; the third multiple is the square root of the ratio of the resolution of the photographed stream image to the resolution of the original image; The photographing path converts the third coordinate data into the second coordinate data based on the fourth vertex coordinate data and the third multiple.
12. The method according to claim 11, wherein The fourth vertex coordinate data includes the fifth vertex coordinate data of the second vertex of the captured stream image in the third coordinate system; the capturing path converts the third coordinate data into the second coordinate data based on the fourth vertex coordinate data and the third multiple, including: The capturing path multiplies the coordinate values of the third coordinate data by the third multiple to obtain sixth coordinate data; The capturing path subtracts the coordinate values of the fifth vertex coordinate data from the coordinate values of the sixth coordinate data to obtain the second coordinate data.
13. The method according to claim 11, wherein The fourth vertex coordinate data includes the sixth vertex coordinate data of the second vertex of the captured stream image in the third coordinate system; the capturing path converts the third coordinate data into the second coordinate data based on the fourth vertex coordinate data and the third multiple, including: The capturing path subtracts the coordinate values of the sixth vertex coordinate data from the coordinate values of the third coordinate data to obtain seventh coordinate data; The capturing path multiplies the coordinate values of the seventh coordinate data by the third multiple to obtain the second coordinate data.
14. The method according to any one of claims 1 to 13, characterized in that The captured stream image is obtained through the following steps: The capturing path reduces the field of view angle of the original image to obtain the captured stream image.
15. The method according to claim 14, wherein The capturing path reduces the field of view angle of the original image to obtain the captured stream image, including: The capturing path obtains the zoom ratio of the camera application; The capturing path determines the second coordinate point of the original image based on the zoom ratio; The capturing path crops the original image based on the second coordinate point to obtain the captured stream image.
16. The method according to any one of claims 1-15, characterized in that, Both the detected stream image and the captured stream image are generated based on the first original image obtained by the camera application.
17. The method according to any one of claims 1-15, characterized in that, The captured stream image is generated based on the first original image obtained by the camera application, and the detected stream image is generated based on the second original image obtained by the camera application; the position of a specific object in the first original image is the same as the position of the specific object in the second original image.
18. The method according to claim 17, wherein The first original image is obtained through the following steps: Performing image fusion processing on multiple original images including the second original image to obtain the first original image.
19. The method according to any one of claims 1 to 18, characterized in that, The specific object includes a human face; the capturing path uses the specific object detection result to perform display effect processing on the specific object in the captured stream image to generate a captured picture, including: The capturing path uses the specific object detection result to perform image retouching on the human face in the captured stream image to generate a captured picture.
20. An electronic device, characterized in that, Including a memory and a processor; The memory is coupled to the processor, and the memory is used to store computer program code. The computer program code includes computer instructions, and one or more of the processors call the computer instructions to cause the electronic device to execute the image processing method according to any one of claims 1-19.
21. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, it implements the image processing method according to any one of claims 1-19.
Citation Information
Patent Citations
Internet-based face beautifying system
CN107392110A
Visual precise positioning control method based on industrial camera
CN113538565A
Low-resolution star map target identification method based on cyclic matching
CN116343056A
Image processor and image processing method
JP2008060844A
Image Processor and Image Processing Method
US20100027914A1