Image processing method and related device
By cutting pictures in real time and adjusting the shooting angle in the same interface of electronic devices, the problem of cushioning of pictures in the existing technology is solved, and an efficient cutting effect is achieved that meets the user's expectations.
Patent Information
- Application Number
- CN202410061466.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-15
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-01-15
AI Technical Summary
The cutting method of existing electronic devices requires users to perform tedious operations, resulting in inefficient cutting and seriously affecting the user experience.
The electronic device displays the input interface and the shooting preview interface in the same interface, displays the camera framing screen in real time, and automatically cuts out the cutout object. Users can adjust the cutout effect by adjusting the shooting angle and scene state until it meets expectations.
Simplify the cutout operation, improve the cutout efficiency, make the cutout image more in line with user expectations, and improve the user experience.
Smart Images

Figure CN120358408A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of terminals, and in particular to an image processing method and related devices. Background Art
[0002] Matting is one of the most common operations in image processing, which refers to separating a part of a picture or video from the original picture or video to form a separate layer. During the use of an electronic device, users often have a need to perform matting on images. For example, they may cut out an object in a captured image and put it into a note or upload it to a third-party application. However, the current matting methods supported by electronic devices require users to perform very cumbersome operations to cut out the image and place it in the target area, resulting in extremely low matting efficiency and seriously affecting the user experience.
[0003] Therefore, it is necessary to study a matting method with simpler operations. Summary of the Invention
[0004] The purpose of this application is to provide an image processing method and related devices. By implementing this method, an electronic device can simultaneously display an input interface and a shooting preview interface on the same interface. The shooting preview interface can display the picture obtained by the camera's real-time viewfinder. The electronic device can perform a matting operation on the matting object in this picture and pre-display the obtained matted image in the information input interface. This method simplifies the operations required by users for matting and can make the effect of the matted image more in line with the user's expectations, enhancing the user experience.
[0005] The above objectives and other objectives will be achieved by the features in the independent claims. Further implementation means are reflected in the dependent claims, the description, and the drawings.
[0006] In a first aspect, this application provides an image processing method applied to an electronic device, including: displaying a first interface, where the first interface includes a first edit box and a first control; in response to a first operation of the user on the first control, displaying a second interface, where the second interface includes a first area and a second area, the first area includes a first picture captured in real time by the electronic device's camera, the first picture includes a first object, the second area includes the first edit box, and the first edit box in the second area includes a second object, and the second object is obtained by matting the first object.
[0007] In this method, the first interface may be an application interface of applications such as a memo, a note, a notepad, WeChat, an office software, etc. The second area in the first interface may be used to receive the input of an image. Specifically, the input area of the first edit box may be the entire second area. When a user wants to input an image of an object existing in the environment into the first edit box, the user can perform the first operation on the first control through the method provided by this application to activate the camera of the electronic device, so that the electronic device displays the second interface. Specifically, the first operation may include a click operation.
[0008] After the second interface is displayed, the content displayed on the first interface may be reduced and displayed in the second area of the second interface. In addition, the user can aim the camera at a physical entity corresponding to the first object for shooting, and the obtained first picture can be displayed in real time in the first area. The electronic device will automatically extract the object corresponding to the physical entity in the picture (i.e., the first object), and display the extracted image, that is, the second object, in the second area.
[0009] When implementing this method, the electronic device can simultaneously display the first area as an input interface and the second area as a shooting preview interface on the first interface. The second area may display the picture obtained by the camera's real-time viewfinder. The electronic device can automatically perform a matte extraction operation on the matte extraction object in this picture, that is, the first object, and pre-display the obtained matte extraction image in the second area.
[0010] In a possible implementation manner, when the entity corresponding to the first object does not move out of the field of view angle range of the camera, the specific style (such as shape, size) of the second object presented in the second area may change with the change of the first object in the first area. For example, the user can change the imaging effect of the first object in the first area by adjusting the shooting angle of the device or adjusting the state of the object to be photographed in the shooting scene. After the imaging effect changes, the electronic device can re-perform matte extraction on the first object in the picture displayed on the first interface, and display the newly obtained matte extraction image in the second area until the matte extraction effect of the matte extraction image meets the user's expectation. In this way, the user only needs to perform simple operations to extract the image and place it in the area where they hope to place it, and can effectively ensure that the effect of the matte extraction image is more in line with the user's expectation.
[0011] In combination with the first aspect, in a possible implementation manner, after displaying the second interface, the method further includes: displaying, in the first area, a second picture captured in real time by the camera, the second picture including a first changing object that is different from the first object and both the first changing object and the first object correspond to a first entity; and displaying, in the first editing box, a third object that is obtained by changing the second object by a first proportional value.
[0012] Since the matting process imposes a certain burden on the computing power and power consumption of the electronic device, in this implementation manner, the electronic device can use a tracking algorithm (such as a tracking algorithm based on machine learning, such as single-object tracking or multi-object tracking based on a Siamese network) to determine whether the tracking of the matted object (i.e., the first object) is lost. Specifically, the electronic device can periodically obtain the picture in the viewfinder and determine whether the change amplitude of the first object in the viewfinder is within a certain range, specifically reflected in whether the similarity between the first object and the first changing object in the two pictures reaches a first threshold. When the similarity between the first object and the first changing object is greater than or equal to the first threshold, the electronic device can consider that the first object in the first picture and the first changing object in the second picture are actually the same object, which corresponds to the same physical entity, and the electronic device still maintains the tracking state of the original first object; and since these two objects are similar, the electronic device does not perform re-matting on the first changing object, but directly changes the scale (such as zooming) of the second object already displayed in the first editing box and then displays it in the second area.
[0013] In combination with the first aspect, in a possible implementation manner, the ratio of the area of the first changing object to the area of the first object is a second proportional value, and the first proportional value is positively correlated with the second proportional value.
[0014] It should be understood that even when the similarity between the first object and the first changing object is greater than or equal to the first threshold, the areas of these two objects may still be different. For example, the user can change the distance between the camera and the physical entity of the object to be cut out by adjusting the camera focal length or moving the camera, so as to change the area of the object to be cut out in the picture while keeping the specific shape of the object to be cut out in the picture unchanged. In this embodiment, in order to facilitate the user to adjust the size of the cut-out image in the second area by simply adjusting the focal length or moving the electronic device, when it is determined that the similarity between the first object and the first changing object is greater than or equal to the first threshold, the electronic device may determine the first ratio value according to the ratio of the area of the first changing object in the second picture to the area of the first object in the first picture. Specifically, the first ratio value may be the ratio of the electronic device to scale the first object.
[0015] Combined with the first aspect, in a possible implementation manner, after displaying the second interface, the method further includes: displaying a third picture captured in real time by the camera in the first area, the third picture includes a second changing object, the second changing object is different from the first changing object and the first object, and both the second changing object and the first object, the first changing object correspond to the first entity, and the first editing box in the second area includes a fourth object, and the fourth object is obtained by cutting out the second changing object.
[0016] Combined with the foregoing description, similarly, when the similarity between the first object and the second changing object is less than the first threshold, the electronic device may consider that the first object in the first picture and the second changing object in the third picture are not similar, and the tracking state of the original first object lost by the electronic device; and since these two objects are not similar, the electronic device will perform an operation of re-cutting the second changing object to obtain the fourth object, and replace the second object already displayed in the first editing box with the fourth object.
[0017] Combined with the first aspect, in a possible implementation manner, the first object corresponds to a first entity. After displaying the second interface, the method further includes: displaying a fourth picture captured in real time by the camera in the first area, the fourth picture includes a fifth object, the fifth object corresponds to a second entity, and the second entity is different from the first entity; replacing the second object in the first editing box in the second area with a sixth object, and the sixth object is obtained by cutting out the fifth object.
[0018] As can be seen from the foregoing description, if the form of the object to be cropped previously determined by the electronic device does not change significantly in the screen, the electronic device will not re-determine the object to be cropped. For example, when the user activates the camera through the corresponding user operation, the user does not place the target object in the central area of the shooting range of the camera. However, at this time, there is an interfering object in the central area of the shooting range, and the electronic device can automatically determine the object corresponding to the interfering object as the object to be cropped according to the method described in the foregoing description, and pre-display the cropped image corresponding to it in the editing box. However, after the user places the target object in the central area of the shooting range, according to the user's expectation, the electronic device should re-determine the object corresponding to the above target object in the viewfinder as the object to be cropped and re-crop the image. In this case, the user can only switch it to the new object to be cropped by manually clicking on the object corresponding to the above target object, and the operation becomes more cumbersome.
[0019] However, in this embodiment, the electronic device can detect in real time whether the entire screen content in the second area has changed significantly. The "change" mentioned here includes the change in the number of target objects in the screen and the change in the form of the target object in the screen. When the similarity between two screens in the second area, such as the first screen and the fourth screen, is less than the second threshold, the electronic device can determine that the screen content in the second area has changed significantly, and the electronic device can trigger the operation of re-determining the object to be cropped and re-cropping, and refresh the newly obtained cropped image, that is, the sixth object, into the second area, reducing the user operation while making the cropping effect further meet the user's expectation. That is to say, in this embodiment, in order to make the image cropped by the electronic device meet the user's expectation, the user can change the number of physical entities of the objects in the scene photographed by the camera, the relative positions of the physical entities, and the relative position between the physical entity and the camera, resulting in an increase or decrease in the number of objects in the screen photographed by the camera, or a change in the specific position of the objects in the screen, so that the electronic device can re-trigger the cropping operation.
[0020] Combined with the first aspect, in a possible implementation manner, the first screen further includes a seventh object, the seventh object corresponds to a third entity, and the third entity is different from the first entity; the distance between the center point of the first object and the center point of the first area is less than the distance between the center point of the seventh object and the center point of the first area.
[0021] In this embodiment, the first screen may include multiple objects, including the first object and the seventh object, and the first object is the object with the smallest distance from the center point of the first region among the multiple objects. It can be understood that generally, the center of the screen is the content that the user pays the most attention to, and the user generally places the object of interest at the center of the screen when taking a picture. Therefore, in order to determine which object among the multiple objects is the object for matting, the electronic device can determine the object for matting based on the distance between the center point of the first region and each object. Specifically, if the center point of the first region falls on an object (in this case, it can be considered that the distance between the center point of the first region and this object is 0), the electronic device can determine this object as the first object, that is, the object for matting; if the center point of the first region does not fall on any object, the electronic device can further compare the distance between the center point of the first region and the center points of each object, and determine the object with the smallest distance from the center point of the first region among these objects as the first object, that is, the object for matting.
[0022] Combined with the first aspect, in a possible implementation manner, the first screen further includes an eighth object, the eighth object corresponds to a fourth entity, the fourth entity is different from the first entity, and the area of the first object is larger than the area of the eighth object.
[0023] In this embodiment, the first screen may include multiple objects, including the first object and the seventh object, and the first object is the object with the smallest distance from the center point of the first region among the multiple objects. It can be understood that generally, the region where the object with the largest area in the screen is located is very likely to be the region of interest to the user. Therefore, the electronic device can further compare the areas of each object in the first screen, and determine the object with the largest area among these objects as the first object, that is, the object for matting.
[0024] Optionally, the electronic device may also comprehensively consider the distances between various objects in the first screen and the center point of the first area, as well as the areas of various objects in the first screen, to determine the object to be cut out from the multiple objects. Specifically, the electronic device may first compare the distances between various objects in the first screen and the center point of the first area. If the center point of the first area falls on a certain object, the electronic device may determine that object as the first object, that is, the object to be cut out. Otherwise, the electronic device may further compare the Manhattan distance between the center point of the first area and the center points of various objects. If the Manhattan distances between the center point of the first area and the center points of various objects are not significantly different, the electronic device may determine the object with the largest area among these objects as the first object. Otherwise, the electronic device may determine the object with the smallest distance from the center point of the first area among these objects as the first object, that is, the object to be cut out.
[0025] Combined with the first aspect, in a possible implementation manner, the first screen further includes a ninth object. After the second interface is displayed, the method further includes: in response to a second operation by the user on the ninth object, replacing the second object displayed in the first edit box with a tenth object, where the tenth object is obtained by cutting out the ninth object.
[0026] In some embodiments, if the object to be cut out automatically determined by the electronic device does not meet the user's expectations, the electronic device may respond to an interaction operation between the user and the screen, such as a click or double-click operation, to switch, add, or reduce the target object serving as the object to be cut out.
[0027] In this implementation manner, the first object is the first object automatically determined by the electronic device. If the object to be cut out actually expected by the user is the ninth object in the first screen, the user may, through a second operation on the ninth object in the first screen, determine the fourth object as the new object to be cut out. Correspondingly, the electronic device will cut out the fourth object in the first screen and refresh the obtained tenth object into the first edit box, that is, replace the second object and display it in the first edit box. Specifically, the second operation may be a click operation.
[0028] Combined with the first aspect, in a possible implementation manner, the first screen further includes an eleventh object. After the second interface is displayed, the method further includes: in response to a third operation by the user, displaying a twelfth object and the second object together in the first edit box, where the twelfth object is obtained by cutting out the eleventh object.
[0029] In this embodiment, the first object is the first object automatically determined by the electronic device. If the user actually expects the cutout objects to be the first object and the eleventh object in the first screen, the user can determine the first object and the fifth object as the cutout objects together through the third operation on the eleventh object in the first screen. Correspondingly, the electronic device will cut out the eleventh object in the first screen, and display the obtained twelfth object together with the second object in the first edit box. Specifically, the third operation can be a double-click operation.
[0030] In combination with the first aspect, in a possible implementation, the second object in the first edit box is a preview image. After the second interface is displayed, the method further includes: in response to a fourth operation on the first object in the first screen, replacing the second object displayed in the first edit box with a thirteenth object, the thirteenth object being obtained by cutting out the first object, and the display manner of the thirteenth object in the first edit box is different from the display manner of the second object in the first edit box.
[0031] In this embodiment, the second object displayed in the first edit box is a pre-display. The "pre-display" mentioned here means that the second object in the first edit box of the first interface is displayed. It is not the cutout image formally input into the second area by the user, but a preview image displayed in the second area by the electronic device for the convenience of the user to view the cutout effect of the cutout image. When the user determines that the cutout effect of the second object in the first interface meets the user's expectations, the user can formally input the image obtained by cutout into the second area through the first operation. Specifically, the fourth operation can be an operation of long pressing the first object in the first area and dragging it to the second area in the first interface.
[0032] In this embodiment, in order to prompt the user that the second object is a pre-displayed image, the second object in the first interface may be displayed differently from the thirteenth object. Specifically, the second object may be blurred or blurred, that is, in the second area, the clarity and / or transparency of the second object may be lower than the clarity and / or transparency of the thirteenth object, to prompt the user that the image is a preview image, and at the same time to prompt the user that the cutout image has not yet been formally added to the edit box. Optionally, the difference between the display of the second object and the thirteenth object in the second area may also include or be reflected in other aspects, such as the thickness of the lines, the overall grayscale value of the image, etc., which is not limited in this application.
[0033] In combination with the first aspect, in a possible implementation manner, before displaying the second picture captured in real time by the camera in the first area, the method further includes: receiving a fifth operation input by the user, where the fifth operation is an operation of adjusting the focal length of the camera or adjusting the distance between the camera and the first entity, so that the camera captures the second picture.
[0034] In combination with the first aspect, in a possible implementation manner, before displaying the third picture captured in real time by the camera in the first area, the method further includes: receiving a sixth operation input by the user: the sixth operation is an operation of changing the relative position between the camera and the first entity, so that the camera captures the third picture.
[0035] In a second aspect, the present application provides an image processing method, including: receiving a first instruction, where the first instruction is an instruction generated in response to a first operation of the user on a first control in a first interface, and the first instruction is used to start the camera of the electronic device, and the first interface includes a first edit box; displaying a second interface, where the second interface includes a first area and a second area, the first area includes the first picture captured in real time by the camera, the first picture includes a first object, and the second area includes the first edit box; using a matte algorithm to extract the first object from the first picture to obtain a second object; displaying the second object in the first edit box in the first area.
[0036] In combination with the second aspect, in a possible implementation manner, after displaying the second interface, the method further includes: displaying a second picture captured in real time by the camera in the first area, where the second picture includes a first changing object, the first changing object is different from the first object and both correspond to a first entity; comparing the similarity between the first object in the first picture and the first changing object in the second picture; when the similarity between the first object in the first picture and the first changing object in the second picture is greater than or equal to a first threshold, determining a first ratio value according to the ratio of the area of the first object in the first picture to the area of the first changing object in the second picture; scaling the second object by the first ratio value to obtain a third object; replacing the second object displayed in the first edit box with the third object.
[0037] In combination with the second aspect, in a possible implementation manner, after displaying the second interface, the method further includes: displaying, in the first region, a third picture captured in real time by the camera, where the third picture includes a second changing object; comparing the similarity between the first object in the first picture and the second changing object in the third picture; when the similarity between the first object in the first picture and the second changing object in the third picture is less than a first threshold, using a matting algorithm to extract the second changing object in the third picture to obtain a fourth object, and replacing the fourth object with the second object displayed in the first editing box.
[0038] In combination with the second aspect, in a possible implementation manner, after displaying the second interface, the method further includes: displaying, in the first region, a fourth picture captured in real time by the camera, where the fourth picture includes the first object and a fifth object; comparing the similarity between the first picture and the fourth picture; when the similarity between the first picture and the fourth picture is less than a second threshold, using a matting algorithm to extract the fifth object to obtain a sixth object, and replacing the sixth object with the second object displayed in the first editing box.
[0039] In combination with the second aspect, in a possible implementation manner, before receiving and displaying the second interface, the method further includes: detecting the first picture by using a target detection algorithm, and when at least one target object is detected in the first picture, determining the first object from the first picture, where the first object corresponds to a first category.
[0040] In some scenarios, the category of the object to be matted expected by the user may be relatively fixed. For example, the user may often perform matting on human portraits, or often perform matting on images of flowers, cats, dogs and other animals. Therefore, in this implementation manner, the target detection algorithm used by the electronic device can specify the category of the target object, and this category can include categories such as people, dogs, cats, flowers, etc., and can also include other categories, which can be specifically determined according to the user's matting preference and the specific usage scenario, and this application does not make any limitations. When the electronic device discriminates from the significant objects included in the image based on the target detection algorithm, if the electronic device detects that the category to which an object belongs does not belong to the category it has pre-specified based on the target detection algorithm, then this object will not be regarded as a target object, and therefore the image corresponding to this object cannot be extracted. In this way, it can effectively prevent non-user-expected objects from being extracted as target objects, thereby saving the ineffective operations of the electronic device.
[0041] In combination with the second aspect, in a possible implementation, the first screen includes multiple objects. Before displaying the second interface, the method further includes: when the center point of the first area exists on one of the multiple objects, determining the object where the center point of the first area is located as the first object; when the center point of the first area does not exist on any of the multiple objects, calculating the Manhattan distance between each object in the multiple objects and the center point, and sorting the Manhattan distances between each object in the multiple objects and the center point; when the difference between every two adjacent Manhattan distances is less than a preset threshold, determining the object with the largest area among the multiple objects as the first object; when the difference between every two adjacent Manhattan distances is not all less than the preset threshold, determining the object with the smallest Manhattan distance from the center point among the multiple objects as the first object.
[0042] In combination with the second aspect, in a possible implementation, using the matting algorithm to cut out the first object from the first screen includes: when the first category is a portrait, using a portrait matting algorithm to matte the first object; when the first category is not a portrait, using a non-portrait matting algorithm to matte the first object.
[0043] In this implementation, before matting, the electronic device can determine the category of the first object, which can be the category determined by the electronic device according to the image content corresponding to the matting object and the user's use of the matted image. Since the user generally has higher requirements for the matting effect of portraits than for other types of images, when the first object is a portrait, the electronic device can use a portrait matting algorithm (i.e., a hair-level matting algorithm) to matte the first object to ensure that the image effect of the first image is fine enough; when the first object is not a portrait, the electronic device can use a non-portrait matting algorithm to matte the first object to reduce the computing power burden and power consumption of the electronic device.
[0044] In combination with the second aspect, in a possible implementation, before comparing the similarity between the first object in the first screen and the second changed object in the second screen, the method further includes: extracting multiple frames of images from the images captured by the camera; sorting the multiple frames of images according to the capture time of the camera; and determining two consecutive adjacent frames or the first frame and the last frame among the multiple frames of images as the first screen and the second screen.
[0045] It can be understood that if the object to be cutout is in a state of slow but constant change in the picture displayed in the first area all the time, the similarity between the cutout objects in two adjacent pictures may always be less than the first threshold. Then, even after a long period of time when the change range of the cutout object in the picture has accumulated to a large enough extent, that is, the form of the cutout object determined by the electronic device from the current picture is significantly different from the second object, the electronic device may still not re-cutout the cutout object in the picture and refresh the newly obtained cutout image to the second area.
[0046] Therefore, in this embodiment, after obtaining the multiple frames of pictures captured by the camera, in addition to comparing the similarity of the cutout objects included in every two consecutive adjacent pictures, the electronic device also takes every N consecutive frames of images as a group and compares the similarity of the cutout objects in the first frame and the last frame of every N consecutive pictures. In this way, in the case where the cutout object suddenly changes greatly or changes slowly continuously in the viewfinder, the electronic device can re-cutout the cutout object in the picture in time when the change degree reaches a certain level, and refresh the newly obtained cutout image to the first editing frame for the user to view.
[0047] For example, the first picture and the second picture can be two consecutive pictures among the multiple frames of images, and during the extremely short time period of collecting the first picture and the second picture, the first entity corresponding to the first object and the second changing object changes greatly rapidly in the shooting scene (such as a huge change in the relative position with the camera), resulting in the dissimilarity between the first object in the first picture and the second changing object in the second picture. Then, at this time, the electronic device can re-cutout the second changing object and refresh the cutout image to the first editing frame. In addition, if the first picture and the second picture can be the first frame and the last frame among the consecutive N frames of images included in the multiple frames of images, and during the relatively long time period of collecting the first picture and the second picture, the first entity corresponding to the first object and the second changing object changes extremely slowly in the shooting scene, but due to the long-term accumulation, the dissimilarity between the first object in the first picture and the second changing object in the second picture is caused. Then, at this time, the electronic device can also re-cutout the second changing object and refresh the cutout image to the first editing frame.
[0048] In combination with the second aspect, in a possible implementation manner, before comparing the similarity between the first picture and the fourth picture, the method further includes: extracting multiple frames of pictures from the pictures captured by the camera; sorting the multiple frames of pictures according to the capture time of the camera; and determining two consecutively adjacent frames or the first frame and the last frame in the multiple frames of pictures as the first picture and the fourth picture.
[0049] Similarly, in this implementation manner, after obtaining the multiple frames of pictures captured by the camera, in addition to comparing the similarity of the picture content of every two consecutively adjacent frames, the electronic device also groups every consecutive M frames of images as a group and compares the similarity of the picture content between the first frame and the last frame in every M consecutive frames. In this way, in the case where the picture content in the viewfinder suddenly changes greatly or changes slowly continuously, the electronic device can re-matte the object to be matted in the picture in a timely manner when the degree of change reaches a certain level, and refresh the newly obtained matted image into the first editing frame for the user to view.
[0050] In a third aspect, an embodiment of the present application provides an electronic device, where the electronic device includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the electronic device to execute the method in the first aspect or any possible implementation manner of the first aspect, or the method in the second aspect or any possible implementation manner of the second aspect.
[0051] In a fourth aspect, a chip system is provided, where the chip system is applied to an electronic device, and the chip system includes one or more processors, and the processors are used to call computer instructions to enable the electronic device to execute the method in the first aspect or any possible implementation manner of the first aspect, or the method in the second aspect or any possible implementation manner of the second aspect.
[0052] In a fifth aspect, a computer-readable storage medium is provided, including instructions, and when the above instructions run on an electronic device, the above electronic device is enabled to execute the method in the first aspect or any possible implementation manner of the first aspect, or the method in the second aspect or any possible implementation manner of the second aspect.
[0053] In a sixth aspect, a computer program product including instructions, when the above computer program product runs on an electronic device, enables the above electronic device to execute the method in the first aspect or any possible implementation manner of the first aspect, or the method in the second aspect or any possible implementation manner of the second aspect. Description of the Drawings
[0054] Figure 1 Schematic diagram of Euler angles provided by an embodiment of the present application;
[0055] Figure 2 Process schematic diagram of a matte extraction method provided by an embodiment of the present application;
[0056] Figure 3 Process schematic diagram of a matte extraction method provided by an embodiment of the present application;
[0057] Figure 4 Flowchart of a matte extraction method provided by an embodiment of the present application;
[0058] Figure 5 User interface diagram for outputting a first prompt message provided by an embodiment of the present application;
[0059] Figure 6 User interface diagram for determining a matte extraction object of a specified type provided by an embodiment of the present application;
[0060] Figure 7 User interface diagram for outputting a second prompt message provided by an embodiment of the present application;
[0061] Figure 8 Flowchart of a method for determining a matte extraction object provided by an embodiment of the present application;
[0062] Figure 9 User interface diagram for displaying the positional relationship between the center point of the viewfinder frame and the matte extraction object provided by an embodiment of the present application;
[0063] Figure 10 Process schematic diagram for determining a matte extraction object provided by an embodiment of the present application;
[0064] Figure 11 User interface diagram for switching the matte extraction object provided by an embodiment of the present application;
[0065] Figure 12 User interface diagram for switching the matte extraction object provided by an embodiment of the present application;
[0066] Figure 13 User interface diagram for switching the matte extraction object provided by an embodiment of the present application;
[0067] Figure 14 User interface diagram for one - key selection of matte extraction objects of the same type provided by an embodiment of the present application;
[0068] Figure 15 Flowchart of a method for determining a matte extraction strategy provided by an embodiment of the present application;
[0069] Figure 16A user interface diagram provided by an embodiment of the present application to show the effects of different matte extraction strategies;
[0070] Figure 17 A flowchart of a method provided by an embodiment of the present application for determining whether to perform matte extraction again;
[0071] Figure 18 A schematic diagram of the process of obtaining multiple frames of images and performing similarity comparison provided by an embodiment of the present application;
[0072] Figure 19 A schematic diagram of a scenario for comparing the similarity of matte extraction objects provided by an embodiment of the present application;
[0073] Figure 20 A schematic diagram of a scenario for comparing the similarity of matte extraction objects provided by an embodiment of the present application;
[0074] Figure 21 A schematic diagram of a scenario for comparing the similarity of matte extraction objects provided by an embodiment of the present application;
[0075] Figure 22 A flowchart of a method provided by an embodiment of the present application for determining whether to perform matte extraction again;
[0076] Figure 23 A user interface diagram provided by an embodiment of the present application to show the change of the content in the viewfinder frame;
[0077] Figure 24 A user interface diagram provided by an embodiment of the present application to show the change of the content in the viewfinder frame;
[0078] Figure 25 A structural diagram of an electronic device provided by an embodiment of the present application;
[0079] Figure 26 A software structure block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0080] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. As used in the specification and claims of the present application, the singular forms "a", "an", "the", "above", "said", "this" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term " / and / " used in the present application refers to and includes any or all possible combinations of one or more of the listed items.
[0081] For ease of understanding, the following first introduces the relevant terms involved in the embodiments of the present application.
[0082] (1) Multi-window interface
[0083] A multi-window interface is a user interface that can be displayed on an electronic device, which is convenient for browsing applications in the form of floating windows or split screens, facilitating parallel multitasking and being more flexible and efficient. In a multi-window interface, the display screen can be divided into two display areas, and both of these two display areas can be used to display the application interfaces of two different applications. These two display areas can be spliced and displayed in the multi-window interface, that is, the upper and lower display interfaces do not overlap, and the areas of the upper display area and the lower display area can be the same or different. Alternatively, these two display areas can also be displayed in the multi-window interface in the form of a floating window (picture-in-picture), that is, the display screen of the device can include two display areas, one large and one small, and the smaller display interface is included in the larger display interface; among them, the larger display area generally fills the screen of the device, and the image in the smaller display area can cover the image in the larger display area.
[0084] For example, in the present application, the multi-window interface displayed on the electronic device may include a first area and a second area. Among them:
[0085] The first area can display the interface of the camera application, which can display the picture captured by the camera on the electronic device. It should be noted that the so-called "shooting" here does not refer to the operation of the user pressing the shutter of the camera, but refers to the operation that the camera converts the optical image into an electrical signal in real time, then transmits the electrical signal to the image processor for processing, and finally generates the viewfinder picture for the user to view on the display screen; that is to say, when the camera changes the focal length, shooting direction and angle under the user's operation or the objects within the shooting range change (such as the change in the number of objects or the change in the position of the objects), the picture content displayed in the first area will also change accordingly. Optionally, the picture displayed in the first area can be obtained by any one of the cameras on the electronic device, such as obtained by the front camera or the rear camera of the electronic device, and the present application does not make any limitation in this regard.
[0086] The second area can be an application interface that supports image input. Before the electronic device displays the multi-window interface including the above-mentioned first area and second area, the application interface displayed in the second area can be full-screen displayed on the display screen of the electronic device. The electronic device can perform user operations on this application interface to activate the camera application of the electronic device and make the electronic device display the multi-window interface including the above-mentioned first area and second area.
[0087] In this application, based on the multi-window interface displayed on the electronic device, when the picture displayed in the first area contains a cutout object (i.e., the object to be cut out from the picture), the electronic device can automatically perform a cutout operation on the cutout object, that is, cut out a layer containing only the cutout object from the picture displayed in the first area, and pre-display a preview effect picture of the cutout image (i.e., the aforementioned layer or image containing only the cutout object, the same below) in the second area. The preview effect picture can change as the size and shape of the cutout object in the first area change. After the cutout effect presented in the preview effect picture meets the user's expectations, the user can display the cutout image of the cutout object in the second area through corresponding operations. Optionally, when the picture displayed in the first area contains multiple objects, the user can switch, add, or reduce the cutout objects through corresponding operations, which will be described in more detail later and will not be elaborated here for now.
[0088] Specifically, the first area and the second area can be displayed in the multi-window interface in the form of floating windows or split screens. For the convenience of description, the multi-window interface displayed in the up-and-down split screen form will be used for explanation later. In addition, terms such as "multi-window interface", "floating window", "split screen display", and "cutout image" are only some names used in the embodiments of this application, and their meanings have been recorded in the embodiments of this application, and their names do not constitute any limitation to this embodiment.
[0089] (2) Target object, cutout object
[0090] The target object can be an object determined by the electronic device from an image based on a target detection algorithm. The target detection algorithm can detect whether there is a target object, the category of the target object, and the position of the target object in the image from the image. For the same image, the target objects it contains may be one or multiple; the target objects it contains may belong to one category or multiple different categories. It should be noted that in some embodiments, the target detection algorithm can pre-specify a finite number of categories to which the target object belongs. These categories can include categories such as people, dogs, cats, flowers, etc., and can also include other categories. When the electronic device discriminates from the significant objects contained in the image based on the target detection algorithm, if the electronic device detects that an object belongs to a category that is not pre-specified by it based on the target detection algorithm, then the object will not be regarded as a target object.
[0091] In this application, the target object can also be referred to as the "salient object". Generally speaking, whether an object is a target object is related to its area ratio in the image, its positional relationship with the center point of the image, and the overall color of the object. Among them, the larger the area ratio, the closer to the center point of the image, and the smaller the color variance, the more likely the object is the salient object in the image. It should be noted that there may be multiple target objects in an image, or there may be no target object at all. For example, for an image taken in a complex environment, the electronic device may be able to detect multiple target objects from the multiple objects included in the image; while for a solid-color image (such as an image taken of a pure blue sky), the electronic device may not be able to detect any target object from the image.
[0092] In this application, the electronic device can automatically determine a target object in the image as the object to be cut out, or the electronic device can also respond to the user's operation on one (or more) target objects in the image (such as a click or touch operation), and determine the one (or more) target objects as the object to be cut out. After that, the electronic device can use the corresponding cut-out algorithm to perform a cut-out operation on the object to be cut out in the original image, that is, the image corresponding to the target object in the original image is separated from the original image to become a separate layer, and this layer is the cut-out image.
[0093] (3) Mask
[0094] The mask can also be referred to as a "masking film". In the fields of computer vision and image processing, a mask is an image used to specify a specific area in an image. Optionally, the electronic device can set the area where the object to be cut out in the image is white (pixel value is 255), and set other areas to black (pixel value is 0), so as to obtain a layer that has an occlusion effect, and this layer is the mask. Simply put, during the cut-out process, the mask can be overlaid on the original image. That is, the white area with a gray value of 255 can be understood as a "light-transmitting area", and through this area, the elements in the corresponding area of the original image (that is, the object to be cut out in the cut-out) can be seen. And the black area with a gray value of 0 has an occlusion effect on the elements in the corresponding area of the original image, so that the elements in the corresponding area of the original image cannot be seen. In this way, by occluding the original image with the mask, the cut-out image of the object to be cut out can be obtained. Of course, in order to pursue a more refined cut-out effect, the electronic device can also perform more refined processing on the edge part of the cut-out image, which will be described later and will not be elaborated here.
[0095] In this application, there may be an overlapping part between two or more target objects in an image (which appears as two target objects being connected together in the image). In response to this situation, after the user selects one of the target objects, the electronic device may only determine the selected target object as the object to be matte cut, and can also recognize the outer contour of the target object during matte cutting to generate a mask only for cutting out the target object; optionally, the electronic device may also determine the selected target object and an entire entity formed by all other target objects overlapping with the target object as the object to be matte cut, and can also recognize the outer contour of the entire entity during matte cutting to generate a mask for completely cutting out the above-mentioned entire entity, which can be specifically determined according to the implementation logic of the matte cutting algorithm. That is to say, in this application, the object to be matte cut can be any target object that can be segmented in the image, and the present invention does not make any limitations.
[0096] (4) Tracking box
[0097] After the camera application in the electronic device is launched, the screen of the electronic device will display the picture obtained by the camera for shooting for the user to preview. In the preview picture, the electronic device usually tracks the target objects (such as human faces, cats, and dogs) existing in the picture, and displays corresponding tracking boxes for each target object in the preview picture. The tracking box can be in the form of a rectangular box or other styles, and the present application does not make any limitations in this regard. When the electronic device (camera) moves, the size of the target object in the preview picture may also change accordingly, and the electronic device will also automatically adjust the size of the corresponding tracking box according to the size of the target object, so that the user can shoot the target more accurately.
[0098] In this application, all target objects in the viewfinder picture can be framed by their corresponding tracking boxes. Optionally, the specific manifestation form of the tracking box of the object to be matte cut and the tracking boxes of other target objects (target objects other than the object to be matte cut) in the picture can be different. For example, the color of the tracking box of the object to be matte cut can be different from the color of the tracking boxes of other target objects; or, the line of the tracking box of the object to be matte cut can be thicker than the lines of the tracking boxes of other target objects; or, the electronic device can only display the tracking box of the object to be matte cut in the interface, and the tracking boxes of other target objects can not be displayed in the interface. Of course, the tracking box of the object to be matte cut and the tracking boxes of other target objects can also be distinguished in other ways, and the present application does not make any limitations in this regard.
[0099] In addition, in some embodiments, the tracking box of the object to be matte cut and the tracking boxes of other target objects can all exist in a form that can only be perceived by the electronic device, and they will not be displayed in the preview picture in a form visible to the user.
[0100] (5) Euler angle
[0101] Euler angles are a way to represent rotation. Formally, it is a three-dimensional vector, and its values respectively represent the rotation angles of an object around the three axes (x-axis, y-axis, z-axis) of the coordinate system.
[0102] As Figure 1 shown, assuming that the plane where the electronic device is located is the xoy plane, the angle of rotation of the electronic device around the x-axis can be called the pitch angle, the angle of rotation of the electronic device around the y-axis can be called the yaw angle, and the angle of rotation of the electronic device around the z-axis can be called the roll angle.
[0103] When the object being photographed does not move at the same frequency as the electronic device, the movement of the electronic device around any spatial axis during shooting may cause changes in the position, shape, or size of the object being photographed in the preview screen, and may even cause the original object being photographed to move out of the field of view of the camera, and / or a new object being photographed to enter the field of view of the camera, resulting in changes in the number and type of target objects in the preview screen.
[0104] (6) TOF camera
[0105] The TOF camera calculates the distance between the object and the lens by emitting a laser beam or an LED flash and detecting the echo time, so as to achieve three-dimensional perception of the environment. The TOF camera can achieve fast and high-precision spatial measurement and three-dimensional imaging. For example, in this application, when the electronic device can measure the distance between the electronic device (camera) and the object being photographed through the TOF camera.
[0106] "Matting" is one of the most common operations in image processing, that is, separating a part of a picture or video (the matting object) from the original picture or video to become a separate layer (the matting image), which can be used as information required by the user in an electronic note later, or for making promotional posters, etc. For example, in some application scenarios, users often perform a matting operation on a certain object in the photo they take and put the obtained matting image into an electronic note as part of the note to make their notes more vivid.
[0107] Figure 2 shows the specific process of the electronic device performing a matting operation on the matting object in the image through the currently supported matting method and adding the obtained matting image to the application "Note".
[0108] As Figure 2As shown in (A) therein, the electronic device displays the user interface 21. The user interface 21 is the application interface of the camera application in the electronic device, and it can be displayed after the electronic device responds to a click operation on the icon of the camera application. The user interface 21 may include a viewfinder area 211, a shooting control 212, an echo control 213 for the captured image, and a camera switching control 214. Among them:
[0109] The viewfinder area 211 is used to display a real-time preview of the image obtained by the camera currently taking pictures of the shooting scene. At this time, the electronic device is taking a real-time picture of an object in the scene. Therefore, the image 211A currently displayed in the viewfinder area 211 is the image corresponding to the object. It should be noted that the image corresponding to the object is the image that the user wants to cut out and put into the note. Therefore, it is necessary to take a picture of the object at this time.
[0110] The shooting control 212 is used to respond to the user operation to shoot and save the taken photo.
[0111] The echo control 213 for the captured image is used for the user to view the taken pictures or videos.
[0112] The camera switching control 214 is used to switch the camera for collecting images between the front camera and the rear camera.
[0113] In response to the touch operation of the user on the shooting control 212 as shown in (A) therein, the electronic device can take a picture, save the obtained photo to the gallery of the electronic device, and display the user interface 22 as shown in (B) therein. Correspondingly, a thumbnail of the photo can be displayed in the echo control 213 for the captured image. In some embodiments, after the electronic device takes a picture, the electronic device can also continue to display the user interface 21. After responding to the touch operation of the user on the echo control 213 for the captured image, the electronic device will display the user interface 22 as shown in (B) therein. Figure 2 Figure 2 Figure 2 Figure 2 Figure 2
[0114] As Figure 2 As shown in (B) therein, the user interface 22 can be an image browsing interface of the camera application, which may include the image 221 of the user's historical taken photos, that is, the photo obtained by the electronic device taking pictures of the above-mentioned object after responding to the touch operation of the user on the shooting control 212. The image 221 may include the image 221A corresponding to the above-mentioned object. The electronic device can respond to the user operation on the image 221A, such as Figure 2The long press operation shown in (B) in the figure cuts out image 221A from image 221 to obtain cutout image 221B. If the user does not let go, cutout image 221B can move with the user's fingertips in user interface 22, and cutout image 221B can be displayed on top in user interface 22 (i.e., cutout image 221B can cover any element in user interface 22). It can be understood that the image content displayed by image 221B is the same as the image content displayed by image 221A, but the two can be the same or different in area size, and this application does not limit this.
[0115] Optionally, after image 221B is cut out, the electronic device may display image 221 in the user interface 22 and blur or defocus image 221 , that is, the clarity of image 221 in the user interface 22 may be reduced to prompt the user that the device has completed the cutout operation on image 221 .
[0116] Next, also in the user interface 22, the user can use a multi-finger operation, that is, while pressing the cutout image 221B, slide upward at the lower boundary of the electronic device screen; the electronic device can respond to this operation and display the following Figure 2 User interface 23 shown in (C) in FIG. User interface 23 is an application navigation interface, in which thumbnails of application interfaces of all applications currently opened on the electronic device and a cutout image 221B obtained by the user performing a cutout operation in user interface 22 can be displayed. Among them, application interface thumbnail 231 is a thumbnail of the image browsing interface (i.e., user interface 22) in the application "Camera", and application interface thumbnail 232 is a thumbnail of the note editing interface in the application "Notes". In user interface 23, the user can move the cutout image 221B onto the application interface thumbnail 232 and then release it to place the cutout image 221B into the note he is editing in the note application; and the electronic device can respond to this operation and switch the note application corresponding to the application interface thumbnail 232 to the foreground application, that is, display the following: Figure 2 The user interface 24 shown in (D) in FIG.
[0117] like Figure 2 As shown in (D) in FIG. 1 , the user interface 24 is a note editing interface of the "note" application in the electronic device. It may include an input area 241. The input area 241 may receive text, images, and other information input by the user. Figure 2 As can be seen from (D) in FIG. 2 , at this time, the input area 241 already displays the cutout image 221B that the user previously cut out from other interfaces.
[0118] It should be noted here that in the user interface 22, when the user extracts the image 221A from the image 221 through a long - press operation to obtain the cut - out image 221B, the electronic device can set the size of the cut - out image 221B to be the same as that of the image 221A; or, the electronic device can scale the image 221A after extracting it, that is, enlarge or reduce the extracted image while ensuring the image content remains unchanged to obtain the cut - out image 221B. Similarly, in the user interface 24, when the electronic device places the cut - out image 221B into the input area 241, it can also scale the cut - out image 221B; that is to say, the area size of the cut - out image 221B in the user interface 23 and the cut - out image 221B in the user interface 24 can be the same or different.
[0119] From Figure 2 As can be seen from the specific operation process of the cut - out method shown, this method requires multiple operation steps when executed, and involves the jump of multiple applications and pages. The interaction efficiency between the user and the electronic device is very low, seriously affecting the user experience. In addition, the shooting effect of the electronic device on the photographed object will further affect the quality of the subsequent cut - out image. If the effect of the cut - out image does not meet the user's expectations due to poor shooting quality (for example, the shape of the cut - out image does not meet the user's expectations due to the shooting angle, or the edge of the cut - out image is not clear due to the influence of light intensity during shooting, etc.), then the user needs to start from taking pictures again and execute the cut - out process as shown in Figure 2 shown again, and the overall operation is cumbersome.
[0120] In view of the defects existing in the above - mentioned method, the present application provides an image - processing method and related device. When implementing this method, the electronic device can simultaneously display the input interface and the shooting preview interface on the same interface. Among them, the shooting preview interface can display the picture obtained by the camera's real - time viewfinder. The electronic device can perform a cut - out operation on the cut - out object in this picture and pre - display the obtained cut - out image in the input interface. The user can change the cut - out effect (including image size, image shape, etc.) of the cut - out image pre - displayed in the input interface by adjusting the shooting angle of the device or adjusting the state of the photographed object in the shooting scene. When the cut - out effect of the cut - out image pre - displayed in the input interface meets the user's expectations, the user can formally place the cut - out image into the input interface through corresponding user operations. In this way, the user only needs to perform simple operations to extract the image and place it into the area where they hope to place it. And because the cut - out image will be pre - displayed in the input area in real - time, it can effectively avoid the situation where the effect of the cut - out image does not meet the user's expectations due to poor shooting quality after taking pictures.
[0121] Figure 3It shows the specific process in which an electronic device performs a matting operation on a matting object in the image displayed in the viewfinder through the image processing method provided in this application, and adds the obtained matted image to the "Notes" application.
[0122] As Figure 3 shown in (A) of, the user interface 31 is the note editing interface of the "Notes" application in the electronic device. It may include an editing box 311 and a function list 312. Among them:
[0123] The editing box 311 can receive information such as text and images input by the user. If the user wants to capture an object in the environment at this time and extract the image of the object and put it into the above-mentioned editing box 311, the user can perform a user operation on the blank part of the above input area, such as Figure 3 the long-press operation shown in (A) of, so that an option list 311A is displayed in the user interface 31. The option list 311A may include a paste option, a select all option, and a capture input option. Among them, the paste option is used to paste the text information copied by the user into the editing box 311; the select all option is used to select all the information in the editing box 311 with one click; the capture input option is used to call the camera of the electronic device to capture, and display the captured picture and the input area on the same interface. Any one of the options can be used to respond to the user's operation, such as a click operation, so that the electronic device performs the operation corresponding to the option.
[0124] The function list 312 may include a share option, a favorite option, a delete option, and a more option. Any one of the options can respond to the user's click operation, so that the electronic device performs the function corresponding to the option for the current note.
[0125] As Figure 3 shown in (A) of, the electronic device can respond to the user's click operation on the above capture input option and display the user interface 32 as shown in Figure 3 (B) of. The user interface 32 can be a multi-window interface displayed in a split-screen manner, which includes a first area 321 and a second area 322. Among them, the first area 321 is used to display the preview interface obtained by the camera of the electronic device capturing the scene, that is, the viewfinder corresponding to the camera, and the second area 322 is used to display the application interface of the "Notes" application. Specifically:
[0126] In the first region 321, it can be used to display a preview screen obtained by the camera shooting a shooting scene. The camera can be a front camera of the electronic device or a rear camera of the electronic device, and the present application does not limit this. It should be noted that the above preview screen is not the photo obtained after the user clicks the shooting control, but the screen displayed after the camera takes real-time views of the shooting scene, and its screen content can change with the shooting scene, the shooting angle and direction of the camera, and the shooting parameters of the camera (such as exposure time, filter effect, etc.).
[0127] The second region 322 includes an edit box 3221. The edit box 3221, that is, the edit box 311 in the foregoing description, can be used to receive and display information such as text and images input by the user. In some embodiments, the edit box 3221 can be the entire second region 322. In addition, in some embodiments, the boundary of the edit box 311 may not be displayed in the second region 322, and it may exist in the second region in a form perceptible to the electronic device (such as program code).
[0128] After the camera is started, the user can aim the camera at an object so that the image corresponding to the object is displayed in the screen of the first region 321. Then, the electronic device can extract the image corresponding to the object from the screen of the first region 321 and pre-display the extracted image in the edit box 3221.
[0129] As Figure 3 shown in (B) of, there is a unique target object in the preview screen displayed in the first region 321 at this time, that is, the target object 321A. Therefore, the electronic device can automatically use the target object 321A as the object to be cropped, extract it from the screen displayed in the first region 321 to obtain a cropped image 322A, and pre-display the cropped image 322A in the edit box 3221. It should be understood that the "pre-displayed in the edit box 3221" mentioned here means that the cropped image displayed in the edit box 3221 is not the cropped image officially input by the user into the second region, but a preview image displayed in the second region by the electronic device to facilitate the user to view the cropping effect of the cropped image. In the case where the user determines that the cropping effect of the cropped image in the second region does not meet the user's expectations, the user can change the imaging effect of the target object 321A in the first region by adjusting the shooting angle, shooting parameters, etc., and then change the cropping effect of the cropped image 322A obtained by the electronic device performing a cropping operation on the target object 321A until the cropping effect of the cropped image 322A meets the user's expectations.
[0130] In some embodiments, the cutout image 322A in the editing box 3221 may be blurred or processed to be out of focus. That is, in the second region 322, the clarity of the cutout image 322A may be reduced to prompt the user that the image is a preview image, and at the same time prompt the user that the cutout image has not been officially added to the editing box 3221 yet.
[0131] In some embodiments, the area size of the cutout image 322A pre-displayed in the editing box 3221 may be the same as or different from the area size of the target object 321A in the first region 321. Optionally, after performing a cutout operation on the target object 321A to obtain the cutout image 322A, the electronic device may adaptively adjust the size of the cutout image 322A according to the area of the editing box 3221, and then display the resized image 322A in the editing box 3221.
[0132] When the cutout effect of the cutout image 322A meets the user's expectations, the electronic device may respond to a user operation, such as Figure 3 the operation of long-pressing and dragging the target object 321A and then dragging it into the editing box 3221 as shown in (C) below, and officially input the cutout image 322A into the editing box 3221. At the same time, the camera is turned off, and the Figure 3 user interface 34 as shown in (D) below is displayed.
[0133] In the user interface 34, the user may continue to input text information into the input area 341 through the keyboard 342, or continue to input other cutout images into the input area 341 according to the above operation. It should be noted that since the cutout image 322A in the input area 341 is the cutout image officially input by the user, the electronic device may not perform blurring or out-of-focus processing on the cutout image 322A in the input area 341.
[0134] It should be understood that Figure 3Only the user interfaces provided by the embodiments of the present application are exemplarily shown, which should not constitute a limitation to the embodiments of the present application. For example, in some embodiments, the user can start the camera of the electronic device in other ways; for another example, the electronic device can input the matte image into the input area in other ways; or, after the matte image is officially input into the input area, the electronic device can also not turn off the camera and continue to display the input area and the preview screen captured by the camera on the same interface, so that the user can subsequently input other matte images into the input area by the same method; or, the user can input the matte image into the edit box by means of a voice command instead of the method of long-pressing the image and dragging it to the edit box, or the electronic device can automatically input the matte image into the edit box when no user operation is received for a certain period of time. The present application does not make any limitation thereto.
[0135] In the embodiments of the present application, the user interface 31 may be referred to as the "first interface", the edit box 311 may be referred to as the "first edit box", the shooting input option may be referred to as the "first control", and the click operation of the user on the shooting input option may be referred to as the "first operation". The user interface 32 may be referred to as the "second interface", the object 321A may be referred to as the "first object", the matte images 322A in the user interface 32 and the user interface 33 may be referred to as the "second object", the matte image 322A in the user interface 34 may be referred to as the "thirteenth object", the preview screen captured by the camera of the shooting scene displayed in the first area 321 may be referred to as the "first screen", and the operation of long-pressing and dragging the object 321A to the edit box 3221 in the user interface 33 may be referred to as the "fourth operation".
[0136] In the actual process of the user using the electronic device, the scenarios where the user is located may be relatively complex. When the user uses the image processing method provided by the present application for matte extraction, various possible situations may be involved. For example, there may be no target object in the image captured by the camera, or there may be multiple target objects in the image, or in the case of multiple target objects in the image, the matte object automatically determined by the electronic device is not the matte object that the user really expects, and how to adopt an appropriate matte extraction strategy for different types of matte objects, etc. For these situations, the electronic device provided by the present application is provided with corresponding methods and processing logics, which will be described in combination with specific embodiments hereinafter.
[0137] Next, in combination with Figure 4 the shown method flow chart, the image processing method provided by the present application will be described in more detail.
[0138] As Figure 4As shown, the image processing method provided by this application may include the following steps:
[0139] S401: Invoke the camera viewfinder.
[0140] In response to a user operation, the electronic device invokes the viewfinder of the camera.
[0141] The above-mentioned electronic device may be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), or a dedicated camera (such as a single-lens reflex camera, a compact camera), etc. This application does not impose any restrictions on the specific type of this electronic device. Specifically, the above-mentioned electronic device may be the electronic device shown in the foregoing Figure 3 shown.
[0142] The above-mentioned user operation may be that after the user long-presses on the blank area in the edit box displayed on the electronic device to bring up the corresponding control, and then clicks on the control. Specifically, reference may be made to the relevant description of Figure 3 above. Without being limited thereto, the user operation may also be other operations. For example, the user may also use a voice command to cause the electronic device to invoke the camera viewfinder.
[0143] S402: Display the viewfinder image captured by the camera in real time in the first area of the first interface.
[0144] After invoking the camera, the electronic device may simultaneously display two display areas on the same interface. Among them, the first area is used to display the camera viewfinder, and the preview image obtained by the camera's real-time view of the shooting scene may be displayed in the viewfinder; the second area may include the aforementioned edit box.
[0145] S403: Determine whether the edit box in the second area supports image input.
[0146] It should be understood that in actual application scenarios, the edit box displayed on the electronic device can receive data such as text and images input by the user. For example, the edit box provided by the note application of the electronic device. However, some edit boxes have been specified at the corresponding code level of the electronic device that the edit box can only input text and cannot input other types of data such as images. For example, the account and password input fields provided by some applications.
[0147] Therefore, in the present application, after the user invokes the viewfinder of the camera, the electronic device can determine whether the edit box included in the second region supports image input. When the edit box included in the second region does not support image input, the electronic device can perform the subsequent step S404; when the edit box included in the second region does not support image input, the electronic device can continue to perform the subsequent step S405.
[0148] S404. Output a first prompt message.
[0149] When the edit box included in the second region does not support image input, the electronic device does not need to perform matting on the object included in the picture displayed in the first display region. Correspondingly, the electronic device can output a first message to prompt the user that the edit box does not support image input.
[0150] As Figure 5 shown, the user interface 51 can be a multi-window interface displayed by the electronic device after invoking the camera, which includes a first region 511 and a second region 512. The first region 511 is a viewfinder, which is used to display the preview picture obtained by the real-time view of the camera on the electronic device, and a target object 511A is displayed therein. The second region 512 includes an edit box 5121. Since the edit box 5121 does not support image input, different from the edit box 3221 in the foregoing Figure 3 which displays the preview image of the electronic device's matting, no preview image may be displayed in the edit box 5121. Instead, a prompt box 513 is displayed, and the prompt box 513 may include a prompt message "This edit box does not support image input", and this prompt message is the above-mentioned first prompt message. Specifically, the position where the prompt box 513 is displayed can be at the top, middle or bottom of the user interface 51. After being displayed, the prompt box 513 can automatically disappear after a certain time interval (3s - 5s), or automatically disappear after the user clicks on other positions in the user interface 51.
[0151] S405. Determine whether a target object is detected.
[0152] The electronic device detects whether there is a target object in the picture displayed in the above-mentioned viewfinder through a target detection algorithm. In this method, the target object is also the salient object. The target detection algorithm can be traditional algorithms such as Viola Jones Detector, HOG Detector, DPM Detector, etc., or the target algorithm can be an algorithm based on machine learning or deep learning, such as One-Stage algorithms represented by YOLO, SSD, etc., Two-Stage algorithms represented by R-CNN, Fast R-CNN, Faster R-CNN, etc., transformer-based algorithms represented by Relation Net, DETR, etc., Anchor-free algorithms represented by CornerNet, CenterNet, etc., Trick algorithms represented by FPN, CascadeR-CNN, etc. This application does not make any limitations in this regard.
[0153] For an image taken in a scene with a complex environment, the electronic device may be able to detect multiple target objects from among the multiple objects included in the image; however, for a solid-color image (such as an image obtained by photographing a pure blue sky), the electronic device may not be able to detect any target object from the image. That is to say, in the same picture, the target object may not exist, may exist one, or may exist multiple.
[0154] In the case where no salient object is detected, the electronic device can directly execute the subsequent step S410; otherwise, the electronic device will execute the subsequent step S406.
[0155] In some embodiments, when the target detection algorithm used by the electronic device does not specify the category of the target object, the above-mentioned target object can be all the objects in the image that can attract people's attention.
[0156] Such as Figure 6As shown in (A) of [the figure], the user interface 61 may be a multi-window interface displayed on the electronic device after activating the camera, which includes a first area 611 and a second area 612. The first area 611 is a viewfinder, which is used to display the preview image obtained by the real-time view of the camera on the electronic device, and objects 611A, 611B, 611C, and 611D are shown therein. The second area 612 contains an edit box 6121. In some embodiments, the category of the target object may not be limited. The electronic device will determine objects 611A, 611B, 611C, and 611D as target objects through the target detection algorithm. Correspondingly, the electronic device may generate a tracking box for each object in the target objects to prompt the user that the object is a target object. The user can cut out the image corresponding to any one of these target objects through the corresponding user operation and input it into the edit box 6121. In addition, in the case where the user does not make a selection, the electronic device may also automatically cut out the image corresponding to one of the target objects (for example, the target object with the largest area or the target object closest to the center point of the first area. For specific reference, please refer to the subsequent embodiments and will not be elaborated here first) and pre-display it in the edit box 6121. As Figure 6 As shown in (A) of [the figure], since object 611C is the target object closest to the center point of the first area (i.e., the point where the "+" symbol shown in the second area 612 is located, the same below) among all the target objects, the electronic device may cut out the image corresponding to object 611C and pre-display it in the edit box 6121, that is, the image 612C shown in the edit box 6121.
[0157] In some embodiments, when the target detection algorithm used by the electronic device does not specify the category of the target object, the above target object may be an object belonging to the category specified by the above target detection algorithm among all the objects that can attract people's attention in the image. This category may include categories such as people, dogs, cats, flowers, etc., and may also include other categories, which can be determined according to the user's image cropping preferences and specific usage scenarios. The present application does not limit this. When the electronic device discriminates from the salient objects included in the image based on the target detection algorithm, if the electronic device detects that the category of an object does not belong to the category it has pre-specified based on the target detection algorithm, then this object will not be regarded as a target object, and therefore the image corresponding to this object cannot be cut out.
[0158] For the convenience of description, it is assumed here that the category specified by the above target detection algorithm only includes the category of "cube". Then as Figure 6As shown in (B) of , the user interface 62 may be a multi-window interface displayed by the electronic device after activating the camera, which includes a first area 621 and a second area 622. The first area 621 is a viewfinder, which is used to display the preview image obtained by the real-time view of the camera on the electronic device, and objects 621A, 621B, 621C, and 621D are displayed therein. The physical entities of the objects corresponding to objects 621A and 621B are cubes. The second area 622 contains an edit box 6221. Since the target detection algorithm stipulates that the categories of target objects only include "cube", the electronic device will only determine objects 621A and 621B as target objects, and will not determine objects 621C and 621D as target objects. Correspondingly, the electronic device will also generate a tracking box for each of objects 621A and 621B respectively, to prompt the user that only the images corresponding to these two objects can be cut out and input into the edit box 6221. In addition, since object 621C is not a target object, even if it is the closest to the center point of the first area, the electronic device will not cut out the image corresponding to object 621C and pre-display it in the edit box 6221. Instead, the image corresponding to one of objects 621A and 621B (here it is assumed to be the object with the largest area among the target objects, that is, object 621A) is cut out to obtain a cut-out image 622A, and the cut-out image 622A is pre-displayed in the edit box 6221.
[0159] In some embodiments, the edit box for receiving the cut-out image may also stipulate the category of the image. For example, it is stipulated that the images input therein can only be images of categories such as portraits and plants. Still taking the user interface 62 and the edit box 6221 as an example, assuming that the categories stipulated by the above target detection algorithm only include multiple categories such as "cube", "cylinder", and "sphere", the electronic device may determine objects 621A, 621B, 621C, and 621D as target objects. However, in the case where the category of the image that the edit box 6221 can receive is stipulated to only include "cylinder", even if the electronic device determines object 621C as the cut-out object and has cut out the corresponding image, the edit box 6221 may not receive (i.e., not display) the cut-out image corresponding to object 621C, and prompt the user that "the current cut-out object does not meet the input requirements", prompting the user to replace the cut-out object. Until the user determines object 621D as the cut-out object through corresponding operations, the electronic device can cut out the image corresponding to object 621D and successfully input it into the edit box 6221 for pre-display.
[0160] In the embodiments of the present application, the user interface 62 may be referred to as the "second interface", the object 621A may be referred to as the "first object", the screen displayed in the first area 621 may be referred to as the "first screen", the category "cube" may be referred to as the "first category", and the objects 621A and 621B may be collectively referred to as "at least one target object existing in the first screen".
[0161] S406. Obtain the matte object's matte image.
[0162] In this method, the electronic device can determine the matte object from the target objects in two ways.
[0163] First, when the user just activates the viewfinder of the camera through a user operation and has not operated on any of the target objects, the electronic device can determine the matte object from all the target objects according to parameters such as the distance of each target object from the center point of the screen (i.e., the center point of the viewfinder, which is also the center point of the aforementioned first area) and the area of each target object. Specifically, reference can be made to the relevant descriptions in the foregoing and subsequent embodiments, which will not be elaborated here. Figure 6 And the relevant descriptions of the subsequent embodiments, which will not be elaborated here.
[0164] Second, if the matte object determined by the electronic device does not meet the user's expectations, or the user wants to matte out one or more other target objects together with the target object automatically determined by the electronic device, the user can switch or add the target object as the matte object through a user operation (such as single-clicking or double-clicking on the target object). Specifically, reference can be made to the relevant descriptions of the subsequent embodiments, which will not be elaborated here first.
[0165] S407. Determine a matte strategy according to the matte object.
[0166] It should be understood that the contour complexities and specific uses of different matte objects are different. For example, for matte objects such as human faces, cats, and dogs that may have hair, since there may be hair strands in the images corresponding to these objects, more precise matte strategies (algorithms) are required to identify the contours of these objects. However, for some objects with relatively regular shapes, such as water cups, boxes, etc., only algorithms with general precision can achieve relatively good matte effects.
[0167] However, the higher the precision of the matte strategy, the higher the requirements for the computing power and power consumption of the electronic device. Therefore, in the embodiments of the present application, after determining the above-mentioned matte object, the electronic device can determine a corresponding matte strategy for the matte object according to the type and use of the matte object, so as to balance the quality of the matte image and the computing power and power consumption of the electronic device. Specifically, reference can be made to the relevant descriptions of the subsequent embodiments. This will not be elaborated here first.
[0168] S408. Obtain the matte image of the object to be matted.
[0169] After determining the object to be matted and its corresponding matte strategy, the electronic device can first use the matte strategy to extract the image corresponding to the object to be matted, obtaining the above-mentioned matte image.
[0170] S409. Determine whether the quality of the matte image is too low.
[0171] In an actual application scenario, the area ratio of the object to be matted in the viewfinder, the matte strategy used by the electronic device, and the shooting angle and light intensity when the user takes a photo may all affect the quality of the matte image. For example, when the area ratio of the object to be matted in the viewfinder is too small, the above-mentioned matte image obtained by the electronic device for matting it may have a phenomenon of incomplete extraction, or the matte image may be blurred due to too low scene light intensity, etc.
[0172] Therefore, in the embodiment of the present application, after the electronic device obtains the above-mentioned matte image, it can use a corresponding image quality evaluation algorithm to evaluate the matte image. When the quality of the above-mentioned matte image is too low, the electronic device can default that there are no target objects in the preview screen displayed in the viewfinder and execute step S410. Otherwise, the electronic device can execute subsequent step S411.
[0173] S410. Output a second prompt message.
[0174] In the case where the electronic device determines that there are no target objects in the preview screen displayed in the viewfinder, the electronic device can output the above-mentioned second prompt message to prompt the user to manually select the object to be matted or adjust the shooting parameters.
[0175] Such as Figure 7As shown, when the quality of the above-mentioned matte image is too low, the electronic device can display a user interface 71, which includes a first area 711 and a second area 712. The first area 711 is a viewfinder, which is used to display the preview screen obtained by the real-time viewfinder of the camera on the electronic device, in which an object 711A is displayed. The second area 712 contains an editing box 7121. Although there is an object 711A in the first area 711, since its area ratio in the first area 711 is too small, the electronic device will consider that there is no target object in the first area 711. Correspondingly, no preview image will be displayed in the editing box 7121, and the electronic device can display a prompt box 713 in the user interface 71. The prompt box 713 can contain a prompt message "No target detected. Please adjust the shooting parameters to change the picture content", and this prompt message is the above-mentioned second prompt message. It can prompt the user that there is no target object in the picture displayed in the first area 711, which may be because the lens is too far from the object being photographed, resulting in too small an area of the target object, or the parameters of the electronic device during shooting cause the image to be unclear, etc. Therefore, the user can adjust shooting parameters such as the shooting angle, aperture, and focal length to change the picture content in the preview screen, so that the electronic device can find the target object from it and obtain a matte image for pre-display in the editing box 7121.
[0176] S411. Pre-display the matte image of the matte object in the editing box of the second area.
[0177] When the quality of the above-mentioned matte image meets the evaluation criteria of the electronic device, the electronic device can pre-display the matte image cut out for the matte object in the editing box of the above-mentioned second area.
[0178] S412. Real-time track the matte object in the viewfinder screen.
[0179] S413. Determine whether the matte object is lost during tracking.
[0180] In the embodiments of the present application and subsequent embodiments, "pre-displayed in the editing box of the second area" all means that the matte image displayed in the editing box of the second area is not the matte image officially input by the user into the second area, but a preview image displayed in the second area by the electronic device to facilitate the user to view the matte effect of the matte image. That is to say, in the case where the user determines that the matte effect of the matte image in the above-mentioned editing box does not meet the user's expectations, the user can change the imaging effect of the above-mentioned matte object in the viewfinder of the above-mentioned first area by adjusting the shooting angle, shooting parameters, etc., and then change the matte effect of the above-mentioned matte image cut out by the electronic device for the matte object, until the matte effect of the matte image displayed in the above-mentioned editing box meets the user's expectations, and the user can officially input the matte image into the editing box of the above-mentioned second area through user operations.
[0181] That is to say, when the object to be cropped remains unchanged, the electronic device can track the object to be cropped in the above-mentioned viewfinder in real time. When the object to be cropped changes in the viewfinder (such changes include changes in the area, imaging angle, etc. of the object to be cropped), the electronic device can re-crop the object to be cropped in the picture, or scale the cropped image obtained historically by a certain proportion, so as to simultaneously change the preview effect of the cropped image displayed in the above-mentioned editing box. Correspondingly, when the main part or the whole of the cropped image moves out of the above-mentioned second area (i.e., the above-mentioned viewfinder), the electronic device can determine at this time that the original tracking of the object to be cropped is lost, and the electronic device will start from the foregoing step S405 and continue to execute this method according to the foregoing method logic. For related embodiments of tracking the object to be cropped, reference can be made to the subsequent description, which will not be elaborated here first.
[0182] S414. In response to a first operation on the object to be cropped in the first area, the cropped image of the object to be cropped is officially input into the editing box in the second area.
[0183] When the cropping effect of the cropped image displayed in the editing box in the above-mentioned second area meets the user's expectation, the user can officially input the cropped image into the editing box in the above-mentioned second area through user operations. Specifically, reference can be made to the relevant descriptions of (C) and (D) above. Embodiments of the present application and subsequent embodiments will not be elaborated. Figure 3 For the relevant descriptions of (C) and (D) above, embodiments of the present application and subsequent embodiments will not be elaborated.
[0184] Next, in combination with Figure 8 the specific process of the electronic device determining the object to be cropped in the preview screen will be described.
[0185] Combined with the foregoing description, it can be seen that in the image processing method provided in the present application, the electronic device can determine the object to be cropped from the target object in two ways. One is that when the viewfinder of the camera is activated, the electronic device can automatically determine the object to be cropped from all target objects according to parameters such as the distance of each target object from the center point of the screen and the area of each target object. The other is that if the object to be cropped automatically determined by the electronic device does not meet the user's expectation, the electronic device can respond to the interactive operation between the user and the screen to switch, add or reduce the target object as the object to be cropped.
[0186] In Figure 8Among them, steps S802 - S807 are the process for the electronic device to automatically determine the object to be cropped from the target object, and steps S806 - S811 are the process for the electronic device to determine the object to be cropped from the target object in response to the user's operation. For the convenience of understanding, hereinafter, it is assumed that the electronic device only displays the tracking frame of the object to be cropped in the interface, and the tracking frames of other target objects are not displayed in the interface. And hereinafter, it is assumed that the first area (that is, the objects displayed in the viewfinder all belong to the categories predefined by the target detection algorithm) is as Figure 8 shown. In the image processing method provided in this application, the steps for the electronic device to determine the object to be cropped in the preview screen may include:
[0187] S801. Invoke the camera viewfinder.
[0188] S802. Whether the center point of the camera viewfinder is within a target object.
[0189] It should be understood that people's line of sight will first be attracted by the object in the center of the screen. In addition, combined with the shooting habits of most users, it can be known that users often place the object they want to shoot (the entity corresponding to the object to be cropped in this application) at the center of the viewfinder. Therefore, the electronic device can first determine whether the center point of the viewfinder is within a certain target object included in the viewfinder screen. If so, step S803 is executed, that is, the target object where the center point is located is determined as the object to be cropped; otherwise, the electronic device will execute subsequent steps S804 - S805, that is, further determine the distance and area of each target object from the center point, so as to determine an object as the object to be cropped from the target objects.
[0190] S803. Determine the object where the center point is located as the object to be cropped.
[0191] As Figure 9 shown, the user interface 81 may be the interface displayed after the electronic device invokes the camera viewfinder, which includes a first area 811 and a second area 812. The first area 811 is the viewfinder, which is used to display the preview screen obtained by the camera on the electronic device in real time, and there are target objects 811A and target object 811B displayed therein. The second area 812 contains an edit box 8121. Among them, the first area 811 can be approximately regarded as a rectangle, and the four vertices of this rectangle are Figure 9 the a point, b point, c point, and d point shown in. Combining basic mathematical knowledge, it can be known that the intersection of the two diagonals of the rectangle is the center point of the rectangle. Therefore, in Figure 9Among them, the center point of rectangle abcd is point P, and this point P happens to fall on the target object 811A. Therefore, the electronic device will automatically determine the target object 811A as the object to be cropped. Correspondingly, the electronic device can generate a tracking frame for the target object 811A, and crop out the image corresponding to the target object 811A and pre-display it in the editing frame 8121, that is, the image 812C shown in the editing frame 8121.
[0192] S804. Calculate the Manhattan distance between each target object in the viewfinder screen and the center point.
[0193] S805. Whether the differences between the Manhattan distances of each target object and the center point are all less than a preset threshold.
[0194] The electronic device can calculate the Manhattan distance between each target object and the center point of the viewfinder frame (actually, take a point on each target object, such as the center point or the center of gravity of the object, and calculate the Manhattan distance between this point and the center point of the viewfinder frame). After sorting the Manhattan distances between each target object and the center point of the viewfinder frame to obtain an ordered Manhattan distance sequence, the electronic device can calculate the differences between adjacent two Manhattan distances in the sequence in turn. If the differences between each adjacent two Manhattan distances are all less than the above preset threshold, the electronic device can determine that there is no target object in the viewfinder frame that is significantly close to the center point of the viewfinder frame, and the electronic device can execute the subsequent step S806, that is, determine the target object with the largest area among all target objects as the object to be cropped (generally speaking, the larger the area of the object, the more attention it attracts, that is, the more likely it is the object to be cropped); if there are differences between adjacent two Manhattan distances in the above sequence that are all greater than or equal to the above preset threshold, the electronic device can execute the subsequent step S807, that is, determine the target object with the smallest Manhattan distance from the center point of the viewfinder frame among all target objects as the object to be cropped.
[0195] S806. Determine the target object with the largest area as the object to be cropped.
[0196] S807. Determine the target object with the smallest Manhattan distance from the center point as the object to be cropped.
[0197] As Figure 10 As shown in (A) of, the user interface 82 can be the interface displayed after the electronic device activates the camera viewfinder frame, which includes a first area 821 and a second area 822. The first area 821 is the above-mentioned viewfinder frame, which is used to display the preview image obtained by the camera on the electronic device in real time, in which there are shown a target object 821A, a target object 821B, and a target object 821C. The second area 822 contains an editing frame 8221.
[0198] Among them, the first area 821 can be approximately regarded as a rectangle, and the schematic diagram after its magnification can be referred toFigure 10 The coordinate diagram shown in (B) in FIG. For ease of explanation and understanding, it is assumed here that the center point of the first area 821 is point O, and its coordinates are O: (0, 0). The center points of the target object 821A, the target object 821B, and the target object 821C are respectively point P1, point P2, and point P3, and their coordinates are respectively P1: (x1, y1), P2: (x2, y2), and P3: (x3, y3). Based on basic mathematical knowledge, it can be known that the Manhattan distances of point P1, point P2, and point P3 from point O are d P1 =|x1|+|y1|, d P2 =|x2|+|y2|, d P2 =|x3|+|y3|. Here we assume that d P2 >d P1 >d P3 , and the imaging areas of the target object 821A, the target object 821B, and the target object 821C in the first region are S 821A , S 821B , S 821C , S 821B >S 8211 >S 821C .
[0199] Combined with the above description, it can be seen that if (d P3 -d P1 )、(d P1 -d P2 ) are all less than the above-mentioned preset threshold, the electronic device will determine the target object with the largest area among all the target objects, that is, the target object 821B, as the cutout object. Then the user interface 82 can be further expressed as follows Figure 10 In the user interface 83 shown in (C), the electronic device automatically determines the target object 821B as the cutout object. Accordingly, the electronic device can generate a tracking frame for the target object 821B, and cut out the image corresponding to the target object 821B and pre-display it in the editing frame 8221, that is, the image 822B shown in the editing frame 8221.
[0200] If (d P3 -d P1 )、(d P1 -d P2 ) is less than the above-mentioned preset threshold, the electronic device will determine the target object with the smallest Manhattan distance from the center point of the view frame among all the target objects, that is, the target object 821C as the cutout object. Then the user interface 83 can be further expressed as follows Figure 10The user interface 84 shown in (D) in []. In the user interface 84, the electronic device automatically determines the object 821C as the object to be cut out. Correspondingly, the electronic device can generate a tracking frame for the target object 821C, and cut out and pre-display the image corresponding to the target object 821C in the editing box 8221, that is, the image 822C shown in the editing box 8221.
[0201] It should be noted that before the user officially inputs the cut-out image into the editing box, the user can change the field of view angle range and direction of the camera (changing the field of view angle range and direction will cause changes in the number, position, and size of the target objects in the picture), so that the electronic device changes the pre-display style of the object to be cut out in the editing box or changes the original object to be cut out into a new object to be cut out.
[0202] In the embodiments of the present application, the user interface 83 and the user interface 84 can be referred to as the "second interface", and the object 821A, the object 811B, and the object 821C can be collectively referred to as "multiple objects included in the first picture". In the user interface 84, the point O can be referred to as the "center point of the first area", the object 821C can be referred to as the "first object", its corresponding cylinder can be referred to as the first entity, and the object 821A and the object 821B can be referred to as the "seventh object", and its corresponding cone or cube can be referred to as the "third entity". In the user interface 83, the point O can be referred to as the "center point of the first area", the object 821B can be referred to as the "first object", its corresponding cube can be referred to as the first entity, and the object 821A and the object 821C can be referred to as the "eighth object", and its corresponding cone or cylinder can be referred to as the "eighth entity".
[0203] S808. The user clicks on other areas in the viewfinder.
[0204] Of course, if the object to be cut out automatically determined by the electronic device does not meet the user's expectations, the electronic device can respond to the interaction operation between the user and the screen, such as a click or double-click operation, to switch, add, or reduce the target object as the object to be cut out.
[0205] S809. Whether the contact point clicked by the user exists on the target object.
[0206] It is understandable that when the user clicks on other areas in the viewfinder, if the contact point between the user and the screen is not on any other target object (i.e., other target objects in the viewfinder except the currently determined object to be cut out), the electronic device can perform the subsequent step S810, that is, keep the object to be cut out as the originally determined object to be cut out; however, if the contact point between the user and the screen is on a certain target object (including the currently determined object to be cut out), the electronic device can perform the subsequent step S811, that is, re-determine the object to be cut out according to the user's specific operation on the target object.
[0207] S810. Keep the object to be cut out unchanged.
[0208] S811. Re-determine the object to be cut out.
[0209] In the embodiments of the present application, the user can switch the object to be cut out through a single-click operation on the target object, or increase or decrease the object to be cut out through a double-click operation on the target object.
[0210] 1) Switch the object to be cut out through a single-click operation. For details, please refer to Figure 11 .
[0211] As Figure 11 shown in (A) of, the user interface 85 can be a multi-window interface displayed by the electronic device after just activating the camera, which includes a first area 851 and a second area 852. The first area 851 is a viewfinder, which is used to display the preview image obtained by the real-time view of the camera on the electronic device, and target object 851A and target object 851B are displayed therein; the second area 852 is a viewfinder, which is used to display the preview image obtained by the real-time view of the camera on the electronic device, and it contains an editing box 8521. Since the center point of the first area (i.e., the viewfinder) happens to fall within target object 851A, the electronic device will automatically determine target object 851A as the object to be cut out. Correspondingly, the electronic device will also generate a tracking box for object 851A, cut out the image corresponding to target object 851A to obtain a cut-out image 852A, and pre-display it in the editing box 8521.
[0212] If target object 851A is not the cut-out image expected by the user, and the image that the user actually wants to cut out is the image corresponding to target object 851B in the first area, the user can change the cut-out image of the electronic device through corresponding user operations. As Figure 11 shown in (A) of, in response to the user's single-click operation on the area where target object 851B is located, the electronic device can switch the object to be cut out from target object 851A to target object 851B (at this time, the object to be cut out only includes target object 851B, and target object 851A is no longer the object to be cut out), and display a figure as Figure 11The user interface 86 shown in (B) of []. In the user interface 86, the electronic device has determined the target object 851B as the object to be cut out. Similarly, the electronic device will also generate a tracking frame for the target object 851B, cut out the image corresponding to the target object 851B to obtain the cut-out image 852B, and pre-display it in the editing box 8521. Correspondingly, since the target object 851A is no longer the object to be cut out at this time, the original cut-out image 852A will no longer be pre-displayed in the editing box 8521.
[0213] In the embodiments of the present application, the user interfaces 85 and 86 can be referred to as the "second interface", the object 851A can be referred to as the "first object", the object 851B can be referred to as the "ninth object", the picture shown in the first area 851 can be referred to as the "first picture", the cut-out image 852A can be referred to as the "second object", and the cut-out image 852B can be referred to as the "tenth object". The click operation of the user on the object 851B in the user interface 85 can be referred to as the "second operation".
[0214] 2) Increase / decrease the object to be cut out through a double-click operation. For details, please refer to Figure 12 .
[0215] As Figure 12 shown in (A) of [], the user interface 87 can be a multi-window interface displayed by the electronic device after just activating the camera. It includes a first area 871 and a second area 872. The first area 871 is a viewfinder, which is used to display the preview picture obtained by the real-time view of the camera on the electronic device, and the target objects 871A, 871B, and 871C are shown therein; it should be noted that for some target detection algorithms with better performance, the target object 871A can be refined into the target objects 871A1 and 871A2; the second area 872 is a viewfinder, which is used to display the preview picture obtained by the real-time view of the camera on the electronic device, and it contains an editing box 8721. Since the center point of the first area (i.e., the viewfinder) happens to fall within the target object 871B, the electronic device will automatically determine the object 871B as the object to be cut out. Correspondingly, the electronic device will also generate a tracking frame for the target object 871B, cut out the image corresponding to the target object 851B to obtain the cut-out image 872B, and pre-display it in the editing box 8721.
[0216] If the user also wants to use other objects together with the target object 851B as the object to be cut out, the user can, through a double-click operation on other objects (objects that have not been determined as the object to be cut out), determine that object together with the object that has already been determined as the object to be cut out as the object to be cut out.
[0217] For example, if the user not only wants to extract the image corresponding to the target object 871B, but also wants to extract the image corresponding to the target object 871A from the first region 871, the user can, through corresponding operations, when the target object 871B is already a cutout object, also determine the target object 871A as a cutout object at the same time. As shown in (A) of Figure 12 in response to the user's double-click operation on the area where the target object 871A is located, the electronic device can also determine the target object 871A as a cutout object (at this time, the cutout objects include the target object 871B and the target object 871A at the same time), and display a user interface 88 as shown in (B) of Figure 12 in. In the user interface 88, the electronic device has already determined the target object 871A and the target object 871B as cutout objects at the same time. The electronic device will also generate a tracking frame for each of the target object 871A and the target object 871B, and extract the images corresponding to the target object 871A and the target object 871B to obtain the cutout images 872A and 872B, and pre-display these two cutout images together in the editing frame 8721. Specifically, the sizes and positions of the cutout image 872A and the cutout image 872B in the editing frame 8721 can correspond to the sizes and positions of the target object 871A and the target object 871B in the first region 871. For example, if in the first region 871, the target object 871A is directly to the left of the target object 871B, and the imaging area of the target object 871A is larger than the imaging area of the target object 871B, then in the editing frame 8721, the cutout image 872A can also be directly to the left of the cutout image 872B, and the area of the cutout image 872A can also be larger than the area of the cutout image 872B.
[0218] Similarly, if the user also wants to extract the image corresponding to the target object 871C from the first region 871, the user can, through corresponding operations, when both the target object 871A and the target object 871B are cutout objects, also determine the target object 871C as a cutout object at the same time. As shown in (B) of Figure 12 in response to the user's double-click operation on the area where the target object 871C is located, the electronic device can also determine the target object 871C as a cutout object (at this time, the cutout objects include the target object 871A, the target object 871B, and the target object 871C at the same time), and display a figure as shown in (B) of Figure 12871B and 871C as cutout objects. The electronic device also generates a tracking frame for each of the target objects 871A, 871B and 871C, and cuts out the images corresponding to the target objects 871A, 871B and 871C to obtain the cutout images 872A, 872B and 872C, and pre-displays the three cutout images together in the edit box 8721. Similarly, the size and position of the cutout images 872A, 872B and 872C in the edit box 8721 may correspond to the size and position of the target objects 871A, 871B and 871C in the first area 871.
[0219] Of course, if the user mistakenly selects an object as a cutout object, the user can also double-click the selected cutout object to no longer determine the object as a cutout object. Alternatively, the user can single-click a cutout object to cancel the previously determined cutout object with one click and newly determine the clicked cutout object as a separate cutout object.
[0220] For example, Figure 12 For example, when target object 871A, target object 871B, and target object 871C are all cutout objects, if the user no longer wants to determine target object 871B as a cutout object, the user can cancel the determination of target object 871B as a cutout object through corresponding operations if target object 871B is already a cutout object. Figure 12 As shown in (C) in FIG. 1 , in response to a user double-clicking the area where the target object 871B is located, the electronic device may no longer determine the target object 871B as a cutout object (at this time, the target object 871A and the target object 871C are still cutout objects), and display the figure as shown in FIG. Figure 12 In the user interface 90 shown in (D) in FIG. 8 , the electronic device no longer determines the target object 871B as the cutout object, and only determines the target object 871A and the target object 871C as the cutout objects. The electronic device only generates a tracking frame for each of the target object 871A and the target object 871C, and cuts out the images corresponding to the target object 871A and the target object 871C to obtain the cutout image 872A and the cutout image 872C, and pre-displays the two cutout images together in the edit box 8721.
[0221] In addition, based on the fact that the electronic device displays the user interface 90, if the user subsequently decides to separately determine the target object 871B as the object to be cut out, the user can, through corresponding operations, cancel the determination of the target object 871A and the target object 871C as the objects to be cut out when both the target object 871A and the target object 871C are the objects to be cut out, and re-determine the target object 871B as the only object to be cut out. As shown in Figure 12 (D) shown below, in response to a click operation on the area where the target object 871B is located by the user, the electronic device re-determines the target object 871B separately as the object to be cut out (at this time, the target object 871A and the target object 871C are no longer the objects to be cut out), and displays a user interface 91 as shown in Figure 12 (E) shown below. In the user interface 91, the electronic device only determines the target object 871B as the object to be cut out. The electronic device generates only one tracking box for the target object 871B, cuts out the image corresponding to the target object 871B, and after obtaining the cut-out image 872B, pre-displays the image 872B in the editing box 8721.
[0222] It should be further noted that for the target object 871A existing in the foregoing user interface 87 - user interface 91, if the performance of the target detection algorithm and the cut-out algorithm used by the electronic device is good enough, the target object 871A can be subdivided into the target object 871A1 and the target object 871A2. When the user clicks or double-clicks on the target object 871A, the electronic device can accurately determine which of these two objects the user has operated on according to whether the contact point when the user clicks on the screen falls on the target object 871A1 or the target object 871A2, so as to generate a corresponding mask for the object being operated on. Even if there is an overlapping part between the target object 871A1 and the target object 871A2, the electronic device can accurately cut out the image corresponding to one of the objects, obtain the corresponding cut-out image and display it in the editing box 8721.
[0223] In the embodiments of the present application, the user interface 87 and the user interface 88 can be referred to as the "second interface", the object 871B can be referred to as the "first object", the object 871A can be referred to as the "eleventh object", the picture displayed in the first area 871 of the user interface 81 can be referred to as the "first picture", and the cut-out image 872B displayed in the second area 872 can be referred to as the "second object"; the cut-out image 872A displayed in the user interface 88 can be referred to as the "twelfth object", and the double-click operation of the user on the object 871A in the user interface 87 can be referred to as the "third operation".
[0224] Figure 13Exemplarily shown is the specific style of the matte object and the generated mask determined by the electronic device in response to the operation after the user clicks on object 871A when the performance of the object detection algorithm and the matte algorithm is different.
[0225] As Figure 13 shown, both user interface 92 and user interface 94 can be the interfaces displayed by the electronic device after the user clicks on object 871A in the aforementioned user interface 87.
[0226] In the case where the performance of the object detection algorithm and the matte algorithm used by the electronic device is insufficient to support the recognition and extraction of a single object from overlapping objects, after the user clicks on the target object 871A, the electronic device can display the user interface 92 as shown in (A) of Figure 13 . In user interface 92, the electronic device will determine the target object 871A (including target object 871A1 and target object 871A2) as the matte object. Correspondingly, the electronic device will also generate a tracking box for the target object 871A, extract the image corresponding to the target object 871A to obtain the matte image 852B, and pre-display it in the editing box 8721. Due to the limitations of the object detection algorithm and the matte algorithm performance, the electronic device can only recognize the entire target object 871A as one object and cannot recognize target object 871A1 or target object 871A2 as a separate object. Therefore, when generating the matte, the electronic device can only generate the corresponding mask for the target object 871A. For details, refer to the mask1 shown in (a) of Figure 13 .
[0227] In the case where the performance of the object detection algorithm and the matte algorithm used by the electronic device can support the recognition and extraction of a single object from overlapping objects, after the user clicks on object 871A, if the contact point of the user's click on the target object 871A with the screen falls on target object 871A2, the electronic device can display the user interface 93 as shown in (B) of Figure 13 . In user interface 93, the electronic device can only determine target object 871A2 (excluding target object 871A1) as the matte object. Correspondingly, the electronic device will also generate a tracking box for the target object 871A2, extract the image corresponding to the target object 871A2 to obtain the matte image 852B2, and pre-display it in the editing box 8721. At this time, target object 871A2 can be recognized as a separate object. Therefore, when generating the matte, the electronic device can generate the corresponding mask for the target object 871A2. For details, refer to the mask2 shown in (b) of Figure 13 . Similarly, if the contact point of the user's click on the target object 871A with the screen falls on target object 871A2, the electronic device can display the user interface as shown in Figure 13The user interface 94 shown in (C) in []. In the user interface 94, the electronic device can determine only the target object 871A1 (excluding the object 871A2) as the object to be cut out. Correspondingly, the electronic device also generates a tracking frame for the target object 871A2, cuts out the image corresponding to the target object 871A1 to obtain the cut-out image 852B1, and pre-displays it in the editing frame 8721. At this time, the target object 871A1 can be recognized as a separate object. Therefore, when cutting out the image, the electronic device can generate a corresponding mask for the target object 871A1. For details, please refer to Figure 13 the mask 3 shown in (c) in [].
[0228] Of course, to improve the efficiency of determining the object to be cut out, when the viewfinder contains multiple similar objects to be cut out expected by the user (for example, the viewfinder contains multiple kinds of flowers, and these flowers are recognized as dozens of target objects with the category of "flowers" by the electronic device in the viewfinder), the user can also input a prompt text in the interface to determine all the target objects belonging to the category corresponding to the text in the viewfinder as the objects to be cut out with one key.
[0229] As Figure 14 shown in (A) in [], the user interface 95 can be a multi-window interface displayed by the electronic device after activating the camera, which includes a first area 951 and a second area 952. The first area 951 is a viewfinder, which is used to display the preview image obtained by the real-time view of the camera on the electronic device. The preview image shows the target object 951A, the target object 951B, the target object 951C, the target object 951D, and the target object 951E. Among them, the target object 951A is a cone, the target objects 951B, 951C, and 951E are spheres, and the target object 951D is a cube. The second area 952 contains an editing frame 9521. Here, it is assumed that the target detection algorithm stipulates that the categories of target objects include spheres, cubes, cylinders, and cones. After activating the camera, since the center point of the first area falls on the target object 951A, the electronic device can automatically determine the target object 951A as the object to be cut out, cut out the corresponding image, and pre-display it in the editing frame 9521, that is, the image 952A shown in the editing frame 9521.
[0230] At this time, if the object to be cut out expected by the user is not the target object 951A, but all the spheres in the first area 951, that is, the target objects 951B, 951C, and 951E, the user can perform a user operation, such as a long press on any area in the first area 951 in the user interface 95 ( Figure 14(not shown in the figure) to display the text box 953 in the user interface 95. The user can input corresponding keywords or key sentences in the text box, so that the electronic device determines all target objects corresponding to the type of the keywords or key sentences as the matte objects. It should be particularly noted that the keywords or key sentences input in the text box 953 should be the image category information specified by the object detection algorithm, such as one or more of the aforementioned sphere, cube, cylinder, and cone.
[0231] As Figure 14 shown in (A) of the figure, the user can input the keyword "sphere" in the text box 953, and then the electronic device can respond to this operation, determine all the objects in the first region 951 that appear as spheres as the matte objects, and display the user interface 96 as shown in Figure 14 (B) of the figure. In the user interface 96, the electronic device can determine the target objects 951B, 951C, and 951E as the matte objects. Correspondingly, the electronic device also generates a tracking box for each of these target objects, and extracts the images corresponding to the target objects 951B, 951C, and 951E to obtain the matte images 952B, 952C, and 952E, and pre - display these three matte images in the editing box 8721.
[0232] In some embodiments, the text input by the user may not exactly match the category of the target object specified by the aforementioned object detection model. For example, when the text input by the user is "sphere", but the category of the target object specified by the object detection model is "ball", the electronic device can utilize the principle of fuzzy query and automatically screen the categories of the objects in the picture according to the keywords input by the user and the synonyms of the keywords. In this way, all objects whose object categories are the same as or similar to the keywords will be determined as the matte objects. For example, when the keyword input by the user is "sphere", the electronic device can automatically determine words such as "ball" and "spherical object" as the synonyms of the keyword "sphere", and then the electronic device can determine all the objects in the picture whose categories are "sphere", "ball", and "spherical object" as the matte objects.
[0233] Not limited to this, the electronic device can also comprehensively judge based on other types of interactions, preferences, portraits of its users, the position of the target object in the viewfinder, the positional relationship between target objects, the type of the upper - layer editing box, the device type, time, location, etc., and take one object or multiple objects in the viewfinder as the matte objects. For example, the user can also perform positive selection or negative selection on the target objects in the viewfinder by drawing lines, box - selecting, etc., which will not be elaborated one by one here.
[0234] Next, in combination with Figure 15 and Figure 16Introduce the process of an electronic device determining a corresponding matting strategy for different types of matting objects.
[0235] Figure 15 A method for determining a matting strategy provided by an embodiment of the present application. It should be noted in advance that this method is executed when the electronic device has already determined the matting object, and it may include the following steps:
[0236] S1501. Determine the category of the matting object.
[0237] The electronic device determines the category of the matting object. It should be noted that the "category" mentioned here may not be the "category" defined by the object detection algorithm in the foregoing description. It may be a category determined by the electronic device according to the image content corresponding to the matting object and the purpose of the user for the matted image. In the embodiment of the present application, the category of the matting object may include sensitive / prohibited objects, and portraits; in the case where the category of the matting object does not include sensitive / prohibited objects and is not a portrait, the matting object may be further divided into a matting object requiring low-latency operation and a matting object not requiring low-latency operation according to its purpose.
[0238] S1502. Whether the matting object is a sensitive object or a prohibited object.
[0239] S1503. Output a second prompt message.
[0240] In the embodiment of the present application, the electronic device may first determine whether the matting object is a sensitive object or a prohibited object. If the matting object is a sensitive object or a prohibited object, the electronic device may execute step S1503 and will not execute the subsequent steps S1504 - S1509. That is, the electronic device outputs the above-mentioned second prompt message to prompt the user to manually select the matting object or adjust the shooting parameters. For details, reference may be made to the relevant description of Figure 7 which is not elaborated here.
[0241] Specifically, the above-mentioned sensitive objects or prohibited objects may include, but are not limited to, registered trademarks, one-dimensional codes or two-dimensional codes, and images containing user privacy information (such as ID card numbers, property ownership certificates, household registers).
[0242] S1504. Whether the matting object is a portrait.
[0243] Specifically, the electronic device may pre-specify some specific categories of matting objects. For these specific categories of matting objects, the electronic device may use a matting algorithm exclusive to this category for matting, while for matting objects that do not belong to these specific categories, the electronic device may use a unified matting algorithm to perform matting on them.
[0244] For example, in the embodiments of the present application, the electronic device may determine a human portrait as a specific type of matte object category. When the matte object is a human portrait, the electronic device may perform the subsequent step S1505; otherwise, the electronic device may perform the subsequent step S1506.
[0245] S1505: Generate a mask using a human portrait matting algorithm.
[0246] In the case where the matte object is a human portrait, the electronic device may use an algorithm dedicated to matting human portraits to generate the mask required during the matting process.
[0247] It should be understood that users generally have higher requirements for the matting quality of human portraits. If the matte object is a human portrait, since it is very likely to contain the hair of the human body, in this case, the electronic device may use a better-performing matting algorithm (a hair-level matting algorithm) to generate the corresponding mask. Specifically, the human portrait matting algorithm may be a matting algorithm based on the P3Mnet model.
[0248] Here, a brief introduction is given to the matting principle of the human portrait matting algorithm adopted in the present application. After determining the matte object, the electronic device may first obtain a partial image of the area where the matte object is located (this area completely contains the matte object, specifically, it may be an external rectangle determined by the tracking box corresponding to the matte object) from the viewfinder image displayed in the viewfinder. For this partial image, it may consist of a foreground and a background. The area where the matte object is located (the region of interest for the matting algorithm) is the foreground (such as a human portrait). The most important operation of the matting algorithm is to separate the foreground and the background of this partial image. When determining the background and the foreground, the formula used by the matting algorithm adopted in the present application may be specifically as follows:
[0249] I = α i F i + (1 - α i )B i ;
[0250] Among them, i represents the pixel index, that is, the number of the pixel point in the image, F represents the foreground, B represents the background, and α is a predicted probability value of the pixel point being the foreground by the matting algorithm, and its value range is 0 to 1 (the value of α in the human portrait segmentation task is 0 or 1).
[0251] The electronic device may predict the α i value of each pixel point in this partial image based on the matting algorithm. For example, if the α iis 0.9, then the electronic device can consider that the probability of this pixel being the foreground is 90%. Then when the electronic device generates the corresponding mask of the cutout object, the higher the probability of a certain pixel being the foreground, the higher the transparency of the corresponding pixel in the generated mask. For example, when the α of a pixel is i is 1, then the electronic device can consider that the probability of this pixel being the foreground is 100%, and the grayscale value of this pixel in the mask is 255, which is completely transparent; correspondingly, when the α of a pixel is i is 1, then the electronic device can consider that the probability of this pixel being the foreground is 0%, and the grayscale value of this pixel in the mask is 0, which is completely opaque.
[0252] For the hair part of the portrait, since the hair is on the outer contour of the portrait, the electronic device will adjust the α value of the pixel points in these areas. i Most of the predicted values are within the range of (0, 1), and in the generated mask, the closer the pixels in the hair part are to the periphery, the higher the transparency will be. Then, after the above mask is superimposed on the original image to cut out the foreground image, that is, the cutout image in the above description, the electronic device can successfully retain the hair part in the image while improving the rendering effect of the person's hair in the cutout image, avoiding the outer contour of the cutout image being too rigid.
[0253] Of course, in some embodiments, in addition to portraits, the electronic device can also set exclusive cutout algorithms for other types of images. For example, the cutout algorithms supported by the electronic device can also include a special algorithm for cutting out "flower" type cutout objects (hereinafter referred to as "flower type cutout algorithm"). When the cutout object is a flower, the electronic device can use the "flower type cutout algorithm" to generate a mask for the cutout object. In this case, when the cutout object is neither a portrait nor a flower type image, the electronic device will execute the subsequent step S1506, that is, use a general non-portrait cutout algorithm to generate a mask for the cutout object.
[0254] S1506: Generate a mask using a non-portrait cutout algorithm.
[0255] When the object to be matte-extracted is not a human portrait, regardless of which category the object belongs to, the electronic device can use the same matte-extraction algorithm (i.e., the non-human portrait matte-extraction algorithm) to generate the mask required during the matte-extraction process. It should be noted that the non-human portrait matte-extraction algorithm mentioned here is just a name different from the aforementioned human portrait matte-extraction algorithm, and it can specifically be any matte-extraction algorithm with a lower fineness than the aforementioned human portrait matte-extraction algorithm. For example, the matte-extraction algorithm based on the PFAN model. This main matte-extraction algorithm can support matte-extracting objects to be matte-extracted of types such as human portraits, animals, plants, buildings, vehicles, food, commodities, anime cartoon characters, graphic logo signs, etc.
[0256] S1507. Whether the matte-extraction process requires low latency.
[0257] S1508. Perform refined post-processing on the mask generated by the non-human portrait matte-extraction algorithm.
[0258] When the object to be matte-extracted is not a human portrait, the electronic device can determine whether the matte-extraction process requires low latency according to the content and / or use of the matte-extracted image.
[0259] When the matte-extraction process requires low latency, the electronic device can directly execute step S1509, that is, directly use the mask generated by the non-human portrait matte-extraction algorithm to matte-extract the object to be matte-extracted.
[0260] When the matte-extraction process does not require low latency, the electronic device can execute step S1508, that is, perform refined post-processing on the mask generated by the non-human portrait matte-extraction algorithm. The refined post-processing can include, but is not limited to, one or more operations such as image closing operation and image mean filtering. Performing refined post-processing on the mask will make the time required for the matte-extraction operation longer, but at the same time it can also make the mask more refined. Although it may not reach the fineness of the mask obtained by using the human portrait matte-extraction algorithm, it can still make the contour of the matte-extracted image smoother and more regular.
[0261] For example, the electronic device can determine the specific application or specific application scenario to which the edit box for receiving the matte-extracted image belongs. When the edit box is the edit box in a note, the electronic device can determine that the matte-extraction process does not need to have low latency. Therefore, after performing refined post-processing on the mask generated by the non-human portrait matte-extraction algorithm, the electronic device can use the obtained mask for matte-extraction; when the edit box is the edit box in a browser (for example, the edit box for receiving the user-input image during image search), the user may need to quickly obtain the searched information data in this scenario. Then the electronic device can determine that the matte-extraction process needs to have low latency. Therefore, the electronic device can directly use the mask generated by the non-human portrait matte-extraction algorithm for matte-extraction.
[0262] S1509. Obtain a matte image corresponding to the object to be matted based on the generated matte.
[0263] Figure 16 The matte images obtained by the electronic device for different objects to be matted based on the above three matte extraction algorithms are shown.
[0264] As Figure 16 shown in (A) therein, the user interface 97 may be a multi-window interface displayed by the electronic device after the user activates the camera in the editing box provided by the note application, which includes a first area 971 and a second area 972. The first area 971 is a viewfinder for displaying the preview image obtained by the real-time view of the camera on the electronic device, in which the target object 971A is displayed, and the image corresponding to the object 971A is a portrait. The second area 972 includes an editing box 9721. After determining that the target object 971A is the object to be matted, the electronic device can recognize that the object to be matted is a portrait. Therefore, the electronic device can use a hair-level matte extraction algorithm to generate a corresponding matte for the object to be matted, use the matte to perform matte extraction to obtain a matte image 972A, and pre-display the matte image 972A in the editing box 9721. It should be noted that in order to more clearly reflect the matte extraction effect of the matte image, although the matte image 972A is not blurred in the user interface 95, it is still pre-displayed in the editing box 9721. The same applies hereinafter.
[0265] It can be seen from the target object 971A in the first area and the matte image 972A in the second area 972 that the matte image obtained by using the hair-level matte extraction algorithm can successfully retain extremely small main parts in the portrait, such as human hair.
[0266] Figure 16 Shown in (B) therein are the matte images obtained by the electronic device for matte extraction based on the matte extraction algorithm with refined post-processing (that is, first obtaining a matte based on the aforementioned non-portrait matte extraction algorithm, and then performing refined post-processing on the matte, and using the obtained matte for matte extraction. The same applies hereinafter). As Figure 16As shown in (B) thereof, the user interface 98 may be a multi-window interface displayed by the electronic device after the user activates the camera in the editing box provided by the note application, which includes a first area 981 and a second area 982. The first area 981 is a viewfinder, which is used to display the preview image obtained by the real-time view of the camera on the electronic device, in which a target object 981A is displayed, and the image corresponding to the target object 981A is an image of a dog. The second area 982 contains an editing box 9821. After determining that the target object 981A is a matte object, the electronic device can recognize that the matte object is not a human figure, and the editing box 9821 is an editing box in the note application. The entire matte process is likely to have relatively low latency requirements. Therefore, the electronic device can use a matte algorithm with refined post-processing to generate a corresponding mask for the matte object, use the mask to perform matting to obtain a matted image 982A, and pre-display the matted image 982A in the editing box 9821.
[0267] It can be understood that, as shown by the target object 981A, there are also many hairs on the outer contour of the dog. However, although the matte algorithm with refined post-processing used in this application is superior to the human figure matte algorithm in terms of power consumption and latency, its matte effect (i.e., matte accuracy) is far inferior to the aforementioned human figure matte algorithm. As can be seen from the target object 981A in the first area and the matted image 982A in the second area 982, there are actually many serrated lines formed by the dog's hairs on the outer contour of the target object 981A. However, the entire outer contour of the matted image 982A obtained by using a non-hair-level matte algorithm with refined post-processing fails to retain these serrated contours formed by the hairs.
[0268] Figure 16 (C) therein shows the matted image obtained by the electronic device performing matting based on a matte algorithm without refined post-processing (i.e., an algorithm that directly uses the aforementioned non-human figure matte algorithm to obtain a mask for matting, the same below). As Figure 16As shown in (C) in [reference], the user interface 99 may be a multi-window interface displayed by the electronic device after the user activates the camera in the editing box provided by the browser application, which includes a first area 991 and a second area 992. The first area 991 is a viewfinder, which is used to display the preview image obtained by the real-time view of the camera on the electronic device, and a target object 991A is displayed therein. The image corresponding to the object 991A is an image of a car. The second area 992 contains an editing box 9921. After determining that the target object 991A is a matte object, the electronic device can recognize that the matte object is not a portrait, and the editing box 9921 is an editing box in the browser search interface. The entire matte process is likely to have high requirements for latency. Therefore, the electronic device can use a matte algorithm without refined post-processing to generate a corresponding mask for the matte object, use the mask to perform matting to obtain a matted image 992A, and pre-display the matted image 992A in the editing box 9921.
[0269] It can be understood that although the accuracy of the matte algorithm without refined post-processing is relatively low, the latency of the electronic device to implement this algorithm is low. For the matted image used for image search, the electronic device only needs to identify the general outline of the object to be searched to quickly provide the searched information for the user. As shown in the matted image 992A, although the matted image obtained by the matte algorithm without refined post-processing used in this application cannot accurately reflect the outline of the matte object, it has low requirements for the performance of the electronic device, and the electronic device can quickly perform matting on the matte object to obtain a matted image for the user to use.
[0270] Optionally, when there are multiple matte objects in the preview image, the electronic device can use the matte algorithms corresponding to the respective matte objects according to the categories of the respective matte objects according to the above method to obtain the matted images of the respective matte objects. For example, when a portrait, a dog, and a vehicle are simultaneously displayed in the viewfinder, and all three objects (there is no overlapping part between the three objects) are determined to be matte objects, the electronic device can use a portrait matte algorithm to perform matting on the portrait, use a matte algorithm with refined post-processing to perform matting on the dog, and use a matte algorithm without refined post-processing to perform matting on the vehicle; or, the electronic device can start from the perspective of low power consumption and uniformly use a matte algorithm without refined post-processing to perform matting on the portrait, the dog, and the vehicle; or, the electronic device can also start from the perspective of accurate matting and uniformly use a portrait matte algorithm to perform matting on the portrait, the dog, and the vehicle. This application does not make any limitations in this regard.
[0271] In some embodiments, if there is an overlapping part between multiple objects (for example, a person holding a cat), the electronic device may identify it as a target object. If such a target object is determined to be a matte object, in this case, the electronic device can determine the matte algorithm applicable to the matte object based on the category discrimination of the matte object in step S1501, and determine the matte algorithm applicable to the matte object according to its category. For example, when the image corresponding to the matte object is a person holding a cat, if the electronic device determines that the matte object is a portrait, the electronic device can use a portrait matte algorithm to matte it; if the electronic device determines that the matte object is a cat, the electronic device can use a matte algorithm without refined post-processing or a matte algorithm with refined post-processing to matte it.
[0272] Next, a method for the electronic device to automatically trigger an adjustment operation on the matte object displayed in the edit box will be described in conjunction with Figures 17 - 24 the following.
[0273] In some embodiments, before the user officially inputs the matte image into the edit box, the electronic device can track the determined matte object in real time. For each frame of the picture containing the target object displayed in the viewfinder, the electronic device can determine the matte object for this picture, and extract the image corresponding to the matte object, and use the matte images extracted from each frame of the image to refresh the matte image pre-displayed in the edit box in real time. For example, when the camera transmits 30 frames of images per second to the viewfinder for display, the electronic device will sequentially obtain these 30 frames of images. For each frame of image obtained, the electronic device will re-determine the matte image in this frame of image automatically according to the method described above. After extracting the image to obtain the matte image, the electronic device will pre-display the matte image in the edit box. That is to say, when the performance of the electronic device is good enough, the matte image displayed in the edit box in the subsequent 1 s actually includes 30 frames, and these 30 frames of images actually come from the 30 frames of images displayed in the viewfinder in the past 1 s.
[0274] In some embodiments, since the matting process imposes a certain burden on the computing power and power consumption of the electronic device, the electronic device can use a tracking algorithm (such as a machine learning-based tracking algorithm, such as single-object tracking or multi-object tracking based on a Siamese network) to determine whether the tracking of the matting object is lost. Specifically, the electronic device can periodically obtain the images in the viewfinder and determine whether the change range of the determined matting object in the viewfinder is within a certain range. If the change range of the matting object in the viewfinder is within a certain range, the electronic device can consider that it still maintains the tracking state of the matting object; otherwise, it is considered that the tracking of the matting object is lost. When the tracking of the matting object is lost, the electronic device will re-determine the matting object, re-mat the matting object, and then update the historical pre-displayed matting image in the editing box with the obtained matting image.
[0275] Figure 17 The figure shows a method for the electronic device to trigger a re-matting operation, which may include the following steps:
[0276] S1701. Obtain multiple frames of images from the images captured per second.
[0277] The electronic device periodically obtains the images displayed in the viewfinder, that is, periodically obtains images from the images captured by the camera. It should be noted that when performing this step, the electronic device has already determined a matting object from the historical display images in the viewfinder, and currently is displaying the matting image corresponding to the matting object in the editing box.
[0278] In this method, the electronic device does not need to obtain all the frames of images captured by the camera. It can select several frames of images from the images captured by the camera per second. For example, 5 frames are extracted from the images captured by the camera every 1 second (usually the camera outputs 30 frames of image data per second). The specific extraction method can be centralized extraction from the images captured by the camera per second (such as extracting the 1st - 5th frames or the 26th - 30th frames from the 30 frames of images captured per second), or interval extraction from the images captured by the camera per second (such as extracting the 6th, 12th, 18th, 24th, and 30th frames from the 30 frames of images captured per second). This application does not make any limitations in this regard. In addition, among the above-mentioned multiple frames of images, any one frame of image may not be a frame of image completely displayed in the viewfinder. It can be a partial image intercepted by the electronic device from the rectangular area formed by the tracking frame of the matting object in each extracted frame of image.
[0279] S1702. Determine whether the similarity of the matting object in every two adjacent frames of images is less than a first threshold.
[0280] S1703. Determine whether the similarity of the matting object in the last frame of image and the matting object in the first frame of image in every continuous N frames of images is less than a first threshold.
[0281] After obtaining the above-mentioned multiple frames of images, the electronic device can arrange the above-mentioned multiple frames of images in order according to the generation time of the images (or the time displayed on the screen). After that, the electronic device can start from the first frame of image and compare the similarity of the object to be cropped in each adjacent two frames of images in turn to determine whether the similarity of the object to be cropped included in the adjacent two images is less than the above-mentioned first threshold; in addition, for the above-mentioned multiple frames of images, the electronic device will also take every consecutive N frames of images as a group and determine whether the similarity of the object to be cropped in the last frame of image in each group and the object to be cropped in the first frame of image in this group is less than the first threshold.
[0282] Taking N as 6 above, and taking the example that the electronic device selects 5 frames of images from the images captured by the camera every second, Figure 18 It shows the process in which the electronic device compares the similarity of the object to be cropped in each adjacent two frames of images, and compares the similarity of the object to be cropped in the last frame of image in each group and the object to be cropped in the first frame of image in this group.
[0283] As Figure 18 shown in (A) in
[0284] After that, as Figure 18As shown in (B) of , the electronic device can arrange the above-mentioned multiple frames of images including pic1 - pic5 and pic6 - pic10 in order according to the generation time of the images (or the time displayed on the screen). Then, starting from the first frame of the image, the electronic device can compare the similarity of the cropped objects in each adjacent two frames of images in turn. That is, the electronic device can compare the similarity of the cropped objects included in pic1 and pic2; if the similarity of the cropped objects included in pic1 and pic2 is greater than or equal to the above-mentioned first threshold, the electronic device can then compare the similarity of the cropped objects included in pic2 and pic3; if the similarity of the cropped objects included in pic2 and pic3 is greater than or equal to the above-mentioned first threshold, the electronic device can then compare the similarity of the cropped objects included in pic3 and pic4, and so on. If until the electronic device compares the similarity of the cropped objects included in pic5 and pic6 and determines that the similarity of the cropped objects included in pic5 and pic6 is still greater than or equal to the above-mentioned first threshold, then at this time, the electronic device can compare the similarity of the cropped object of the last frame of the first group of images (i.e., pic6) and the cropped object of the first frame of this group of images (i.e., pic1). Similarly, the electronic device can continue to complete the comparison of the similarity of the cropped objects in the subsequent images in this way.
[0285] In the embodiments of the present application, the above-mentioned multiple frames of images including pic1 - pic5 and pic6 - pic10 can be referred to as "multiple frames of pictures, the obtained multiple frames of pictures", and pic1 - pic5 and pic6 - pic10 can be referred to as "N consecutive frames of pictures". pic1 and pic2 can be respectively referred to as "the first picture" and "the second picture", or pic1 and pic6 can be respectively referred to as "the first picture" and "the second picture".
[0286] During the entire comparison process, if there are any two adjacent images, and the similarity of the cropped objects they contain is less than the above-mentioned first threshold, it means that the state of the cropped object in the viewfinder has changed suddenly. Therefore, the electronic device can execute the subsequent step S1704. If in a group of images, among the consecutive N images it contains, the similarity of the cropped object of the last frame of the image and the cropped object of the first frame of the image is less than the first threshold, it means that the state of the cropped object in the viewfinder has changed cumulatively. Therefore, the electronic device can also execute the subsequent step S1704.
[0287] It should be noted that in this application, the electronic device determines the similarity between two cutout objects in two images based on the feature points included in the two cutout objects. In image processing, a "feature point" is a set of points analyzed by an algorithm, which refers to a set of points that can represent an image or an object in an identical or at least very similar invariant form in other similar images containing the target. In combination with the two cutout objects required by this application, that is, if the feature points of the two cutout objects are the same or similar, then the similarity between the two cutout objects is relatively high, and the electronic device can consider that the tracking of the cutout object has not been lost.
[0288] However, such feature points will not change significantly due to the object being simply scaled proportionally and reduced, or the change in the position of the object in the entire image. Therefore, in the two pictures displayed in the viewfinder, even if the cutout objects in the two pictures show a scaled state, or the imaging position of the cutout object in the viewfinder changes, the electronic device may still determine that the cutout objects in the two pictures are similar. Specifically, it may include but is not limited to the following situations:
[0289] ① The user changes the distance between the camera and the physical entity corresponding to the cutout object, or adjusts the camera focal length, resulting in the cutout object being scaled.
[0290] As Figure 19 shown in (A) below, the user interface 01 can be a multi-window interface displayed by the electronic device after the user activates the camera in the editing box provided by the note application. It includes a first area 011 and a second area 012. The first area 011 is the viewfinder, which is used to display the preview picture obtained by the camera on the electronic device in real time. A cutout object 011A is displayed therein. The second area 012 contains an editing box 0121. At this time, the camera is at a position with a distance d1 from the cube (the cube is the physical entity corresponding to the cutout object 011A, the same below), and uses a focal length of "2×" (i.e., twice the focal length) to take a picture of the cube, and the taken picture is displayed in the first area 011 in real time. At this time, the electronic device has already cut out the image corresponding to the cutout object 011A and pre-displays it in the editing box 0121, that is, the image 012A shown in the editing box 0121.
[0291] If the user changes the relative position relationship between the camera and the above cube by moving the camera, for example Figure 19Move the camera to a position where the physical distance from the object to be cut out is d2 as shown in (B) in [description], and still use a focal length of "2×" to take a picture of it. The electronic device can display the user interface 02. In the user interface 02, it includes a first area 021 and a second area 022. The first area 021 is a viewfinder, which is used to display the preview image obtained by the real-time viewfinder of the camera on the electronic device. Since the electronic device will always track the object to be cut out 011A, even if the object to be cut out 011A in the first area 011 is scaled to the target object 021A in the first area 021, before the electronic device performs the similarity comparison, it will still determine the target object 021A as the object to be cut out and generate a corresponding tracking frame 0211 for it.
[0292] After that, the electronic device can further determine the similarity between the object to be cut out 021A and the object to be cut out 011A. Figure 19 (C) in [description] shows the image 011B obtained by the electronic device completely cropping out the corresponding image of the object to be cut out 011A according to the tracking frame 0111 in the user interface 01, and the image 021B obtained by completely cropping out the corresponding image of the object to be cut out 021A according to the tracking frame 0211 in the user interface 02. It can be understood that the images 011B and 021B can be two consecutive frames among the multiple frames of images obtained from the images taken by the camera every second, such as Figure 18 pic1 and pic2 in [description]. And as common sense knows, when the shooting focal length remains fixed, if the values of d2 and d1 are different, the area of the object to be cut out 011A is also different from the area of the object to be cut out 021A. However, if the electronic device performs a similarity detection on the object to be cut out 011A and the object to be cut out 021A contained in the obtained images 011B and 021B at this time, the electronic device can determine that the feature points contained in the object to be cut out 011A and the object to be cut out 021A are basically matched. Therefore, the similarity between the object to be cut out 011A and the object to be cut out 021A is greater than or equal to the above first threshold. So the electronic device can consider that the object to be cut out 021A and the object to be cut out 011A are the same object to be cut out, and the tracking of the original object to be cut out 011A has not been lost. Then the electronic device does not trigger the process of re-determining the object to be cut out, nor will it re-cut the object to be cut out 021A and refresh the newly obtained cut-out image to the editing box 0221.
[0293] That is to say, the cut-out image 022A in the user interface 02 is actually the cut-out image 012A in the user interface 01. Optionally, the cut-out image 022A in the user interface 02 can be obtained by the electronic device scaling the cut-out image 012A in the user interface 01 to a certain extent and then displaying it in the user interface 02; specifically, the scaling ratio can be determined according to the ratio of the area of the cut-out object 011A in the first region 011 to the area of the cut-out object 021A in the first region 021, the ratio of the areas of the rectangular regions respectively framed by the tracking frame 0111 and the cut-out tracking frame 0211, the ratio of the diagonal lengths of the rectangles respectively framed by the tracking frame 0111 and the cut-out tracking frame 0211, etc. This application does not limit this.
[0294] Taking the ratio of areas as an example for illustration, it is assumed here that the area of the cut-out object 011A in the first region 011 is S 011A , and the area of the cut-out object 021A in the first region 021 is S 021A , and S 021A / S 011A = N1 (N1 > 0), then the electronic device can scale the cut-out object 011A by N2 times according to the specific value of N1 to obtain the cut-out image 012A (N2 > 0, and when N2 < 1, it means shrinking the cut-out object 011A, and when N2 > 1, it means enlarging the cut-out object 011A). Specifically, N2 can be equal to N1; or, N2 can be positively correlated with N1, that is, the larger N1 is, the larger N2 is (for example, when N1 = 2, N2 = 1.2; when N1 = 3, N2 = 1.3). The mapping relationship between the values of N1 and N2 can be stored in the electronic device. After obtaining the above N1, the electronic device can obtain the specific value of the corresponding N2, scale the cut-out object 011A by N2 times to obtain the cut-out image 012A, and display it in the user interface 02.
[0295] In one embodiment, when the cut-out object remains unchanged, it can be set that only when N1 is greater than a preset threshold will the scaling of the cut-out image be triggered, so as to avoid the change of the cut-out image being triggered by the user's slight adjustment of the electronic device and affecting the user experience.
[0296] In addition, the user can also cause the cut-out object to be scaled by adjusting the focal length of the camera. Similarly, taking Figure 19Taking the user interface 01 shown in (A) as an example, in combination with the foregoing description, it can be known that at this time, the camera is at a position with a distance of d1 from the cube, and uses a focal length of "2×" (i.e., twice the focal length) to photograph the cube, and the photographed image is displayed in real time in the first area 011 of the user interface 01. If at this time, the user adjusts the focal length of the camera from "2×" to "1×" (i.e., single focal length) by performing relevant operations on the focal length adjustment control 0113 in the first area 011, although the camera is still at a position with a distance of d1 from the cube at this time, the imaging area of the cube in the electronic device will also change. Specifically, reference can be made to Figure 19 the user interface 03 shown in (D) thereof.
[0297] The user interface 03 includes a first area 031 and a second area 032. The first area 031 is a viewfinder, which is used to display the preview image obtained by the camera on the electronic device in real time, and a matte object 031A is displayed therein. The second area 032 includes an edit box 0321. Similarly, before the electronic device performs the similarity comparison, it will determine the object 031A as the matte object and generate a corresponding tracking box 0311 for it.
[0298] After that, the electronic device can further determine the similarity between the matte object 031A and the matte object 011A. After determining that the feature points contained in the matte object 011A and the matte object 031A are basically matched, the electronic device can determine that the similarity between the matte object 011A and the matte object 031A is greater than or equal to the above-mentioned first threshold. Therefore, the electronic device can consider that the matte object 031A and the matte object 011A are the same matte object, and the tracking of the original matte object 011A has not been lost. Then the electronic device does not trigger the process of re-determining the matte object, nor does it re-matte the matte object 031A.
[0299] Correspondingly, the electronic device can scale the matte image 012A in the user interface 01 to a certain extent according to the ratio of the area of the matte object 011A in the first area 011 to the area of the matte object 031A in the first area 031, and then display it in the user interface 03.
[0300] In an embodiment of the present application, the user interface 01 may be referred to as the "second interface", the object 011A may be referred to as the "first object", the screen displayed in the first area 011 may be referred to as the "first screen", the matte image 012A may be referred to as the "second object", the screen displayed in the second area 021 (the third area 031) may be referred to as the "second screen", the matte object 021A (the matte object 031A) may be referred to as the "first changing object", and the matte image 022A (the matte image 032A) may be referred to as the "third object". The ratio of the area of the matte object 021A (the matte object 031A) to the area of the matte object 011A may be referred to as the "second ratio value", the ratio of the area of the matte image 022A (the matte image 032A) to the area of the matte image 012A may be referred to as the "first ratio value", and the physical entities corresponding to the objects 011A, 021A, and 031A may be referred to as the "first entities". An operation in which the user changes the distance between the camera and the entity corresponding to the matte object or adjusts the camera focal length, resulting in the scaling of the matte object, may be referred to as the "fifth operation".
[0301] ② When the entity corresponding to the matte object moves in the scene, the imaging position of the matte object in the viewfinder changes, but the size of the matte object remains unchanged.
[0302] As Figure 20 As shown in (A) of
[0303] While keeping the distance between the camera and the cube still at d1, the user can slightly move the camera. For example, the user can translate the camera downward by a small distance to change the display position of the cube on the electronic device, and make the electronic device display as Figure 20The user interface 05 shown in (B) of []. The user interface 05 includes a first area 051 and a second area 052. The first area 051 is a viewfinder frame, which is used to display the preview image obtained by the real-time viewfinder of the camera on the electronic device. Since the electronic device will always track the object to be cut out 041A, even if the position of the object to be cut out 041A in the first area 041 moves upward and is displayed as the object 051A in the first area 051, before the electronic device performs the similarity comparison, it will still determine the object 051A as the object to be cut out and generate a corresponding tracking frame for it.
[0304] After that, the electronic device can further determine the similarity between the object to be cut out 051A and the object to be cut out 041A. When performing the similarity detection, the electronic device can determine that the feature points contained in the object to be cut out 051A and the object to be cut out 041A are basically matched. Therefore, the similarity between the object to be cut out 051A and the object to be cut out 041A is greater than or equal to the above-mentioned first threshold. So the electronic device can consider that the object to be cut out 051A and the object to be cut out 041A are the same object to be cut out, and the tracking of the original object to be cut out 041A has not been lost. Then the electronic device does not trigger the process of re-determining the object to be cut out, nor will it re-cut the object to be cut out 051A and refresh the newly obtained cut-out image into the editing frame 0521.
[0305] That is to say, the cut-out image 042A in the user interface 04 is actually the cut-out image 052A in the user interface 05. Optionally, the electronic device can determine that the area of the object to be cut out 041A in the first area 041 is the same as the area of the object to be cut out 051A in the first area 051. Therefore, the electronic device does not need to scale the object to be cut out 041A, that is, the area of the cut-out image 052A in the user interface 05 is set to be the same as the area of the cut-out image 042A.
[0306] In some embodiments, when the similarity of the object to be cut out is greater than or equal to the above-mentioned first threshold, the electronic device can also scale the cut-out image pre-displayed in the editing frame according to the distance between the lens and the main body of the object to be cut out; the specific scaling ratio can be obtained through hardware, software, or a combination of hardware and software; for example, the distance between the lens and the main body of the object to be cut out can be calculated through a TOF camera, monocular / multiple-eye distance estimation, etc., or the scaling ratio can be determined by calculating the change ratio before and after the tracking frame corresponding to the object to be cut out.
[0307] S1704. Re-determine the object to be cut out.
[0308] Among multiple frames of images captured by a camera of an electronic device, if there are any two adjacent images whose similarity of the cutout object included is less than the first threshold; or there is any group of images, and among the consecutive N images included, the similarity of the cutout object in the last frame image and the cutout object in the first frame image is less than the first threshold, it indicates that the state of the cutout object in the viewfinder has changed significantly, and thus it is determined that the tracking of the original cutout object is lost. Therefore, the electronic device can re-determine the cutout object, extract the image corresponding to the new cutout object from the screen where the new cutout object is located to obtain a new cutout image. After that, the electronic device can pre-display the new cutout image in the editing box.
[0309] As Figure 21 shown in (A) of Figure 21 , the user interface 06 can be a multi-window interface displayed by the electronic device after the user activates the camera in the editing box provided by the note application. It includes a first area 061 and a second area 062. The first area 061 is a viewfinder, which is used to display the preview screen obtained by the real-time view of the camera on the electronic device, and a cutout object 061A and its corresponding tracking frame 0611 are displayed therein. The second area 062 includes an editing box 0621. At this time, the camera is shooting the cube on the same horizontal line as the cube, and the captured image is displayed in the first area 061 in real time. As can be seen from Figure 21 (A) of
[0310] During the shooting process, the user can change the direction of the field of view angle of the camera by controlling the electronic device to rotate a certain Euler angle (the rotation direction can include one or more of Pitch, Yaw, Roll, and specific reference can be made to the relevant description of Figure 1 ). For example, the user controls the electronic device to rotate the electronic device so that the camera shoots the cube in a top-down posture. At this time, the electronic device can shoot the upper surface of the cube, that is, the surface of the cube containing the digital serial number "③", and the display is as shown in Figure 21The user interface 07 shown in (B) of [the figure]. In the user interface 07, it includes a first area 071 and a second area 072. The first area 071 is a viewfinder, which is used to display the preview image obtained from the real-time view of the camera on the electronic device. Since the electronic device will continuously track the object to be cut out 061A, even if the object to be cut out 061A in the first area 061 has changed significantly and becomes the object 071A in the first area 071, before the electronic device performs the similarity comparison, it will still first determine the object 071A as the object to be cut out and generate a corresponding tracking frame 0711 for it.
[0311] After that, the electronic device can further determine the similarity between the object to be cut out 061A and the object to be cut out 071A. Figure 21 (C) in [the figure] shows the image 061B obtained by the electronic device completely cropping out the corresponding image of the object to be cut out 061A according to the tracking frame 0611 in the user interface 06, and the image 071B obtained by completely cropping out the corresponding image of the object to be cut out 071A according to the tracking frame 0711 in the user interface 07. It can be understood that the image 061B and the image 071B can be two consecutive frames among the multiple frames of images obtained from the images captured by the camera every second, such as Figure 18 pic1 and pic2 in [the figure], or the first frame image and the last frame image among the consecutive N frame images respectively, such as Figure 18 pic1 and pic6 in [the figure]. At this time, when the electronic device performs similarity detection on the object to be cut out 061A and the object to be cut out 071A respectively included in the obtained images 061B and 071B, the electronic device can determine that the feature points contained in the object to be cut out 061A and the object to be cut out 071A do not match, that is, the similarity between the object to be cut out 061A and the object to be cut out 071A is less than the above-mentioned first threshold. Therefore, the electronic device can consider that the object to be cut out 071A and the object to be cut out 061A are the same object to be cut out, that is, the tracking of the original object to be cut out 061A has not been lost. Then the electronic device will trigger the process of determining a new object to be cut out, detect the image displayed in the first area 071 based on the foregoing method, re-determine the object 071A as the object to be cut out, and after re-cutting the object to be cut out 071A, refresh the new cut-out image to the editing box 0721. Specifically, reference can be made to the cut-out image 072A displayed in the editing box 0721.
[0312] In the embodiments of the present application, the user interface 06 may be referred to as the "second interface", the object 061A may be referred to as the "first object", the screen displayed in the first area 061 may be referred to as the "first screen", the matte image 062A may be referred to as the "second object", the screen displayed in the first area 071 may be referred to as the "third screen", the object 071A may be referred to as the "second variable object", the matte image 072A may be referred to as the "fourth object", and the physical entities corresponding to the object 061A and the object 071A may be referred to as the "first entity". An operation in which the user can control the electronic device to rotate a certain Euler angle to change the direction of the field of view angle of the camera may be referred to as the sixth operation.
[0313] In some embodiments, the electronic device may also combine the acceleration sensor and the gyroscope sensor included therein to obtain the attitude and displacement magnitude during the matte extraction process of the electronic device. In the case where the attitude or displacement change during the matte extraction process of the electronic device reaches a certain degree, the electronic device may also trigger a process of re-determining the matte extraction object and re-performing matte extraction based on the new matte extraction object.
[0314] In addition, during the whole process, if the user performs a corresponding user operation, such as clicking on a certain object included in the viewfinder, to switch to the matte extraction object, the electronic device may also respond to this user operation, trigger a process of re-determining the matte extraction object, perform matte extraction based on the newly determined matte extraction object, and refresh the obtained matte image into the editing box.
[0315] In some scenarios, the number of target objects in the viewfinder may increase or decrease, or the specific position of the target object in the viewfinder may change. Consider such a situation where when the user activates the camera through a corresponding user operation, the user does not place the target object in the central area of the shooting range of the camera. However, at this time, there is an interfering object in the central area of the shooting range, and the electronic device can automatically determine the object corresponding to the interfering object as the matte extraction object according to the method described above, and pre-display the corresponding image in the editing box after extracting it. However, after the user places the target object in the central area of the shooting range as expected, the electronic device should re-determine the object corresponding to the above target object in the viewfinder as the matte extraction object and re-perform matte extraction. However, if the previously determined matte extraction object of the electronic device has not changed significantly, the electronic device will not re-determine the matte extraction object, and naturally will not perform matte extraction on the object corresponding to the above target object. The user can only manually click on the object corresponding to the above target object to switch it to a new matte extraction object, and the operation becomes more cumbersome.
[0316] Regarding the above problems, in the foregoing Figure 17Based on the method for automatically triggering the re-matting operation shown above, the present application also provides another method for automatically triggering the re-matting operation. When implementing this method, the electronic device not only detects in real time whether the matting object has changed significantly, but also detects in real time whether the entire content of the viewfinder has changed significantly. The "change" mentioned here includes the change in the number of target objects in the picture and the change in the shape of the target object in the picture. When the matting object changes significantly, or the entire content of the viewfinder changes significantly, the electronic device will trigger the operations of re-determining the matting object and re-matting, reducing user operations while further meeting the user's expectations for the matting effect.
[0317] As Figure 22 shown, this method may include the following steps:
[0318] S2201. Extract multiple frames of images from the images captured per second.
[0319] S2202. Determine whether the similarity between the matting objects in every two adjacent frames of images is less than a first threshold.
[0320] S2203. Determine whether the similarity between the matting object in the last frame and the matting object in the first frame among every consecutive N frames of images is less than the first threshold.
[0321] For the specific content of steps S2201 - S2203, reference can be made to the relevant descriptions of steps S1701 - S1703 in Figure 17 above, which will not be elaborated here.
[0322] S2204. Determine whether the similarity between every two adjacent frames of images is less than a second threshold.
[0323] S2205. Determine whether the similarity between the last frame and the first frame among every consecutive M frames of images is less than the second threshold.
[0324] It should be noted that the above multiple frames of images may include two parts of images, including a first image set and a second image set. In the first image set, any frame of image may not be a frame of image completely displayed in the viewfinder. It may be a partial image intercepted by the electronic device from the rectangular area formed by the tracking frame of the matting object in each frame of the extracted images. This part of the image may be the image used by the electronic device to execute the aforementioned steps S2202 - S2203; while in the second image set, any frame of image is a frame of image completely displayed in the viewfinder. This part of the image may be the image used by the electronic device to execute the aforementioned steps S2204 - S2205. Specifically, the first image set may be obtained by the electronic device taking screenshots of the images in the second image set.
[0325] It can be understood that in step S2204 and step S2205, the electronic device compares not the similarity of the object to be cropped, but the similarity of the entire content of the screen in the screen. After obtaining the above-mentioned second image set, the electronic device can arrange the images in the second image set in order according to the generation time of the images (or the time displayed on the screen). After that, the electronic device can start from the first frame image in the second image set and compare the similarity of every two adjacent frames of images in turn to determine whether the similarity of the two adjacent images is less than the above-mentioned second threshold; in addition, for the images in the second image set, the electronic device will also use every consecutive M frames of images as a group to determine whether the similarity between the last frame image and the first frame image in each group of images is less than the second threshold. Similarly, in the embodiments of the present application, the electronic device can also determine the similarity of two images based on the feature points of the two images. Specifically, the above-mentioned M can be equal to the above-mentioned N, and the above-mentioned second threshold can be equal to the above-mentioned first threshold.
[0326] S2206. Redetect the object to be cropped.
[0327] In the first image set, if there are any two adjacent images whose similarity of the object to be cropped is less than the above-mentioned first threshold; or if there is any group of images in which the similarity of the object to be cropped between the last frame image and the first frame image in the consecutive N images included is less than the first threshold, it means that the state of the object to be cropped in the viewfinder has changed significantly, and thus it is determined that the tracking of the original object to be cropped has been lost. Therefore, the electronic device can re-determine the object to be cropped, and crop out the image corresponding to the new object to be cropped from the screen where the new object to be cropped is located to obtain a new cropped image. After that, the electronic device can pre-display the new cropped image in the editing frame.
[0328] Similarly, in the second image set, if there are any two adjacent images whose similarity is less than the above-mentioned second threshold; or if there is any group of images in which the similarity between the last frame image and the first frame image in the consecutive N images included is less than the second threshold, it means that the image content has changed significantly. Therefore, the electronic device can also re-determine the object to be cropped, and crop out the image corresponding to the new object to be cropped from the screen where the new object to be cropped is located to obtain a new cropped image. After that, the electronic device can pre-display the new cropped image in the editing frame.
[0329] Figure 23 It shows the process in which the electronic device triggers re-determining the object to be cropped due to the change in the position of the target object in the screen shown by the viewfinder, and obtaining a new cropped image based on the re-determined object to be cropped.
[0330] Such as Figure 23As shown in (A) therein, the camera is at the same horizontal line as the cube and the sphere at this time to photograph the two, and the photographed image is displayed in the viewfinder in real time. From Figure 23 As can be seen from (A) of Figure 23 . The user interface 08 may be a multi-window interface displayed on the electronic device after the user activates the above camera in the editing box provided by the note application, and it includes a first area 081 and a second area 082. The first area 081 is a viewfinder, which is used to display the preview image obtained by the camera on the electronic device in real time. A target object 081A and a target object 081B are displayed therein. Since the center point of the first area 081 is closer to the target object 081A, the target object 081A is determined as the object to be cropped, and a corresponding tracking frame is included. The second area 082 includes an editing box 0821. The electronic device has cropped the image corresponding to the target object 081A from the image displayed in the first area 081 at this time, and the obtained cropped image 082A is pre-displayed in the editing box 0821.
[0331] During the shooting process, if the sphere rolls a certain distance to the left, causing the center point of the viewfinder to fall on the object corresponding to the sphere. Then as Figure 23 shown in (B) of Figure 23 , the electronic device can display the user interface 09 at this time, which includes a first area 091 and a second area 092. The first area 091 is a viewfinder, which is used to display the preview image obtained by the camera on the electronic device in real time. The second area includes an editing box 0921.
[0332] It can be understood that although the similarity between the target object 081A and the target object 091A in the image displayed in the first area 091 and the image displayed in the first area 081 is greater than the above first threshold, the electronic device will further determine the similarity between the entire images in the first area 081 and the first area 091. Since the relative position relationship between the target object 091A and the target object 091B in the first area is significantly different from the relative position relationship between the target object 081A and the target object 081B in the first area 081, the electronic device can determine that the feature points contained in the image in the first area 091 and the image in the first area 081 do not match, that is, the similarity between the two images is less than the above second threshold. Therefore, the electronic device can trigger the process of re-determining the object to be cropped, detect the image displayed in the first area 091 based on the foregoing method, and after determining the target object 091B as the object to be cropped, re-crop the target object 091B and pre-display the new cropped image 092B in the editing box 0921.
[0333] In addition, during the entire process, if the user performs a corresponding user operation, such as clicking on an object included in the viewfinder, to switch to the object to be cut out, the electronic device can also respond to this user operation, trigger the process of re-determining the object to be cut out, perform cutting based on the newly determined object to be cut out, and refresh the obtained cut-out image into the editing box.
[0334] In the embodiments of the present application, the user interface 08 may be referred to as the "second interface", the objects 081A and 091A may be referred to as the "first objects", the picture displayed in the first area 081 may be referred to as the "first picture", and the cut-out image 082A may be referred to as the "second object". The picture displayed in the first area 091 may be referred to as the "fourth picture", the object 091B may be referred to as the "fifth object", the cut-out image 092B may be referred to as the "sixth object", the cube corresponding to the objects 081A and 091A may be referred to as the "first entity", and the sphere corresponding to the objects 081B and 091B may be referred to as the "second entity".
[0335] Figure 24 Respectively show the process in which the electronic device triggers the re-determination of the object to be cut out due to the change in the number of target objects in the picture shown in the viewfinder, and obtains a new cut-out image based on the re-determined object to be cut out.
[0336] As Figure 24 shown in (A) of Figure 24 At this time, the camera is shooting the only cube existing within the field of view angle and displaying the shot picture in the viewfinder in real time. As can be seen from (A) of
[0337] During the shooting process, the user places another sphere within the field of view angle of the camera, and the center point of the viewfinder falls on the object corresponding to the sphere.
[0338] If there is an overlapping part between the object corresponding to the sphere and the object corresponding to the aforementioned cube in the viewfinder, then as Figure 24As shown in (B) therein, the electronic device can present a user interface 42 at this time, which includes a first area 421 and a second area 422. The first area 421 is a viewfinder frame, which is used to display the preview image obtained by the real-time view of the camera on the electronic device; the second area includes an editing frame 4221.
[0339] In the image displayed in the first area 421, since there is an overlapping part between the object corresponding to the sphere and the object corresponding to the cube, these two objects are regarded as the same target object 421A in the image. For the image displayed in the first area 421 and the image displayed in the first area 411, the electronic device will further determine the similarity between the two images. Since the image contents are significantly different, the electronic device can determine that the feature points contained in the image displayed in the first area 421 and the image displayed in the first area 411 do not match, that is, the similarity between the two images is less than the above-mentioned second threshold. Therefore, the electronic device can trigger the process of re-determining the matte object, and detect the image displayed in the first area 421 based on the foregoing method. Since the center point of the first area 421 falls on the target object 421A, the target object 421A will be determined as the new matte object. After that, the electronic device will re-matte the target object 421A and then pre-display the new matte image 422A in the editing frame 4221. From Figure 24 As can be seen from (B) therein, the matte image 422A contains both the sphere and the cube images, and there is also an overlapping part between the sphere image and the cube image.
[0340] If there is no overlapping part between the object corresponding to the sphere and the object corresponding to the cube in the viewfinder frame, then as Figure 24 shown in (C) therein, the electronic device can present a user interface 43 at this time, which includes a first area 431 and a second area 432. The first area 431 is a viewfinder frame, which is used to display the preview image obtained by the real-time view of the camera on the electronic device; the second area includes an editing frame 4321.
[0341] In the image displayed in the first region 431, since there is no overlapping object between the object corresponding to the sphere and the object corresponding to the cube, the objects corresponding to these two objects can be regarded as two separate objects in the image, namely the target object 431A and the target object 431B. Similarly, for the image displayed in the first region 431 and the image displayed in the first region 411, the electronic device will further determine the similarity between the two images. Since the image contents are significantly different, the feature points contained in the image displayed in the first region 431 and the image displayed in the first region 411 do not match, that is, the similarity between the two images is less than the above-mentioned second threshold. Therefore, the electronic device can also trigger the process of re-determining the matte object, and detect the image displayed in the first region 431 based on the foregoing method. Since the center point of the first region 431 falls on the target object 431B at this time, the electronic device will determine the target object 431B as the matte object. After that, the electronic device can re-matte the target object 431B and pre-display the new matted image 432B in the editing box 4321. From Figure 24 As can be seen from (C) in
[0342] It should be understood that in the above embodiments, before the user manually or the electronic device automatically triggers the adjustment of the matte object displayed in the editing box, the matte object corresponding to the image displayed in the editing box will not change.
[0343] Next, the electronic device provided in this application will be introduced.
[0344] Figure 25 The structural schematic diagram of the electronic device 100 is shown.
[0345] The electronic device 100 can be a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, as well as a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) device, a virtual reality (VR) device, an artificial intelligence (AI) device, a wearable device, a vehicle-mounted device, a smart home device, and / or a smart city device. The specific type of the electronic device in this application embodiment is not particularly limited. Specifically, the electronic device 100 can be the electronic device mentioned in the foregoing method embodiments.
[0346] The electronic device 100 may include a processor 110, an internal memory 121, a sensor module 180, a button 190, a camera 193, a display screen 194, etc. Among them, the sensor module 180 may include a gyro sensor 180B, an acceleration sensor 180E, a touch sensor 180K, etc.
[0347] It can be understood that the structure schematically shown in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In some other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0348] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, and a video codec. Among them, different processing units may be independent devices or integrated in one or more processors.
[0349] The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.
[0350] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory may save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0351] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, a mobile industry processor interface (MIPI), etc.
[0352] In some embodiments, the processor 110 may include multiple groups of I2C buses. The processor 110 may be respectively coupled to the touch sensor 180K, the camera 193, etc. through different I2C bus interfaces. For example, the processor 110 may be coupled to the touch sensor 180K through an I2C interface, enabling the processor 110 and the touch sensor 180K to communicate through the I2C bus interface to implement the touch function of the electronic device 100.
[0353] In some embodiments, the MIPI interface may be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI), a display serial interface (DSI), etc. In some embodiments, the processor 110 and the camera 193 communicate through the CSI interface to implement the shooting function of the electronic device 100. The processor 110 and the display screen 194 communicate through the DSI interface to implement the display function of the electronic device 100.
[0354] It can be understood that the interface connection relationship between the modules illustrated in the embodiments of the present invention is only illustrative and does not constitute a structural limitation on the electronic device 100. In some other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0355] The electronic device 100 implements the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, which is connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.
[0356] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. In some embodiments, the electronic device 100 may include 1 or N display screens 194, where N is a positive integer greater than 1.
[0357] The electronic device 100 implements the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc.
[0358] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and light passes through the lens and is transmitted to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also perform algorithm optimization on image noise, brightness, etc. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be provided in the camera 193.
[0359] The camera 193 is used to capture still images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0360] The video codec is used to compress or decompress digital videos. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple coding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0361] The internal memory 121 may include one or more random access memories (RAMs) and one or more non-volatile memories (NVMs). The random access memory can be directly read and written by the processor 110 and can be used to store the operating system or executable programs of other running programs (such as machine instructions), and can also be used to store user and application data, etc. The non-volatile memory can also store executable programs and store user and application data, etc., and can be pre-loaded into the random access memory for direct reading and writing by the processor 110.
[0362] The gyroscope sensor 180B can be used to determine the motion posture of the electronic device 100. In some embodiments, the angular velocity of the electronic device 100 around three axes (i.e., the x, y, and z axes) can be determined by the gyroscope sensor 180B. The gyroscope sensor 180B can be used for anti-shake during shooting. Exemplarily, when the shutter is pressed, the gyroscope sensor 180B detects the shaking angle of the electronic device 100, calculates the distance that the lens module needs to compensate based on the angle, and makes the lens counteract the shaking of the electronic device 100 through reverse movement to achieve anti-shake. The gyroscope sensor 180B can also be used in navigation and somatosensory game scenarios.
[0363] The acceleration sensor 180E can detect the magnitude of the acceleration of the electronic device 100 in various directions (generally three axes). When the electronic device 100 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the posture of the electronic device and is applied to applications such as horizontal and vertical screen switching and pedometers.
[0364] The touch sensor 180K, also known as a "touch control device". The touch sensor 180K can be disposed on the display screen 194. The touch sensor 180K and the display screen 194 form a touch screen, also known as a "touch control screen". The touch sensor 180K is used to detect touch operations acting thereon or nearby. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In some other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device 100, at a different position from that of the display screen 194.
[0365] The keys 190 include a power-on key, volume keys, etc. The keys 190 can be mechanical keys. They can also be touch keys. The electronic device 100 can receive key inputs and generate key signal inputs related to the user settings and function control of the electronic device 100.
[0366] In the embodiments of the present application, the processor 110 of the electronic device 100 can execute the methods in Figure 4 , Figure 8 , Figure 15 , Figure 17 and Figure 22 .
[0367] Specifically, the processor 110 can simultaneously display a first area and a second area on the display screen 194. The first area can display the picture obtained by the real-time viewfinder of the camera 193. The processor 110 can perform matting on the matting object existing in the picture and pre-display the obtained matted image in the second area.
[0368] When pre-displaying the matted image in the second area, the processor 110 can send a corresponding drawing instruction to the GPU. The drawing instruction can set the position, size area, and transparency of the matted image in the second area. For example, the drawing instruction sent by the processor 110 can set the transparency of the matted image to 50% to indicate that the matted image is pre-displayed in the second area.
[0369] The processor 110 may also, in response to a user operation, officially display the above-mentioned matte image in the above-mentioned second area. Similarly, the processor 110 may, in response to this user operation, issue a corresponding drawing instruction to the GPU again. The drawing instruction may set a different transparency for the matte image. For example, the transparency of the matte image is set to 0%, that is, completely opaque, indicating that the matte image has been officially input into the second area at this time.
[0370] The software system of the electronic device 100 may adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In the embodiments of the present invention, the Android system with a layered architecture is taken as an example to exemplarily illustrate the software structure of the electronic device 100.
[0371] Figure 26 It is a software structure block diagram of the electronic device 100 according to the embodiments of the present invention.
[0372] The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into five layers, from top to bottom, namely the application layer, the system framework layer, the algorithm engine layer, the Android runtime and the system libraries, and the kernel layer.
[0373] The application layer may include a series of application packages.
[0374] As Figure 26 shown, the application packages may include multiple application programs that support text and image output, such as SMS / MMS, notes, document editing, etc. The visual input is responsible for determining whether the edit box can input the cropped image and for inputting the extracted image into the current edit box; at the same time, the visual input is responsible for starting the camera preview stream page and loading it onto the input method window by the input method framework.
[0375] The system framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions. At the same time, the system framework layer can provide system support for visual input and is responsible for managing the entire lifecycle of visual input. The Camera2 / x framework provides system support for the visual input preview stream and manages the camera lifecycle and data acquisition. The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, and so on. The notification manager enables applications to display notification messages in the status bar. It can be used to convey informative messages, which can automatically disappear after a short stay without user interaction. For example, the notification manager is used to inform that the download is complete, message reminders, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as the notification of a background running application, or a notification that appears on the screen in the form of a dialogue window. For example, it can prompt text information in the status bar, emit a prompt tone, vibrate the electronic device, blink the indicator light, etc.
[0376] The algorithm engine layer may include multiple algorithms for assisting in implementing the image processing method provided in this application. For example, the subject detection algorithm can be used to detect the significant subject (target object) in the picture, the matte object tracking algorithm can track the subject to be matted (i.e., the matte object) in the picture, and the matte strategy selection algorithm can identify the category of the matte object and select different matte algorithms according to different categories of the matte object. When the category is a portrait, the portrait matte algorithm is selected; for other categories, the matte algorithm for other types of subjects is selected. The subject detection algorithm can be used to detect the image quality of the matted image and not display the matted image when the image quality is poor.
[0377] The Android runtime includes a core library and a virtual machine and is responsible for the scheduling and management of the Android system.
[0378] The system library can include multiple functional modules. For example: surface manager, Media Libraries, 3D graphics processing library (e.g., OpenGL ES), 2D graphics engine (e.g., SGL), etc. The surface manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications. The Media Libraries support the playback and recording of multiple common audio and video formats, as well as static image files, etc. The Media Libraries can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc. The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc. The 2D graphics engine is the drawing engine for 2D drawing.
[0379] The kernel layer is the layer between hardware and software. The kernel layer includes at least a display driver, a camera driver, an audio driver, and a sensor driver.
[0380] The following takes the scenario of shooting and matting preview as an example to exemplarily illustrate the working processes of the software and hardware of the electronic device 100.
[0381] When the touch sensor 180K receives a touch operation, the corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including information such as touch coordinates and the timestamp of the touch operation). The raw input event is stored in the kernel layer. The system framework layer obtains the raw input event from the kernel layer and identifies the control corresponding to the input event. Taking the touch operation as a touch click operation and the control corresponding to the click operation being the shooting input option in the note as an example, the note application calls the interface of the system framework layer to start the camera application, and then starts the camera driver by calling the kernel layer, and captures a static image or video through the camera 193. At the same time, the poem input calls the interface of the system framework layer, and then calls the algorithm in the algorithm engine layer to perform matting on the image displayed by the camera, and sends the matted image back to the application layer for display.
[0382] The embodiment of the present application also provides an electronic device, which includes: one or more processors and a memory; wherein, the memory is coupled to the one or more processors, and the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the electronic device to execute the method shown in the foregoing embodiment.
[0383] As used in the foregoing embodiments, depending on the context, the term "when..." may be interpreted to mean "if...", or "after...", or "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if detecting (the stated condition or event)" may be interpreted to mean "if determining...", or "in response to determining...", or "when detecting (the stated condition or event)", or "in response to detecting (the stated condition or event)".
[0384] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive), etc.
[0385] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware with a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The foregoing storage medium includes: various media that can store program codes such as ROM or random access memory RAM, magnetic disk, or optical disc.
Claims
1. An image processing method, applied to an electronic device, characterized in that Including: Display a first interface, where the first interface includes a first edit box and a first control; In response to a first operation of the user on the first control, display a second interface, where the second interface includes a first area and a second area, the first area includes a first picture captured in real time by the camera of the electronic device, the first picture includes a first object, the second area includes the first edit box, and the first edit box in the second area includes a second object, and the second object is obtained by matting the first object.
2. The method according to claim 1, wherein After displaying the second interface, the method further includes: Display a second picture captured in real time by the camera in the first area, where the second picture includes a first changing object, the first changing object is different from the first object and both correspond to a first entity; Display a third object in the first edit box, and the third object is obtained by changing the second object by a first proportional value.
3. The method according to claim 2, wherein The ratio of the area of the first changing object to the area of the first object is a second proportional value, and the first proportional value is positively correlated with the second proportional value.
4. The method according to claim 2, characterized in that, After displaying the second interface, the method further includes: Display a third picture captured in real time by the camera in the first area, where the third picture includes a second changing object, the second changing object is different from the first changing object and the first object, and both the second changing object and the first changing object correspond to the first entity, and the first edit box in the second area includes a fourth object, and the fourth object is obtained by matting the second changing object.
5. The method according to claim 1, characterized in that The first object corresponds to a first entity. After displaying the second interface, the method further includes: Display a fourth picture captured in real time by the camera in the first area, where the fourth picture includes a fifth object, and the fifth object corresponds to a second entity, and the second entity is different from the first entity; Replace the second object in the first edit box in the second area with a sixth object, and the sixth object is obtained by matting the fifth object.
6. The method according to any one of claims 1-5, characterized in that, The first picture further includes a seventh object, and the seventh object corresponds to a third entity, and the third entity is different from the first entity; The distance between the center point of the first object and the center point of the first area is less than the distance between the center point of the seventh object and the center point of the first area.
7. The method according to any one of claims 1-5, characterized in that, The first picture further includes an eighth object, and the eighth object corresponds to a fourth entity, and the fourth entity is different from the first entity, and the area of the first object is greater than the area of the eighth object.
8. The method according to claim 1, characterized in that, The first picture further includes a ninth object. After displaying the second interface, the method further includes: In response to a second operation of the user on the ninth object, replace the second object displayed in the first edit box with a tenth object, and the tenth object is obtained by matting the ninth object.
9. The method according to claim 1, wherein The first picture further includes an eleventh object. After displaying the second interface, the method further includes: In response to a third operation of the user, the twelfth object and the second object are displayed together in the first editing box, and the twelfth object is obtained by matting the eleventh object.
10. The method according to claim 1, characterized in that, The second object in the first editing box is a preview image. After displaying the second interface, the method further includes: In response to a fourth operation on the first object in the first screen, the second object displayed in the first editing box is replaced with a thirteenth object, which is obtained by matting the first object, and the display mode of the thirteenth object in the first editing box is different from that of the second object in the first editing box.
11. The method according to claim 2 or 3, characterized in that, Before displaying the second screen captured in real time by the camera in the first area, the method further includes: Receiving a fifth operation input by the user, where the fifth operation is an operation to adjust the focal length of the camera or the distance between the camera and the first entity, so that the camera captures the second screen.
12. The method according to claim 4, characterized in that Before displaying the third screen captured in real time by the camera in the first area, the method further includes: Receiving a sixth operation input by the user: the sixth operation is an operation to change the relative position between the camera and the first entity, so that the camera captures the third screen.
13. An image processing method, characterized in that, including: Receiving a first instruction, which is an instruction generated in response to a first operation on a first control in a first interface. The first instruction is used to start the camera of the electronic device, and the first interface includes a first editing box; Displaying a second interface, which includes a first area and a second area. The first area includes a first screen captured in real time by the camera, and the first screen includes a first object. The second area includes the first editing box; Using a matting algorithm to extract the first object from the first screen to obtain a second object; Displaying the second object in the first editing box in the first area.
14. The method according to claim 13, characterized in that, After displaying the second interface, the method further includes: Displaying a second screen captured in real time by the camera in the first area, where the second screen includes a first changing object, and the first changing object is different from the first object and both correspond to the first entity; Comparing the similarity between the first object in the first screen and the first changing object in the second screen; When the similarity between the first object in the first screen and the first changing object in the second screen is greater than or equal to a first threshold, determining a first ratio value according to the ratio of the area of the first object in the first screen to the area of the first changing object in the second screen; Scaling the second object by the first ratio value to obtain a third object; Replacing the second object displayed in the first editing box with the third object.
15. The method according to claim 13, characterized in that, After displaying the second interface, the method further includes: Displaying a third screen captured in real time by the camera in the first area, where the third screen includes a second changing object; Comparing the similarity between the first object in the first screen and the second changing object in the third screen; When the similarity between the first object in the first screen and the second changing object in the third screen is less than a first threshold, the second changing object in the third screen is cut out using a matting algorithm to obtain a fourth object, and the second object displayed in the first editing box is used to replace the fourth object.
16. The method according to claim 13, characterized in that, After displaying the second interface, the method further includes: Displaying a fourth screen captured in real time by a camera in the first area, the fourth screen including the first object and a fifth object; Comparing the similarity between the first screen and the fourth screen; When the similarity between the first screen and the fourth screen is less than a second threshold, the fifth object is cut out using a matting algorithm to obtain a sixth object, and the second object displayed in the first editing box is used to replace the sixth object.
17. The method according to any one of claims 13 - 16, characterized in that, Before receiving and displaying the second interface, the method further includes: Detecting the first screen using an object detection algorithm; When at least one target object is detected in the first screen, determining the first object from the first screen, the first object corresponding to a first category.
18. The method according to any one of claims 13-17, characterized in that The first screen includes multiple objects. Before displaying the second interface, the method further includes: When the center point of the first area lies on one of the multiple objects, determining the object on which the center point of the first area lies as the first object; When the center point of the first area does not lie on any of the multiple objects, calculating the Manhattan distance between each object in the multiple objects and the center point, sorting the Manhattan distances between each object in the multiple objects and the center point; when the difference between every two adjacent Manhattan distances is less than a preset threshold, determining the object with the largest area among the multiple objects as the first object; when the difference between every two adjacent Manhattan distances is not less than the preset threshold, determining the object with the smallest Manhattan distance from the center point among the multiple objects as the first object.
19. The method according to claim 17 or 18, characterized in that, The cutting out of the first object from the first screen using the matting algorithm includes: When the first category is a portrait, using a portrait matting algorithm to matte the first object; When the first category is not a portrait, using a non - portrait matting algorithm to matte the first object.
20. The method according to claim 14, wherein Before comparing the similarity between the first object in the first screen and the second changing object in the second screen, the method further includes: Extracting multiple frames of images from the images captured by the camera; Sorting the multiple frames of images according to the acquisition time of the camera; Determining two consecutive adjacent frames or the first frame and the last frame among the multiple frames of images as the first screen and the second screen.
21. The method according to claim 16, wherein Before comparing the similarity between the first screen and the fourth screen, the method further includes: Extracting multiple frames of images from the images captured by the camera; Sorting the multiple frames of images according to the acquisition time of the camera; Determine two consecutive adjacent frames or the first frame and the last frame among the multiple frames as the first frame and the fourth frame.
22. An electronic device, characterized in that, The electronic device includes: one or more processors, a memory, and a display screen; The memory is coupled to the one or more processors, and the memory is used to store computer program code. The computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the electronic device to execute the method according to any one of claims 1-21.
23. A chip system, characterized in that, The chip system is applied to an electronic device. The chip system includes one or more processors, and the processors are used to call computer instructions to cause the electronic device to execute the method according to any one of claims 1-21.
24. A computer-readable storage medium, comprising instructions, characterized in that, When the instruction runs on the electronic device, it causes the electronic device to execute the method according to any one of claims 1-21.
Citation Information
Patent Citations
Methods for previewing objects during drag-and-drop, and client-side implementation.
CN102279692A
Processing, displaying method, device and mobile terminal for three-dimensional image data
CN109068063A
Image display method, device and electronic equipment
CN112584040A
Display device for executing a plurality of applications and method for controlling the same
US20140164957A1