A method for removing image content and related device
By replacing the selfie stick through terminal recognition and generation of repair images, the problem of the selfie stick being unable to be completely removed is solved, and the user experience and image display effect are improved.
Patent Information
- Application Number
- CN202211521236.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-30
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2040-05-30
AI Technical Summary
The existing technology cannot completely remove the selfie stick when the selfie stick is not within the blind spot of shooting, and requires special camera hardware configuration and selfie stick placement, which has limited applicability.
The preview image and reference frame image captured by the camera are obtained through the terminal, the object to be removed is identified and a repair image is generated to replace the obscured image content and realize the removal of the selfie stick.
On terminals without special cameras, the selfie stick is effectively removed, improving user experience and the display effect of image content.
Smart Images

Figure CN115914826B_ABST
Abstract
Description
[0001] This application is a divisional application. The application number of the original application is 202010481007.5, and the original application date is May 30, 2020. The entire content of the original application is incorporated into this application by reference. Technical Field
[0002] The present application relates to the field of computer vision, and in particular to a method for removing image content and related devices. Background Art
[0003] With the development of smartphones, taking photos and videos has become one of their most important features. As smartphone camera capabilities become increasingly powerful, more and more people are using smartphones instead of cameras to take photos. To achieve a wider range of shooting angles, smartphones are often mounted on retractable selfie sticks. By freely adjusting the extension and retraction of the stick, users can capture selfies from various angles. However, when using a selfie stick, there is a risk of the stick being partially captured, meaning that the selfie stick may appear in the captured photo or video, affecting the user experience.
[0004] In existing solutions, to remove the selfie stick from captured photos or videos, the camera terminal is equipped with a dual fisheye lens. Specifically, the camera terminal is equipped with two cameras with 180° shooting angles, which together form a shooting range of approximately 200°. When the selfie stick is in the blind spot of the two cameras, the camera terminal can hide the selfie stick by cropping and splicing the images captured by the two 180° cameras. However, if there is a certain angle between the selfie stick and the two cameras, a portion of the selfie stick will still be visible in the cropped and spliced image, which cannot be completely hidden. In existing solutions, the camera terminal must have special camera hardware configuration and the selfie stick must be placed in a special position to completely remove the selfie stick. These stringent removal conditions are not applicable to most camera terminals. Summary of the Invention
[0005] The present application provides a method and related device for removing image content, which can remove unwanted image content from pictures or videos taken by the user on a terminal without a special camera, improve the display effect of the image content that the user wants in the picture or video, and enhance the user experience.
[0006] In a first aspect, the present application provides a method for removing image content, comprising: a terminal launching a camera application. The terminal displays a photo preview interface of the camera application. The terminal obtains a first preview screen and a first reference frame screen captured by the camera, wherein the first preview screen and the first reference frame screen both include image content of a first object and image content of a second object, and in the first preview screen, the image content of the first object obscures part of the image of the second object. The terminal determines that the first object in the first preview screen is the object to be removed. The terminal determines content to be filled in the first preview screen based on the first reference frame screen, wherein the content to be filled is the image content of the second object obscured by the first object in the first preview screen. The terminal generates a first repair screen based on the content to be filled and the first preview screen, wherein in the first repair screen, the image content of the first object is replaced with the obscured image content of the second object. The terminal displays the first repair screen in the photo preview interface.
[0007] Through the image content removal method provided in this application, the terminal can obtain a preview picture and a reference frame picture through the camera when taking a photo, and remove the image content that the user does not want in the preview picture (such as a selfie stick) through the reference frame picture, thereby improving the display effect of the image content that the user wants in the picture or video and improving the user experience.
[0008] In one possible implementation, after the terminal displays the first restoration image on the photo preview interface, the method further includes: displaying, by the terminal, a removal control on the photo preview interface. The terminal receives a first user input for the removal control. In response to the first input, the terminal obtains a second preview image captured by the camera. The terminal displays the second preview image on the photo preview interface. In this way, the terminal can disable the removal function for a specified object in the preview image as desired by the user.
[0009] In one possible implementation, before the terminal obtains the first preview image and the reference frame image captured by the camera, the method further includes: the terminal displays a third preview image on the photo preview interface. After the terminal recognizes that the third preview image includes the object to be removed, it displays a removal confirmation control. The terminal receives a second input from the user regarding the removal confirmation control. The terminal obtains the first preview image and the first reference frame image captured by the camera, specifically including: in response to the second input, the terminal obtains the first preview image and the first reference frame image captured by the camera. In this way, the terminal can remove the first object in the preview image after the user confirms.
[0010] In one possible implementation, the method further includes: in response to the third input, the terminal displays a countdown of a specified duration on the photo preview interface. In this way, the countdown can be displayed before removing the first object from the preview screen, allowing the user to perceive the processing time.
[0011] In one possible implementation, before the terminal displays the first repair screen on the photo preview interface, the method further includes: the terminal displays a third preview screen on the photo preview interface. The terminal receives a click operation from the user on the third preview screen. The terminal determines that the first object in the first preview screen is the object to be removed, specifically including: in response to the click operation, the terminal identifies the click location in the third preview screen. The terminal determines that the first object is the object to be removed based on the image content at the click location in the third preview screen. In this way, the terminal can determine the object that the user wants to remove based on the user's click operation.
[0012] In a possible implementation, before the terminal displays the first repair screen on the photo preview interface, the method further includes: the terminal displays a third preview screen on the camera application interface. The terminal identifies the image content of one or more removable objects in the third preview screen, and displays a removal control corresponding to the removable object. The terminal receives a fourth input from the user for a first removal control among the one or more removal controls. The terminal determines that the first object in the first preview screen is the object to be removed, specifically including: in response to the fourth input, the terminal determines the first object corresponding to the first removal control as the object to be removed. In this way, the terminal can identify all removable objects in the preview screen and prompt the user to select the object to be removed.
[0013] In one possible implementation, before the terminal obtains the first preview image and the first reference frame image captured by the camera, the method further includes: the terminal displaying a first shooting mode control on the camera preview interface. The terminal receives a fifth input from the user regarding the first shooting mode control. The terminal obtains the first preview image and the first reference frame image captured by the camera, specifically including: in response to the fifth input, the terminal obtains the first preview image and the first reference frame image captured by the camera. In this way, the terminal can activate the object removal function in a specific shooting mode.
[0014] In one possible implementation, before the terminal obtains the first preview image and the first reference frame image captured by the camera, the method further includes: when the terminal determines that the terminal's captured image has significantly moved, the terminal displays a screen shake prompt, the screen shake prompt being used to notify the user of the significant movement of the terminal's captured image. In this way, the terminal can encourage user cooperation to ensure the quality of object removal.
[0015] In one possible implementation, the terminal determines that a captured image of the terminal has significantly moved by: obtaining, by the terminal, angular velocity data and acceleration data of the terminal through an inertial measurement unit. When the angular velocity in any direction of the angular velocity data exceeds a specified angular velocity value, or the acceleration in any direction of the acceleration data exceeds a specified acceleration value, the terminal determines that the captured image of the terminal has significantly moved. In this way, the terminal can detect the magnitude of image motion using motion data.
[0016] In a possible implementation, before the terminal obtains the first preview picture and the first reference frame picture captured by the camera, the method further includes: the terminal displays a third preview picture on the camera application interface. When the terminal recognizes that the third preview picture includes the specified image content, the terminal displays a movement operation prompt, and the movement operation prompt is used to prompt the user to move the terminal in a specified direction. The terminal determines the content to be filled in the first preview picture based on the first reference frame picture, specifically including: when the terminal determines that the picture motion amplitude between the first preview picture and the first reference frame picture exceeds a specified threshold, the terminal determines the content to be filled in the first preview picture based on the first reference frame picture. In this way, the terminal can prompt the user to move the terminal in the specified direction to ensure the removal effect of the object in the preview picture.
[0017] In one possible implementation, the terminal determines that the amplitude of the motion between the first preview image and the first reference frame image exceeds a specified threshold, specifically including: generating, by the terminal, a first mask image after segmenting the first object in the first preview image. Generating, by the terminal, a second mask image after segmenting the first object in the first reference frame image. Calculating, by the terminal, an intersection-and-union (IoU) ratio between the first mask image and the second mask image. When the IoU ratio between the first mask image and the second mask image is less than a specified IoU ratio value, the terminal determines that the amplitude of the motion between the first preview image and the first reference frame image exceeds the specified threshold.
[0018] In a possible implementation, the terminal determines that the amplitude of the motion between the first preview picture and the first reference frame picture exceeds a specified threshold, specifically including: the terminal identifies the first object in the first preview picture and segments the first object in the first preview picture. The terminal identifies the first object in the first reference frame picture and segments the first object in the first reference frame picture to obtain the second reference frame picture. The terminal encodes the first preview picture after segmenting the first object into a first target feature map. The terminal encodes the second reference frame picture into a first reference feature map. The terminal calculates the similarity between the first target feature map and the first reference feature map. When the similarity between the first target feature map and the first reference feature map is less than a specified similarity value, the terminal determines that the amplitude of the motion between the first preview picture and the first reference frame picture exceeds the specified threshold.
[0019] In a possible implementation, the method further includes: the terminal receiving a fifth input from the user. In response to the fifth input, the terminal locally saves the first repair screen.
[0020] In a possible implementation, the terminal determines the content to be filled in the first preview picture based on the first reference frame, specifically including: the terminal identifies the first object in the first preview picture and segments the first object in the first preview picture. The terminal identifies the first object in the first reference frame and segments the first object in the first reference frame to obtain the second reference frame. The terminal calculates the missing optical flow information between the first preview picture and the second reference frame after segmenting the first object. The terminal completes the missing optical flow information based on the second reference frame using an optical flow completion model to obtain complete optical flow information between the first preview picture and the second reference frame after segmenting the first object. The terminal determines the content to be filled in the first preview picture from the second reference frame using the complete optical flow information. In this way, the terminal can repair the preview picture through the optical flow field.
[0021] In one possible implementation, the terminal determines content to be filled in the first preview image based on the first reference frame, specifically including: the terminal identifying the first object in the first preview image and segmenting the first object in the first preview image. The terminal also identifies the first object in the first reference frame and segmenting the first object in the first reference frame to obtain the second reference frame. The terminal encodes the first preview image after segmenting the first object into a first target feature map. The terminal encodes the second reference frame into a first reference feature map. The terminal determines, from the first reference feature map, features to be filled that are similar to features surrounding the first region in the first target feature map. The terminal generates a first repaired image based on the content to be filled and the first preview image, specifically including: the terminal replacing the features to be filled with the region in the first target feature map where the first object is located to obtain a second target feature map. The terminal decodes the second target feature map to obtain the first repaired image. In this way, the terminal can repair the preview image at the feature level using the reference frame.
[0022] In one possible implementation, the terminal generates a first repaired image based on the content to be filled and the first preview image, specifically including: filling the area of the first preview image where the first object is located with the content to be filled, thereby obtaining a rough repaired image; and generating a detailed texture of the filled area in the rough repaired image, thereby obtaining the first repaired image. In this manner, the terminal can further generate a detailed texture of the filled area.
[0023] In one possible implementation, after the terminal determines the content to be filled in the first preview image based on the first reference frame, the method further includes: the terminal obtaining a fourth preview image captured by the camera. The terminal obtaining the terminal's motion angle and rotation angle between the time the camera captured the first preview image and the fourth preview image. The terminal determines the area within the fourth preview image where the first object is located based on the terminal's motion angle, rotation angle, and the area within the first preview image. The terminal segments the first object in the fourth preview image. The terminal determines content to be filled in the fourth preview image based on the area within the fourth preview image where the first object is located. The terminal fills the area within the fourth preview image where the first object is located with the content from the fourth preview image, thereby obtaining a second restored image. The terminal displays the second restored image on the shooting preview interface. In this way, when removing an object from consecutive frames, the terminal infers the position of the selfie stick in subsequent frames based on motion data, thereby determining the content to be filled in the selfie stick area in the subsequent frames, thus saving removal time.
[0024] In a possible implementation, the first object includes a selfie stick or a background person.
[0025] In a second aspect, the present application provides a terminal comprising a camera, one or more processors, and one or more memories. The one or more memories and the camera are coupled to the one or more processors, and the one or more memories are used to store computer program code, the computer program code comprising computer instructions. When the one or more processors execute the computer instructions, the terminal performs the method for removing image content in any possible implementation of any of the above aspects.
[0026] In a third aspect, the present application provides a terminal, comprising: one or more functional modules, wherein the one or more functional modules are used to execute the method for removing image content in any possible implementation of any of the above aspects.
[0027] In a fourth aspect, an embodiment of the present application provides a computer storage medium comprising computer instructions. When the computer instructions are executed on a terminal, the terminal executes the method for removing image content in any possible implementation of any of the above aspects.
[0028] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when running on a computer, enables the computer to execute the method for removing image content in any possible implementation of any of the above aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1A-1B A schematic diagram of the principle of removing a selfie stick in the prior art;
[0030] Figure 2A A hardware structure diagram of a terminal provided in an embodiment of the present application;
[0031] Figure 2B A schematic diagram of the software architecture of a terminal provided in an embodiment of the present application;
[0032] Figure 3A-3F A set of interface schematic diagrams provided for embodiments of the present application;
[0033] Figure 4A-4G Another set of interface schematic diagrams provided for an embodiment of the present application;
[0034] Figures 5A-5C Another set of interface schematic diagrams provided for an embodiment of the present application;
[0035] Figures 6A-6C Another set of interface schematic diagrams provided for an embodiment of the present application;
[0036] Figures 7A-7F Another set of interface schematic diagrams provided for an embodiment of the present application;
[0037] Figures 8A-8C Another set of interface schematic diagrams provided for an embodiment of the present application;
[0038] Figures 9A-9F Another set of interface schematic diagrams provided for an embodiment of the present application;
[0039] Figures 10A-10G Another set of interface schematic diagrams provided for an embodiment of the present application;
[0040] Figures 11A-11C Another set of interface schematic diagrams provided for an embodiment of the present application;
[0041] Figures 12A-12D Another set of interface schematic diagrams provided for an embodiment of the present application;
[0042] Figure 13 A schematic diagram of the architecture of an image content removal system provided in an embodiment of the present application;
[0043] Figure 14A A schematic diagram of a first target image provided in an embodiment of the present application;
[0044] Figure 14B A schematic diagram of a second target image provided in an embodiment of the present application;
[0045] Figure 14C A schematic diagram of a first reference image provided in an embodiment of the present application;
[0046] Figure 14D A schematic diagram of a second reference image provided in an embodiment of the present application;
[0047] Figure 14E A mask image of a second target image provided in an embodiment of the present application;
[0048] Figure 14F A schematic diagram of a third target image provided in an embodiment of the present application;
[0049] Figure 14G A schematic diagram of a fourth target image provided in an embodiment of the present application;
[0050] Figure 15 A schematic diagram of a rough optical flow restoration process provided in an embodiment of the present application;
[0051] Figure 16 A schematic diagram of a multi-frame feature rough restoration process provided in an embodiment of the present application;
[0052] Figure 17A A schematic diagram of a first target characteristic graph provided in an embodiment of the present application;
[0053] Figure 17B A schematic diagram of a first reference characteristic graph provided in an embodiment of the present application;
[0054] Figure 18 A schematic diagram of a single-frame feature rough restoration process provided in an embodiment of the present application;
[0055] Figure 19 A schematic diagram of a method for selecting a rough repair process provided in an embodiment of the present application;
[0056] Figure 20 A flowchart of a method for removing image content provided in an embodiment of the present application. DETAILED DESCRIPTION
[0057] The following is a clear and detailed description of the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in the text is only a description of the association relationship between related objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.
[0058] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of the technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of this application, unless otherwise specified, "plurality" means two or more.
[0059] The following describes a method for removing the selfie stick from a photograph in an existing solution.
[0060] Figure 1A and Figure 1B A schematic diagram showing the principle of a method for removing a selfie stick for taking pictures in an existing solution is shown.
[0061] like Figure 1A As shown, in the existing solution, the camera terminal is equipped with two 180° cameras. After the camera terminal captures two images using the two 180° cameras, it can crop the common display area of the two images and then stitch them together into a single image. When a user attaches the camera terminal to a selfie stick to take a photo, the selfie stick needs to be placed within the camera terminal's blind spot. Only when the camera terminal crops and stitches the two images captured by the two 180° cameras can the selfie stick be completely removed from the image.
[0062] like Figure 1BAs shown, when the selfie stick is not completely within the blind spot of the shooting, when the shooting terminal crops and splices the two pictures taken by the two 180° cameras, the part of the selfie stick that is not within the blind spot of the shooting cannot be removed and will also appear in the spliced picture.
[0063] It can be seen from the above existing solutions that the selfie stick can only be completely removed when the shooting terminal has a special camera hardware configuration and the selfie stick has a special placement. The conditions for removing the selfie stick are harsh and cannot be applied to most shooting terminals.
[0064] Therefore, an embodiment of the present application provides a method for removing image content, which can remove image content that the user does not want in pictures or videos taken by the user (such as a selfie stick) on a terminal without a special camera, improve the display effect of the image content that the user wants in the picture or video, and improve the user experience.
[0065] Figure 2A A schematic structural diagram of the terminal 100 is shown.
[0066] The embodiment will be described in detail below using terminal 100 as an example. It should be understood that Figure 2A The terminal 100 shown is only an example, and the terminal 100 may have more Figure 2A The more or less components shown in the figure can be combined with two or more components, or can have different component configurations. The various components shown in the figure can be implemented in hardware, software, or a combination of hardware and software including one or more signal processing and / or application specific integrated circuits.
[0067] The terminal 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0068] It should be understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on terminal 100. In other embodiments of the present application, terminal 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0069] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.
[0070] The controller may be the nerve center and command center of the terminal 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.
[0071] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.
[0072] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface 130, among others.
[0073] The charging management module 140 is configured to receive charging input from a charger, which may be a wireless charger or a wired charger.
[0074] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to provide power to the processor 110, the internal memory 121, the external memory, the display 194, the camera 193, and the wireless communication module 160.
[0075] The wireless communication function of the terminal 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.
[0076] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in terminal 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.
[0077] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied on the terminal 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.
[0078] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the receiver 170B, etc.) or displays an image or video through the display screen 194. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 110 and be set in the same device as the mobile communication module 150 or other functional modules.
[0079] The wireless communication module 160 can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. applied on the terminal 100. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.
[0080] In some embodiments, the antenna 1 of the terminal 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the terminal 100 can communicate with the network and other devices through wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS may include a global positioning system (GPS), a global navigation satellite system (GLONASS), a Beidou navigation satellite system (BDS), a quasi-zenith satellite system (QZSS) and / or a satellite based augmentation system (SBAS).
[0081] Terminal 100 implements display functions through a GPU, display screen 194, and an application processor. The GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.
[0082] Display screen 194 is used to display images, videos, and the like. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, or a quantum dot light-emitting diode (QLED). In some embodiments, terminal 100 may include one or N display screens 194, where N is a positive integer greater than one.
[0083] The terminal 100 can realize the shooting function through the ISP, camera 193, video codec, GPU, display screen 194 and application processor.
[0084] The ISP processes data fed back by camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and converted into a visible image. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. It can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within camera 193.
[0085] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the terminal 100 may include 1 or N cameras 193, where N is a positive integer greater than 1.
[0086] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the terminal 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.
[0087] Video codecs are used to compress or decompress digital video. Terminal 100 may support one or more video codecs. This allows terminal 100 to play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.
[0088] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU enables intelligent cognitive applications in the terminal 100, such as image recognition, face recognition, speech recognition, and text comprehension.
[0089] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the terminal 100. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.
[0090] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the terminal 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area can store data created during the use of the terminal 100 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0091] The terminal 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.
[0092] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals.
[0093] The speaker 170A, also called a "horn", is used to convert audio electrical signals into sound signals.
[0094] The receiver 170B, also called a "handset", is used to convert audio electrical signals into sound signals.
[0095] Microphone 170C, also called "microphone" or "microphone", is used to convert sound signals into electrical signals.
[0096] The headphone jack 170D is used to connect a wired headphone and can be the USB interface 130 or a 3.5mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0097] The pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, the pressure sensor 180A can be set on the display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. A capacitive pressure sensor can be a device comprising at least two parallel plates with conductive material. When a force acts on the pressure sensor 180A, the capacitance between the electrodes changes. The terminal 100 determines the intensity of the pressure based on the change in capacitance. When a touch operation is applied to the display screen 194, the terminal 100 detects the intensity of the touch operation based on the pressure sensor 180A. The terminal 100 can also calculate the position of the touch based on the detection signal of the pressure sensor 180A. In some embodiments, touch operations acting on the same touch position but with different touch operation intensities can correspond to different operation instructions.
[0098] The gyro sensor 180B can be used to determine the motion posture of the terminal 100. In some embodiments, the angular velocity of the terminal 100 around three axes (i.e., x, y, and z axes) can be determined by the gyro sensor 180B. The gyro sensor 180B can also be used for anti-shake photography.
[0099] The air pressure sensor 180C is used to measure air pressure.
[0100] The magnetic sensor 180D includes a Hall sensor, and the terminal 100 can use the magnetic sensor 180D to detect whether the flip cover is opened or closed.
[0101] The acceleration sensor 180E can detect the magnitude of acceleration of the terminal 100 in various directions (generally three axes), and can detect the magnitude and direction of gravity when the terminal 100 is stationary.
[0102] The distance sensor 180F is used to measure distance. The terminal 100 can measure distance using infrared or laser. In some embodiments, when shooting a scene, the terminal 100 can use the distance sensor 180F to measure distance to achieve fast focusing.
[0103] The proximity light sensor 180G may include, for example, a light emitting diode (LED) and a light detector, such as a photodiode. The light emitting diode may be an infrared light emitting diode.
[0104] The ambient light sensor 180L is used to sense the ambient light brightness. The terminal 100 can adaptively adjust the brightness of the display screen 194 based on the sensed ambient light brightness. The ambient light sensor 180L can also be used to automatically adjust the white balance when taking photos.
[0105] The fingerprint sensor 180H is used to collect fingerprints. The terminal 100 can use the collected fingerprint characteristics to implement fingerprint unlocking, access application locks, fingerprint photography, fingerprint call answering, etc.
[0106] The temperature sensor 180J is used to detect temperature.
[0107] The touch sensor 180K is also called a "touch panel." The touch sensor 180K can be disposed on the display screen 194. The touch sensor 180K and the display screen 194 form a touch screen, also called a "touch screen." The touch sensor 180K is used to detect touch operations applied thereto or in the vicinity thereof. The touch sensor can transmit the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operations can be provided via the display screen 194. In other embodiments, the touch sensor 180K can also be disposed on the surface of the terminal 100, in a location different from that of the display screen 194.
[0108] The bone conduction sensor 180M can acquire vibration signals.
[0109] Keys 190 include a power button, a volume button, etc. Keys 190 may be mechanical keys or touch keys. Terminal 100 may receive key inputs and generate key signal inputs related to user settings and function control of terminal 100.
[0110] Motor 191 can generate vibration prompts.
[0111] The indicator 192 may be an indicator light, which may be used to indicate the charging status, power level changes, messages, missed calls, notifications, etc.
[0112] The SIM card interface 195 is used to connect a SIM card.
[0113] The software system of the terminal 100 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture. In the embodiment of the present invention, the Android system with a layered architecture is used as an example to illustrate the software structure of the terminal 100.
[0114] Figure 2B It is a software structure block diagram of the terminal 100 according to an embodiment of the present invention.
[0115] A layered architecture divides software into several layers, each with distinct roles and responsibilities. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.
[0116] The application layer can include a series of application packages.
[0117] like Figure 2B As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc.
[0118] The application framework layer provides an application programming interface (API) and programming framework for the applications in the application layer. The application framework layer includes some predefined functions.
[0119] like Figure 2B As shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like.
[0120] The window manager is used to manage window programs. The window manager can obtain the display size, determine whether there is a status bar, lock the screen, take screenshots, etc.
[0121] Content providers are used to store and retrieve data and make it accessible to applications. The data may include videos, images, audio, calls made and received, browsing history and bookmarks, phone books, etc.
[0122] The view system includes visual controls, such as those for displaying text and images. The view system is used to build applications. A display interface can consist of one or more views. For example, a display interface containing a text notification icon might include a view for displaying text and a view for displaying images.
[0123] The phone manager is used to provide communication functions of the terminal 100, such as management of call status (including answering, hanging up, etc.).
[0124] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.
[0125] The Notification Manager allows applications to display notifications in the status bar. These messages can be displayed briefly and then disappear automatically, without requiring user interaction. For example, the Notification Manager can be used to notify users of completed downloads and message reminders. The Notification Manager can also display notifications in the top status bar as icons or scrolling text, such as notifications from background applications, or as dialog windows on the screen. Examples include displaying text messages in the status bar, emitting alert sounds, vibrating the terminal, or flashing indicator lights.
[0126] Android Runtime includes core libraries and a virtual machine. Android runtime is responsible for scheduling and management of the Android system.
[0127] The core library consists of two parts: one is the function that needs to be called by the Java language, and the other is the Android core library.
[0128] The application layer and application framework layer run in a virtual machine. The virtual machine executes Java files in the application layer and application framework layer as binary files. The virtual machine manages object lifecycles, stack management, thread management, security and exception management, and garbage collection.
[0129] The system library can include multiple functional modules, such as surface manager, media library, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.
[0130] The surface manager is used to manage the display subsystem and provide fusion of 2D and 3D layers for multiple applications.
[0131] The media library supports playback and recording of a variety of common audio and video formats, as well as static image files. The media library can support a variety of audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0132] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0133] A 2D graphics engine is a drawing engine for 2D drawings.
[0134] The kernel layer is the layer between hardware and software. The kernel layer includes at least display driver, camera driver, audio driver, and sensor driver.
[0135] The following describes the workflow of the software and hardware of the terminal 100 in conjunction with the capture and photo shooting scene.
[0136] When the touch sensor 180K receives a touch operation, the corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including touch coordinates, touch operation timestamp, and other information). The raw input event is stored in the kernel layer. The application framework layer obtains the raw input event from the kernel layer and identifies the control corresponding to the input event. For example, if the touch operation is a touch single-click operation and the control corresponding to the single-click operation is the control of the camera application icon, the camera application calls the interface of the application framework layer to start the camera application, and then starts the camera driver by calling the kernel layer to capture a still image or video through the camera 193.
[0137] The following describes in detail a method for removing image content according to an embodiment of the present application in combination with application scenarios.
[0138] In some application scenarios, before the user uses the terminal 100 to take a photo, the terminal 100 can automatically identify whether there is specified image content (such as a selfie stick) in the preview screen captured by the camera. When the specified image content is identified, the terminal 100 can automatically remove the specified image content in the preview screen and output a removal prompt. The removal prompt is used to remind the user that the specified image content in the preview screen has been removed. After the user presses the shooting button, the terminal 100 can save the preview screen after removing the specified image content as a picture and store the picture in the gallery. When the user turns off the terminal 100's function of removing the specified image content, the terminal 100 can restore the display of the specified image content in the preview screen. In this way, the user can remove the image content that the user does not want when taking a photo, improve the display effect of the image content that the user wants in the photo, and improve the user experience.
[0139] For example, Figure 3AAs shown, the terminal 100 can display an interface 310 of a home screen, in which a page with application icons is displayed, and the page includes multiple application icons (for example, a weather application icon, a stock application icon, a calculator application icon, a settings application icon, an email application icon, a gallery application icon 312, a music application icon, a video application icon, a browser application icon, etc.). A page indicator is also displayed below the multiple application icons to indicate the positional relationship between the currently displayed page and other pages. Below the page indicator are multiple tray icons (for example, a dialing application icon, a message application icon, a contact application icon, a camera application icon 311), and the tray icon remains displayed when the page is switched. In some embodiments, the above-mentioned page may also include multiple application icons and a page indicator. The page indicator may not be part of the page and may exist separately. The above-mentioned tray icon is also optional, and the embodiments of the present application are not limited to this.
[0140] The terminal 100 may receive an input operation (eg, a single click) from the user on the camera application icon 311. In response to the input operation, the terminal 100 may display the following information: Figure 3B The shooting interface 320 is shown.
[0141] like Figure 3B As shown, the shooting interface 320 may include an echo control 321, a shooting control 322, a camera switching control 323, a preview screen 324, a settings control 325, a zoom ratio control 326, and one or more shooting mode controls (e.g., a "night mode" control 372A, a "portrait mode" control 372B, a "cloud enhanced mode" control 372C, a "normal shooting mode" control 372D, a "video recording mode" control 372E, a "professional mode" control 372F, and a "more mode" control 327G). The echo control 321 may be used to display captured images. The shooting control 322 may be used to trigger saving of images captured by a camera. The camera switching control 323 may be used to switch the camera used for shooting. The settings control 325 may be used to set the shooting function. The zoom ratio control 326 may be used to set the zoom ratio of the camera. The shooting mode controls may be used to trigger the start of the image processing process corresponding to the shooting mode. For example, the "night mode" control 372A may be used to trigger increasing the brightness and color richness of the captured image. The “portrait mode” control 372B can be used to trigger the blurring of the background of the person in the captured image. The “cloud enhancement mode” control 372C can be used to trigger the enhancement of the image effect of the captured image with the processing power of the cloud server. Figure 3B As shown, the shooting mode currently selected by the user is "normal shooting mode".
[0142] The terminal 100 can identify whether there is specified image content (such as a selfie stick) in the preview screen. If so, the terminal 100 can remove the specified image content in the preview screen and output an identification prompt, which is used to prompt the user that the specified image content has been identified and is being removed.
[0143] For example, Figure 3C As shown, after the terminal 100 recognizes the selfie stick in the preview screen 324, it can display a prompt 331. The prompt 331 can be used to prompt the user that the selfie stick in the preview screen 324 has been recognized and is being removed from the preview screen. The prompt 331 can be a text prompt (for example, "Selfie stick recognized, removing..."). In some possible implementations, the prompt 331 can also be a prompt of the type of image, video, sound, etc.
[0144] Optionally, during the process of removing specified image content from the preview screen, the terminal 100 can detect whether the movement amplitude of the preview screen is too large. If the movement amplitude is too large, the terminal 100 can output a screen shaking prompt, which can be used to prompt the user to hold the device steady and reduce the shaking amplitude of the preview screen.
[0145] For example, Figure 3D As shown, when the terminal 100 detects whether the movement of the preview image is too large during the process of removing the specified image content in the preview image, it can display a prompt 332. The prompt 332 can be a text prompt (for example, "Selfie stick is being removed. The current image has a large movement. Please hold the device steady."). In some possible implementations, the prompt 331 can also be a prompt of the type of picture, video, sound, etc.
[0146] After removing the specified image content in the preview screen, the terminal 100 can receive user input for the shooting control (such as a single click). In response to this operation, the terminal 100 can save the preview screen after removing the specified image content as a picture and store the picture in the gallery.
[0147] After removing the specified image content from the preview screen, the terminal 100 may display the preview screen after removing the specified image content and a removal close control, which may be used to trigger the terminal 100 to cancel the removal of the specified image content from the preview screen.
[0148] For example, Figure 3EAs shown, after removing the selfie stick from the preview screen 324, the terminal 100 may display a preview screen 328 after the selfie stick is removed and a prompt box 341. Compared with the preview screen 324, the preview screen 328 has the selfie stick removed. The prompt box 341 includes a text prompt (for example, the selfie stick has been removed) and a removal close control 342. The terminal 100 may receive a user input (for example, a single click) on the removal close control 342. In response to the input, the terminal 100 may cancel the removal of the selfie stick from the preview screen 328 and display the following. Figure 3F The preview screen 324 shown in FIG. Figure 3F As shown, the preview image 324 includes a selfie stick.
[0149] In one possible implementation, when the terminal 100 recognizes that there is designated image content in the preview screen of the camera application interface, the terminal 100 may display a removal confirmation control, which can be used to trigger the terminal 100 to remove the designated image content in the preview screen. In this way, before removing the designated image content from the preview screen, the terminal 100 can confirm with the user whether to remove the designated image content. After the user confirms the removal, the terminal 100 removes the designated image content from the preview screen, thereby improving the user experience.
[0150] For example, Figure 4A As shown, when the terminal 100 recognizes that there is a selfie stick in the preview screen 324 of the camera application interface 320, the terminal 100 can display a prompt box 410. The prompt box 410 includes a text prompt (for example, "Selfie stick recognized, do you want to remove it?"), a removal confirmation control 411 and a removal denial control 412. The removal confirmation control 411 can be used to trigger the terminal 100 to remove the specified image content in the preview screen. The removal denial control 412 can trigger the terminal 100 to cancel the removal of the specified image content in the preview screen. The terminal 100 can receive the user's input for the removal confirmation control 411 (for example, a single click). In response to the input, the terminal 100 can remove the selfie stick in the preview screen 324 and replace the preview screen 324 with the following display. Figure 4B The preview screen 328 shown in FIG. The selfie stick is not included in the preview screen 328. Optionally, after the terminal 100 removes the selfie stick from the preview screen 324, a prompt box 421 may be displayed. The prompt box 421 includes a text prompt (e.g., "Selfie stick removed") and a removal close control 422. The removal close control 422 can be used to trigger the terminal 100 to cancel the removal of the specified image content from the preview screen.
[0151] In some embodiments, since the terminal 100 can adopt a solution of removing specified image content (such as a selfie stick) from the preview screen through adjacent frames, the terminal 100 needs to find the portion of the preview screen that is blocked by the specified image content from the adjacent frames. Therefore, the position of the specified image content in the adjacent frames and the position in the preview screen need to be different. When the terminal 100 recognizes that there is specified image content in the preview screen of the camera application interface, the terminal 100 can output an operation prompt, which can be used to prompt the user to move the terminal 100 in a specified direction. In this way, the effect of removing the specified image content can be better.
[0152] For example, Figure 4C As shown, when the terminal 100 recognizes that a selfie stick is present in the preview screen 324 of the camera application interface 320, the terminal 100 may display an operation prompt box 430. The operation prompt box 430 includes a text prompt (e.g., "Selfie stick detected. Please move the phone in the indicated direction") and a direction mark 431 (e.g., a left direction mark). The user can follow the operation prompt box 430 to complete the operation corresponding to the operation prompt box 430 (e.g., move the terminal 100 to the left).
[0153] In a possible implementation, the terminal 100 may display multiple operation prompts in sequence, gradually guiding the user to complete the specified operation. Figure 4D As shown, after the user moves the terminal 100 to the right, the terminal 100 can display the captured frame image 442. After detecting that the terminal 100 has completed the operation corresponding to the above-mentioned operation prompt box 430, the terminal 100 can continue to display the operation prompt box 440 on the camera application interface 320. The operation prompt box 440 includes a text prompt (for example, "Please continue to move the phone in the indicated direction") and a direction mark 441 (for example, a mark for the right direction). After completing the operation corresponding to the operation prompt box 430 (for example, moving the terminal 100 to the right), the user can complete the operation corresponding to the operation prompt box 440 (for example, moving the terminal 100 to the right).
[0154] While the user is moving the terminal 100, the terminal 100 can obtain part of the content in the preview screen that is blocked by the specified image content. After the terminal 100 obtains part of the content in the preview screen that is blocked by the specified image content, the terminal 100 can output an operation completion prompt. The operation completion prompt can be used to prompt the user that the indicated operation has been completed and the specified image content is being removed.
[0155] For example, Figure 4EAs shown, after the terminal 100 obtains the partial content blocked by the specified image content in the preview screen, the terminal 100 can display an operation completion prompt 450. The operation completion prompt 450 can be a text prompt (for example, "You have completed the indicated operation and are removing the selfie stick...").
[0156] In some embodiments, before removing the specified image content from the preview screen, the terminal 100 needs to spend a certain amount of time to first obtain the portion of the preview screen that is obscured by the specified image content. After triggering the start of removing the specified image content from the preview screen, the terminal 100 can output a pre-processing countdown, which can be used to indicate the remaining time for the terminal 100 to complete the removal of the specified image content from the preview screen. This allows the user to experience the removal process of the specified image content.
[0157] For example, Figure 4F As shown, the terminal 100 can display a countdown prompt 460 in the camera application interface 320 after triggering the start of removing the selfie stick in the preview screen. The countdown prompt 460 can be a text prompt (for example, "Preparing to remove the selfie stick, waiting for countdown: 5s").
[0158] like Figure 4G As shown, when the countdown ends, the terminal 100 can complete the removal of the selfie stick from the preview screen 324 and display the preview screen 328 after the selfie stick is removed and a prompt box 471. The prompt box 471 includes a text prompt (e.g., "Preparation is complete, selfie stick removed") and a removal close control 472. The removal close control 472 can be used to cancel the removal of the selfie stick from the preview screen 328.
[0159] In some embodiments, after removing the specified image content in the preview screen, the terminal 100 can display image content with augmented reality (AR) effects (including AR static effect image content and AR dynamic effect image content) in the area where the specified image content was previously located in the preview screen.
[0160] Specifically. The user can place the terminal 100 on a selfie stick, and the user can adjust the shooting angle of the terminal 100 through the selfie stick. The specified image content removed by the terminal 100 can be the selfie stick that appears in the preview screen of the camera application. The terminal 100 can detect whether the user's hand appears around the selfie stick in the preview screen. When the terminal 100 detects that the user's hand appears around the selfie stick in the preview screen, the terminal 100 can display the image content of the AR effect in the area where the hand and the selfie stick are in contact in the preview screen after removing the selfie stick in the preview screen. For example, the terminal 100 can display the flashlight image content in the area where the hand and the selfie stick are in contact in the preview screen using AR technology.
[0161] Specifically, the terminal 100 can apply different AR effects to the area where the selfie stick is located in the preview screen after removing the selfie stick from the preview screen, depending on the scene in the preview screen. For example, when the terminal 100 detects that the scene in the preview screen is daytime, the terminal 100 can remove the selfie stick held by the user from the preview screen and then use AR technology to display a bouquet image in the area where the user's hand is in contact with the selfie stick. When the terminal 100 detects that the scene in the preview screen is nighttime, the terminal 100 can remove the selfie stick held by the user from the preview screen and then use AR technology to display a flashlight image in the area where the user's hand is in contact with the selfie stick.
[0162] In one possible implementation, when a user holds a selfie stick, the terminal 100 may be unable to capture the portion of the hand blocked by the selfie stick in the preview image. When the terminal 100 detects the user's hand around the selfie stick in the preview image, the terminal 100 may remove the selfie stick from the preview image and then perform hand restoration on the hand in the preview image using a separate hand restoration network, generating hand details in the portion blocked by the selfie stick.
[0163] In some embodiments, before removing the specified image content, the terminal 100 may detect whether the light intensity in the preview image is lower than a preset threshold. If so, the terminal 100 may output a fill light prompt, which is used to prompt the user to turn on the fill light to enhance the exposure of the preview image subsequently captured by the terminal 100. Optionally, the terminal 100 may also automatically turn on the fill light on the terminal 100 when detecting that the light intensity in the preview image is lower than a preset threshold, thereby enhancing the exposure of the preview image subsequently captured by the terminal 100. The terminal 100 may also adjust the automatic exposure (AE) strategy of the terminal 100 when detecting that the light intensity in the preview image is lower than a preset threshold, thereby increasing the contrast of the preview image subsequently captured by the terminal 100. In this way, the terminal 100 can improve the removal effect of the specified image content even in scenes with low light intensity (such as at night).
[0164] In one possible implementation, terminal 100 may remove noise from the preview image before removing the specified image content from the preview image. Terminal 100 then removes the specified image content from the preview image using the image content removal process described in subsequent embodiments. The image content removal process can be found in subsequent embodiments and is not detailed here.
[0165] In some embodiments, one or more image contents in the preview screen of the camera application interface can be removed by the terminal 100. The terminal 100 can receive a click operation on the preview screen in the camera application interface. In response to the click operation, the terminal 100 can identify the specific image content selected by the user in the preview screen and display a removal confirmation control. The removal confirmation control can be used to trigger the terminal 100 to remove the specific image content in the preview screen. In this way, the user can select the image content to be removed from the preview screen and remove it.
[0166] For example, Figure 5A As shown, the terminal 100 can receive a user's clicking operation (such as single-click, double-click, long press, etc.) on the preview screen 324 in the camera application interface 320. In response to the clicking operation, the terminal 100 can identify the specified image content selected by the user as a selfie stick based on the position of the clicking operation in the preview screen 324.
[0167] like Figure 5B As shown, after the terminal 100 recognizes that the designated image content selected by the user is a selfie stick, it can display a marking box 511 and a prompt box 520 around the selfie stick. The marking box 511 can be used to prompt the user that the selfie stick in the marking box 511 has been selected. The prompt box 520 includes a text prompt (e.g., "Selfie stick detected, do you want to remove it?"), a removal confirmation control 521, and a removal denial control 522. The removal confirmation control 521 can be used to trigger the terminal 100 to remove the designated image content (e.g., selfie stick) from the preview screen. The removal denial control 522 can trigger the terminal 100 to cancel the removal of the designated image content (e.g., selfie stick) from the preview screen.
[0168] The terminal 100 may receive a user input (e.g., a single click) for the removal confirmation control 521. In response to the input, the terminal 100 may remove the selfie stick from the preview screen 324 and replace the preview screen 324 with the following display: Figure 5C Preview screen 328 is shown.
[0169] like Figure 5C As shown, the preview screen 328 does not include the selfie stick. Optionally, after the terminal 100 removes the selfie stick from the preview screen 324, a prompt box 531 may be displayed. Prompt box 531 includes a text prompt (e.g., "Selfie stick removed") and a removal close control 532. The removal close control 532 can be used to trigger the terminal 100 to cancel the removal of the specified image content from the preview screen.
[0170] In some embodiments, one or more image contents in the preview screen of the camera application interface can be removed by the terminal 100. After identifying the one or more image contents that can be removed in the preview screen, the terminal 100 can mark the one or more image contents that can be removed. The terminal 100 can receive input from the user to select to remove specific image contents from the one or more removable image contents. In response to the input, the terminal 100 can remove the specific image contents from the preview screen. This makes it easier for the user to select the image content they want to remove from the preview screen and remove it.
[0171] For example, Figure 6A As shown, terminal 100 may receive user input (e.g., a single click) selecting object-removal mode control 327H. In response to this input, terminal 100 may switch to object-removal shooting mode. In object-removal shooting mode, after identifying that one or more removable image contents in preview screen 324 include a background person and a selfie stick, terminal 100 may display a label 631 around the background person in the preview screen and a label 621 around the selfie stick in the preview screen. Label 631 may include a descriptive text "background person" and a removal control 632. Removal control 632 may be used to trigger terminal 100 to remove the background person in preview screen 324. Label 621 may include a descriptive text "selfie stick" and a removal control 622. Removal control 622 may be used to trigger terminal 100 to remove the selfie stick in preview screen 324. Optionally, after identifying that one or more removable image contents in the preview screen 324 include the background person and the selfie stick, the terminal 100 may further display a prompt 611, which may be used to prompt the user that the removable image contents in the preview screen have been identified. The prompt 611 may display the text "Removable objects in the screen have been identified."
[0172] The terminal 100 may receive a user input (e.g., a single click) for a removal control. In response to the input, the terminal 100 may remove the image content corresponding to the removal control from the preview screen. Optionally, the terminal 100 may further display an undo control after removing the image content corresponding to the removal control. The undo control may be used to trigger the terminal 100 to undo the removal of the image content.
[0173] For example, Figure 6BAs shown, terminal 100 may receive a user click operation for removal control 622. In response to this click operation, terminal 100 may remove the selfie stick from preview screen 324 and display preview screen 328. Preview screen 328 does not include the selfie stick. Optionally, in response to the click operation for removal control 622, terminal 100 may also replace removal control 622 with an undo control 623 in selfie stick label 621. Undo control 623 may be used to trigger terminal 100 to undo the removal of the selfie stick.
[0174] like Figure 6C As shown, the terminal 100 can respond to the user input (e.g., a click) on the cancel control 623 to cancel the removal of the selfie stick in the preview screen 328, display the preview screen 324, and replace the cancel control 623 with the removal control 622. The preview screen 324 includes the selfie stick.
[0175] In an embodiment of the present application, after the terminal 100 identifies one or more removable image contents in the preview screen and marks the one or more removable image contents, the terminal 100 may further identify a user's gesture or facial expression in the preview screen. The terminal 100 may determine that the image content corresponding to the user's gesture or facial expression is the above-mentioned designated image content.
[0176] Exemplarily, the terminal 100 can identify two image contents, for example, a selfie stick and a background person. The terminal 100 can label these three image contents, the selfie stick can be labeled 1, and the background person can be labeled 2. When the terminal 100 recognizes that the user makes gesture 1 (for example, extends 1 finger) or facial expression action 1 (for example, blinks 2 times in succession), the terminal 100 can determine the selfie stick in the preview screen as the above-mentioned designated image content to be removed. When the terminal 100 recognizes that the user makes gesture 2 (for example, extends 2 fingers) or facial expression action 2 (for example, blinks 3 times in succession), the terminal 100 can determine the background person in the preview screen as the above-mentioned designated image content to be removed. The above examples are only used to explain this application and should not constitute a limitation.
[0177] Optionally, after the terminal 100 identifies one or more image contents that can be removed from the preview screen and marks the one or more image contents that can be removed, the terminal 100 may also receive a user's voice input. In response to the user's voice input, the terminal 100 may analyze the semantics of the user's voice input and determine the designated image content to be removed based on the semantics.
[0178] For example, terminal 100 may identify two image contents, for example, a selfie stick and a person in the background. Terminal 100 may mark the selfie stick and the person in the background in the preview screen. When terminal 100 receives user voice input indicating "remove the selfie stick," terminal 100 may identify the selfie stick as the designated image content to be removed. The above example is provided for illustrative purposes only and should not be construed as limiting the present application.
[0179] Optionally, after the terminal 100 identifies one or more image contents that can be removed from the preview screen and marks the one or more image contents that can be removed, the terminal 100 may also receive a user's input for selecting image contents via a connected device such as a Bluetooth connection. In response to the input, the terminal 100 may determine the designated image contents to be removed.
[0180] For example, terminal 100 is connected to a Bluetooth headset via Bluetooth. Terminal 100 can identify two image contents from the preview screen, for example, a selfie stick and a person in the background. Terminal 100 can mark the selfie stick and the person in the background in the preview screen. When the user taps the Bluetooth headset twice in succession, the Bluetooth headset can send instruction 1 to terminal 100, and terminal 100 can determine, based on instruction 1, that the selfie stick is the designated image content to be removed. When the user taps the Bluetooth headset three times in succession, the Bluetooth headset can send instruction 2 to terminal 100, and terminal 100 can determine, based on instruction 2, that the person in the background is the designated image content to be removed. Examples are provided solely for illustrative purposes and should not be construed as limiting the present application.
[0181] In some embodiments, the user can enable a shooting mode (e.g., selfie stick mode) in the camera application that removes specified image content (e.g., selfie stick mode). After enabling the shooting mode that removes specified image content, the terminal 100 can automatically identify the specified image content in the preview screen and remove the specified image content from the preview screen. In this way, the user can directly set the selfie stick mode that removes specified image content in the camera application, so that the terminal 100 can automatically remove the specified image content in the preview screen, making it convenient for the user to quickly remove unwanted image content.
[0182] For example, Figure 7A As shown, the terminal 100 can receive an input (e.g., a single click) from the user selecting the selfie stick mode control 327I. In response to the input, the terminal 100 can switch from the "normal photo mode" to the "selfie stick mode". In the selfie stick mode, the terminal 100 can automatically remove the selfie stick from the preview screen 324 after recognizing the selfie stick in the preview screen 324. Figure 7BAs shown, after the terminal 100 removes the selfie stick from the preview screen 324, the terminal 100 may display a preview screen 328. The selfie stick is not included in the preview screen 328. The terminal 100 may receive user input (e.g., a single click) on the shooting control 322. In response to the input, the terminal 100 may save the preview screen 328 as a picture.
[0183] In one possible implementation, a user can enable a shooting mode (e.g., selfie stick mode) in a camera application that removes specified image content (e.g., selfie stick mode). After enabling the shooting mode that removes specified image content, the terminal 100 does not remove the specified image content from the preview screen until the terminal 100 receives user input on the shooting control. In response to the user input received on the shooting control, the terminal 100 can obtain a target image from the preview screen, remove the specified image content from the target image, and save the target image after removing the specified image content to the local device of the terminal 100.
[0184] For example, Figure 7C As shown, in the selfie stick mode, the terminal 100 currently displays a preview screen 324. The terminal 100 may receive a user input (e.g., a single click) on the shooting control 322. In response to the input, the terminal 100 may use the preview screen 324 as the target image and remove the selfie stick from the target image.
[0185] like Figure 7D As shown, during the process of removing the selfie stick from the target picture, the terminal 100 may output a prompt 711, which may be used to remind the user that the selfie stick from the target picture is being removed. The prompt 711 may be a text prompt, for example, "Removing the selfie stick from the picture..."
[0186] like Figure 7E As shown, after the terminal 100 removes the selfie stick from the target picture, the terminal 100 can save the target picture after the selfie stick is removed to the gallery, and display the thumbnail corresponding to the target picture after the selfie stick is removed on the echo control 321. The terminal 100 can receive the user's input (such as a click) on the echo control 321, and in response to the input, the terminal 100 can display the following Figure 7F The picture browsing interface 730 is shown.
[0187] like Figure 7FAs shown, the image browsing interface 730 includes an image 731 and a menu 732. Image 731 is the target image after the selfie stick is removed. Menu 732 may include a share button, a favorite button, an edit button, a delete button, and an more button. The share button can be used to trigger sharing of image 731. The favorite button can be used to trigger adding image 731 to a favorites folder. The edit button can be used to trigger editing functions such as rotating, cropping, adding filters, and blurring image 731. The delete button can be used to trigger deletion of image 731. The more button can be used to trigger access to more functions related to image 731.
[0188] In some embodiments, when a user uses terminal 100 to record a video, terminal 100 can identify whether there is specified image content (such as a selfie stick) in the frame of the recorded video. When the specified image content is identified, terminal 100 can remove the specified image content from the frame of the recorded video and display the frame after the specified image content is removed. In this way, the image content that the user does not want in the recorded video can be removed in real time while the user is recording, improving the display effect of the image content that the user wants in the recorded video, and enhancing the user experience.
[0189] For example, Figure 8A As shown, the terminal 100 may display the camera application interface 320. The terminal 100 may receive input (e.g., a single click) from the user selecting the recording control 327E. In response to this input, the terminal 100 may switch from "normal photo mode" to "recording mode" and replace the shooting control 322 with the recording start control 801. The terminal 100 may also display recording time information 802. In recording mode, the terminal 100 may recognize a selfie stick in the preview screen 324 and output a prompt box 810. The prompt box 810 includes a text prompt (e.g., "Selfie stick detected, do you want to remove it?"), a removal confirmation control 811, and a removal denial control 812. The removal confirmation control 811 may be used to trigger the terminal 100 to remove a specified image content from the preview screen. The removal denial control 812 may trigger the terminal 100 to cancel the removal of the specified image content from the preview screen.
[0190] The terminal 100 may receive a user input (eg, a single click) for the removal confirmation control 811. In response to the input, the terminal 100 may remove the selfie stick from the preview screen 324 and replace the preview screen 324 with the following display: Figure 8B Preview screen 328 is shown.
[0191] like Figure 8BAs shown, the preview screen 328 does not include the selfie stick. Optionally, after the terminal 100 removes the selfie stick from the preview screen 324, a prompt box 821 may be displayed. This prompt box 821 includes a text prompt (e.g., "Selfie stick removed") and a removal close control 822. This removal close control 822 can be used to trigger the terminal 100 to cancel the removal of the specified image content from the preview screen.
[0192] The terminal 100 may receive user input (eg, a single click) on the video start control 801 . In response to the input, the terminal 100 may start video recording and remove designated image content from each frame during the video recording process.
[0193] like Figure 8C As shown, after starting recording, terminal 100 can replace recording start control 801 with recording end control 803. This recording end control 803 can be used to trigger terminal 100 to end recording. After starting recording, terminal 100 can remove the selfie stick from each frame during the recording process. For example, frame 823 displayed by terminal 100 at the 10th second of recording does not contain the selfie stick.
[0194] The terminal 100 may receive an input (eg, a single click) from the user on the video recording end control 803 . In response to the input, the terminal 100 may end the video recording and save the video after the selfie stick is removed.
[0195] In some application scenarios, after taking a picture or video, the terminal 100 can save the captured picture or video locally. The user can view the pictures or videos captured by the terminal 100, as well as pictures or videos obtained from other devices or the Internet, in the terminal 100's gallery application. The terminal 100 can remove specified image content from the stored pictures or videos. This allows the user to remove unwanted image content from the captured picture or video at any time after taking the picture or video.
[0196] For example, Figure 9A As shown, the terminal 100 can display the interface 310 of the main screen. For the text description of the interface 310, please refer to the aforementioned Figure 3A The embodiments shown will not be described in detail here.
[0197] The terminal 100 may receive an input (eg, a single click) from the user on the gallery application icon 312. In response to the input, the terminal 100 may display the following information: Figure 9B The gallery application interface 910 is shown.
[0198] like Figure 9BAs shown, the gallery application interface 910 can display one or more albums (for example, all photo albums, video albums 917, camera albums, continuous shot albums 916, WeChat albums, Weibo albums, etc.). The terminal 100 can display a gallery menu 911 below the gallery album interface 910. Among them, the gallery menu 911 includes a photo control 912, an album control 913, a moment control 914, and a discovery control 915. Among them, the photo control 912 is used to trigger the terminal 100 to display all local pictures in the form of picture thumbnails. The album control 913 is used to trigger the terminal 100 to display the album to which the local pictures belong. As shown Figure 9B As shown, the current album control 913 is in a selected state, and the terminal 100 displays the gallery application interface 910. The moment control 914 can be used to trigger the terminal 100 to display the selected pictures stored locally. The discovery control 915 can be used to trigger the terminal 100 to display the classified albums of pictures.
[0199] The terminal 100 may receive a user input (eg, a single click) for the continuous photo album 916. In response to the input, the terminal 100 may display the following information: Figure 9C The continuous photo album interface 920 is shown.
[0200] like Figure 9C As shown, the continuous shooting album interface 920 may include thumbnails of one or more pictures (e.g., thumbnail 921 and thumbnail 922). In one possible implementation, the picture corresponding to thumbnail 921 and the picture corresponding to thumbnail 922 may be two pictures continuously shot by the terminal 100.
[0201] The terminal 100 may receive an input (such as a click) from the user on the thumbnail 921. In response to the input, the terminal 100 may display the following information: Figure 9D The picture browsing interface 930 is shown.
[0202] like Figure 9D As shown, the picture browsing interface 930 may include a picture 931, a menu 932, and a return control 933. The picture 931 may be the picture corresponding to the thumbnail 921 mentioned above. The menu 932 may include a share button, a favorite button, an edit button, a delete button, and a more button. The share button can be used to trigger the terminal 100 to share the picture 931. The favorite button can be used to trigger the terminal 100 to favorite the picture 931 to a picture favorite folder. The edit button can be used to trigger the terminal 100 to perform editing functions such as rotating, cropping, adding filters, blurring, etc. on the picture 931. The delete button can be used to trigger the deletion of the picture 931. The more button can be used to trigger the opening of more functions related to the picture 931.
[0203] The terminal 100 can identify whether a picture displayed on the picture browsing interface contains a specified pattern (e.g., a selfie stick). If so, the terminal 100 can display an identification prompt and a removal control on the picture browsing interface. The identification prompt can be used to prompt the user that the picture displayed on the picture browsing interface has been identified as containing the specified image content. The removal control can be used to trigger the terminal 100 to remove the specified image content from the picture displayed on the picture browsing interface.
[0204] For example, Figure 9D As shown, when the terminal 100 recognizes that a selfie stick is included in a picture 931 displayed on the picture browsing interface 930, the terminal 100 may display a prompt 941 and a removal control 942. The prompt 941 may be a text prompt, for example, "A selfie stick has been identified in the picture. You can choose to remove it." The text "Remove selfie stick" may be displayed around the removal control 942.
[0205] The terminal 100 can receive user input (such as a single click) on the removal control. In response to the input, the terminal 100 can remove the specified image content (such as a selfie stick) in the picture displayed on the above-mentioned removal picture browsing interface, and display the picture after removing the specified image content.
[0206] For example, the terminal 100 responds to the received Figure 9D When the click operation of removing the control 942 is performed, the terminal 100 can remove the above Figure 9D The selfie stick in the picture 931 is shown in FIG. Figure 9E The picture 934 shown in FIG. 934 is the picture obtained by removing the selfie stick from the picture 931. Figure 9E As shown, the terminal 100 can also remove the above Figure 9D After removing the selfie stick in picture 931, prompt 943, cancel control 944, and save control 945 are displayed. Prompt 943 can be used to prompt the user that specified image content has been removed from the picture displayed on the picture browsing interface. For example, prompt 943 can be a text prompt such as "Selfie stick removed from picture." Cancel control 944 can be used to trigger terminal 100 to cancel the removal of specified image content from the picture displayed on the picture browsing interface.
[0207] The terminal 100 may receive an input (eg, a click) from the user on the save control 945. In response to the input, the terminal 100 may save the image after removing the specified image content (eg, the selfie stick) to the local computer. Figure 9FAs shown, the terminal 100 may display a thumbnail 923 corresponding to the picture after the specified image content is removed in the continuous shot album interface 920 in the gallery application. The terminal 100 may mark the thumbnail 923. For example, the terminal 100 may display a text mark "Selfie 1 (without selfie stick)" below the thumbnail 923.
[0208] In some embodiments, after displaying the image browsing interface, the terminal 100 may identify one or more removable image contents in the image displayed on the image browsing interface and mark the one or more removable image contents. The terminal 100 may receive input from the user selecting to remove a specific image content from the one or more removable image contents. In response to the input, the terminal 100 may remove the specific image content from the image. This makes it easier for the user to select and remove the image content they wish to remove from the image displayed on the image browsing interface.
[0209] For example, Figure 10A As shown, after displaying image browsing interface 930, terminal 100 may identify one or more removable image contents in image 931 as including a background person and a selfie stick. Terminal 100 may display a label 1031 around the background person in image 931 and a label 1021 around the selfie stick in image 931. Label 1031 may include a descriptive text "background person" and a removal control 1032. Removal control 1032 may be used to trigger terminal 100 to remove the background person in image 931. Label 1021 may include a descriptive text "selfie stick" and a removal control 1022. Removal control 1022 may be used to trigger terminal 100 to remove the selfie stick in image 931. Optionally, after identifying one or more removable image contents in image 931 as including the background person and the selfie stick, terminal 100 may also display a prompt 1011 to inform the user that removable image contents in image 931 have been identified. The prompt 1011 may display the text "Removable objects have been identified in the image, you can remove them selectively."
[0210] The terminal 100 may receive a user input (e.g., a single click) for a removal control. In response to the input, the terminal 100 may remove the image content corresponding to the removal control from the image displayed on the gallery browsing interface. Optionally, the terminal 100 may further display an undo control after removing the image content corresponding to the removal control. The undo control may be used to trigger the terminal 100 to undo the removal of the image content.
[0211] For example, Figure 10BAs shown, in response to a single click on the remove control 1022, the terminal 100 can remove the selfie stick from image 931 and display image 934. Image 934 does not include the selfie stick. Alternatively, in response to a single click on the remove control 622, the terminal 100 can replace the remove control 1022 with a cancel control 1023 in the selfie stick label 1021. This cancel control 1023 can be used to trigger the terminal 100 to cancel the removal of the selfie stick.
[0212] In some embodiments, after the terminal 100 turns on the object removal function in the camera application interface, the terminal 100 can identify one or more image contents that can be removed in the preview screen and display the removal mode controls corresponding to each of the one or more image contents. The terminal 100 can receive the user's input for the removal mode control corresponding to the specified image content. In response to the input, the terminal 100 can remove the specified image content in the preview screen. Then, the terminal 100 can receive the user's input for the shooting control. In response to the input, the terminal 100 can save the preview screen after the specified image content is removed as a picture. The user can view the picture after the specified image content is removed through the echo control on the camera application interface. The terminal 100 can mark other removable image contents in the picture after the specified image content has been removed for the user to choose to remove.
[0213] For example, Figure 10C As shown, terminal 100 may have switched to object-removal shooting mode. In object-removal shooting mode, after terminal 100 identifies one or more removable image content in preview screen 324 as including background people and selfie sticks, it may display a removal mode selection box 1040. Removal mode selection box 1040 includes a text prompt, a selfie stick removal control 1041, and a background person removal control 1042. For example, the text prompt may read "Removable content has been identified in the image. You can select the corresponding removal mode."
[0214] The terminal 100 may receive user input (eg, a single click) for a removal mode control. In response to the input, the terminal 100 may enter a removal mode corresponding to the removal mode control and remove image content corresponding to the removal mode in the preview screen.
[0215] For example, Figure 10D As shown, after the terminal 100 receives the user's selection of the selfie stick removal control 1041, the terminal 100 can enter the selfie stick removal mode and remove the selfie stick from the preview screen. The terminal 100 can display a prompt 1051 during the selfie stick removal process, which can be used to notify the user that the terminal 100 is removing the selfie stick from the preview screen 324.
[0216] The terminal 100 may receive a user input (e.g., a single click) for the removal confirmation control 521. In response to the input, the terminal 100 may remove the selfie stick from the preview screen 324 and replace the preview screen 324 with the following display: Figure 10E Preview screen 328 is shown.
[0217] like Figure 10E As shown, the preview screen 328 does not include the selfie stick. Optionally, after the terminal 100 removes the selfie stick from the preview screen 324, a prompt box 1052 may be displayed. This prompt box 1053 includes a text prompt (e.g., "Selfie stick removed") and a removal close control 1053. This removal close control 1053 can be used to trigger the terminal 100 to cancel the removal of the specified image content from the preview screen.
[0218] like Figure 10F As shown, after the terminal 100 removes the selfie stick, the terminal 100 can receive the user's input (e.g., click) on the shooting control 322. In response to the input, the terminal 100 can save the preview image 328 as the target image and display the thumbnail of the target image in the echo control 321. The terminal 100 can receive the user's input (e.g., click) on the echo control 321. In response to the input, the terminal 100 can display the following image: Figure 10G The picture browsing interface 730 is shown.
[0219] like Figure 10G As shown, the image browsing interface 930 includes an image 1061 and a menu 932. Image 1061 is the target image after the selfie stick is removed. Terminal 100 can use the frames buffered during the selfie stick removal process as reference images to identify and mark the removable image content in image 1061. For example, after terminal 100 identifies a removable background person in image 1061, it can display a prompt 1073 and a label 1071 around the background person in image 1061. Prompt 1073 can be used to inform the user that removable image content has been identified in image 1061. Prompt 1071 can display the text "Removable content has been identified in the image. You can further choose to remove it." Label 1071 can include the descriptive text "Background person" and a removal control 1072. Removal control 1072 can be used to trigger terminal 100 to remove the background person in image 1061.
[0220] In some embodiments, one or more image contents in a picture displayed on the picture browsing interface can be removed by the terminal 100. The terminal 100 can receive a click operation on a picture displayed on the picture browsing interface. In response to the click operation, the terminal 100 can identify the specified image content (e.g., a selfie stick) selected by the user in the picture displayed on the picture browsing interface and display a removal confirmation control. The removal confirmation control can be used to trigger the terminal 100 to remove the specified image content from the picture displayed on the picture browsing interface. In this way, the user can select the image content to be removed from the preview screen and remove it.
[0221] For example, Figure 11A As shown, the terminal 100 can receive a user's clicking operation (such as single-click, double-click, long press, etc.) on the picture 931 in the gallery browsing interface 930. In response to the clicking operation, the terminal 100 can identify the specified image content selected by the user as a selfie stick based on the position of the clicking operation in the picture 931.
[0222] like Figure 11B As shown, after the terminal 100 recognizes that the designated image content selected by the user is a selfie stick, it can display a marking box 1111 and a prompt box 1120 around the selfie stick. The marking box 1111 can be used to prompt the user that the selfie stick in the marking box 1111 has been selected. The prompt box 1120 includes a text prompt (e.g., "We recognize that you have selected the selfie stick. Do you want to remove it?"), a removal confirmation control 1121, and a removal negation control 1122. The removal confirmation control 1121 can be used to trigger the terminal 100 to remove the designated image content (e.g., the selfie stick) from the image 931. The removal negation control 1122 can trigger the terminal 100 to cancel the removal of the designated image content (e.g., the selfie stick) from the preview screen.
[0223] The terminal 100 may receive a user input (eg, a single click) for the removal confirmation control 1121. In response to the input, the terminal 100 may remove the selfie stick in the image 931 and replace the image 931 with the following image: Figure 11C Picture 934 shown.
[0224] like Figure 11C As shown, the picture 934 is the picture obtained by removing the selfie stick from the picture 931. Optionally, after removing the selfie stick from the picture 931, the terminal 100 may also display a prompt 943, a cancel control 944, and a save control 945. The text descriptions of the prompt 943, the cancel control 944, and the save control 945 can refer to the aforementioned Figure 9E The embodiments shown will not be described in detail here.
[0225] In some embodiments, the terminal 100 may store a video locally. The video may be shot by the terminal 100, sent by another device, or downloaded from the Internet. Because the video contains specific image content that affects the overall viewing quality of the video, the terminal 100 may remove the specific image content from the stored pictures or videos. This allows the user to easily remove unwanted image content from the video at any time after shooting.
[0226] For example, Figure 12A As shown, the terminal 100 can display the gallery application interface 910. For the text description of the gallery album interface 910, please refer to the aforementioned Figure 9B The embodiments shown will not be described in detail here.
[0227] The terminal 100 may receive an input (such as a click) from the user on the video album 917. In response to the input, the terminal 100 may display the video album 917. Figure 12B The video album interface 1210 is shown.
[0228] like Figure 12B As shown, the video album interface 1210 includes thumbnails corresponding to one or more videos, for example, thumbnail 1211, thumbnail 1212, thumbnail 1213 and thumbnail 1214, and so on. Each thumbnail on the video album interface 1210 can also display the time length of the video corresponding to the thumbnail. For example, the time length of the video corresponding to thumbnail 1211 is 10s, the time length of the video corresponding to thumbnail 1212 is 15s, the time length of the video corresponding to thumbnail 1213 is 30s, and the time length of the video corresponding to thumbnail 1214 is 45s. The above examples are only used to explain this application and should not constitute a limitation.
[0229] The terminal 100 may receive an input (such as a click) from the user on the thumbnail 1211. In response to the input, the terminal 100 may display the following information: Figure 12C The video browsing interface 1220 is shown.
[0230] like Figure 12CAs shown, the video browsing interface 1220 may include a video 1221, a menu 1222, and a return control 1223. Among them, the video 1221 is the video corresponding to the above-mentioned thumbnail 1211. The menu 1222 may include a share button, a favorite button, an edit button, a delete button, and a more button. The share button can be used to trigger the terminal 100 to share the video 1221. The favorite button can be used to trigger the terminal 100 to favorite the video 1221 to a video favorite folder. The edit button can be used to trigger the terminal 100 to perform editing functions such as rotating, trimming, adding filters, blurring, etc. on the video 1221. The delete button can be used to trigger the deletion of the video 1221. The more button can be used to trigger the opening of more functions related to the video 1221.
[0231] The terminal 100 can identify whether a frame of a video displayed on the video browsing interface contains a specified pattern (e.g., a selfie stick). If so, the terminal 100 can display an identification prompt and a removal control on the video browsing interface. The identification prompt can be used to prompt the user that the specified image content has been identified in the frame of the video displayed on the video browsing interface. The removal control can be used to trigger the terminal 100 to remove the specified image content from the video displayed on the video browsing interface.
[0232] For example, Figure 12C As shown, when the terminal 100 recognizes that a selfie stick is present in a video 1221 displayed on the video browsing interface 1220, the terminal 100 may display a prompt 1231 and a removal control 1232. The prompt 1231 may be a text prompt, for example, "A selfie stick has been detected in the video. You may choose to remove it." The text "Remove selfie stick" may be displayed around the removal control 1232.
[0233] The terminal 100 can receive user input (such as a single click) on the removal control. In response to the input, the terminal 100 can remove the specified image content (such as a selfie stick) in the picture displayed on the above-mentioned removal picture browsing interface, and display the picture after removing the specified image content.
[0234] For example, the terminal 100 responds to the received Figure 12C When the click operation of removing the control 1232 is performed, the terminal 100 can remove the above Figure 12C The selfie stick in the video 1231 is shown in FIG. Figure 12D The video 1223 shown is a video obtained by removing the selfie stick from the video 1221. Figure 12D As shown, the terminal 100 can also remove the above Figure 12CAfter removing the selfie stick from video 1221, prompt 1241, undo control 1242, and save control 1243 are displayed. Prompt 1241 can be used to inform the user that specified image content has been removed from the image displayed on the image browsing interface. For example, prompt 1241 can be a text prompt such as "Selfie stick removed from image." Undo control 1242 can be used to trigger terminal 100 to undo the removal of specified image content from the image displayed on the image browsing interface. Save control 1243 can be used to trigger terminal 100 to save video 1223.
[0235] The following describes a process in which the terminal 100 removes specified image content from a picture in an embodiment of the present application.
[0236] Figure 13 FIG. 1 is a schematic diagram showing the architecture of an image content removal system 1300 provided in an embodiment of the present application. The image content removal system 1300 can be applied to the terminal 100 described above.
[0237] like Figure 13 As shown, the image content removal system 1300 may include a picture segmentation module 1301, a coarse restoration module 1302, a mask image generation module 1303 and a fine restoration module 1304.
[0238] Image segmentation module 1301 can be used to segment a first region of a first target image containing designated image content (e.g., a selfie stick) to obtain a second target image. Segmentation module 1301 can also be used to segment a second region of a first reference image containing designated image content (e.g., a selfie stick) to obtain a second reference image.
[0239] The rough restoration module 1302 can be used to find content similar to features around the first region in the second reference image from the second reference image, fill the first region in the second target image, and generate a third target image. The features include texture, color, shape, etc.
[0240] The mask image generating module 1303 may be configured to generate a mask image of the second target image according to the second target image.
[0241] Specifically, the mask image generating module 1303 may be configured to convert the display color of the first region in the second target image to white, and convert the display color of the non-first region in the second target image to black.
[0242] The fine restoration module 1304 can be used to optimize and generate the texture in the first region of the third target image according to the mask image of the second target image and the third target image, so as to obtain a fourth target image.
[0243] For example, Figure 14A As shown in FIG, the first target image includes a selfie stick. Figure 14B As shown in FIG, the selfie stick in the first area of the second target image has been segmented out, and the first area can be filled with black. Figure 14C As shown, the first reference image includes a selfie stick. Figure 14D As shown in FIG, the selfie stick in the second region of the second reference image has been segmented out, and the second region can be filled with black. Figure 14E As shown, the first area in the mask image of the second target image can be filled with white, and the non-first area in the mask image of the second target image can be filled with black. Figure 14F As shown in FIG, the first area in the third target image has been filled with content similar to the features around the first area determined from the second reference image. Figure 14G As shown, the texture, edge, and details in the first region of the fourth target image have been optimized.
[0244] The designated image content may be a default image on the terminal 100 or may be selected and input by the user. The designated image content may include one or more of the following image contents: a selfie stick, a person in the background, glasses, and the like.
[0245] The first target image may be a target frame image captured by the camera of the terminal 100, and the first reference image may be an adjacent frame image of the target frame image. For example, the first target image may be the above Figure 3B The preview screen 324 shown, or the above Figure 8A The preview screen 324 shown, or the above Figures 8B to 8C Each frame of the picture captured by the camera of the terminal 100 during the video recording process, etc. For another example, the first target image can be the above Figure 7C When the terminal 100 receives the user's input for the shooting control 322, the preview image 324 captured by the camera of the terminal 100.
[0246] In some embodiments, the first target image may also be a picture saved in the gallery application of the terminal 100, and the first reference image may be a continuous shot of the saved picture. Figure 9C The picture corresponding to the thumbnail 921 shown above Figure 9C The picture corresponding to the thumbnail 922 is shown, and so on.
[0247] In some embodiments, the first target image may also be any frame during the recording process of the terminal 100, and the first reference image may be an adjacent frame of the frame during the recording process. Figures 8B to 8C Any frame captured by the camera of the terminal 100 during the video recording process, etc.
[0248] In some embodiments, the first target image may also be any frame in a video stored on the terminal 100, and the first reference image may be an adjacent frame of the frame in the video. Figure 12C Any frame in the video 1221 shown, and so on.
[0249] Specifically, the image segmentation module 1301 can perform feature matching with the first target image based on the feature information of the specified image content (for example, a selfie stick) acquired in advance, determine the area in the first target image where the specified image content is located, and segment the area where the specified image content is located from the first target image to obtain the second target image. The image segmentation module 1301 can perform feature matching with the first reference image based on the feature information of the specified image content acquired in advance, determine the area in the first reference image where the specified image content is located, and segment the area where the specified image content is located from the first target image to obtain the second reference image.
[0250] In one possible implementation, the image segmentation module 1301 may further use a trained segmentation neural network to identify, based on the RGB information of the first target image, a first region in the first target image where the specified image content (e.g., a selfie stick) is located, and segment the first region from the first target image to obtain a second target image. The image segmentation module 1301 may further use a trained segmentation neural network to identify, based on the RGB information of the first reference image, a second region in the first reference image where the specified image content (e.g., a selfie stick) is located, and segment the second region from the first reference image to obtain a second reference image.
[0251] In one possible implementation, the image segmentation module 1301 may further use a trained segmentation neural network to identify a first region in the first target image where specified image content (e.g., a selfie stick) is located based on the RGB information, depth of field information, and confidence information of the first target image, and segment the region where the specified image content is located from the first target image to obtain a second target image. The image segmentation module 1301 may further use a trained segmentation neural network to identify a second region in the first reference image where specified image content (e.g., a selfie stick) is located based on the RGB information, depth of field information, and confidence information of the first reference image, and segment the second region from the first reference image to obtain a second reference image.
[0252] In one possible implementation, the image segmentation module 1301 may further use a trained segmentation neural network to identify a first region in the first target image where specified image content (e.g., a selfie stick) is located based on the RGB information and thermal imaging information of the first target image, and segment the region where the specified image content is located from the first target image to obtain a second target image. The image segmentation module 1301 may further use a trained segmentation neural network to identify a second region in the first reference image where specified image content (e.g., a selfie stick) is located based on the RGB information and thermal imaging information of the first reference image, and segment the second region from the first reference image to obtain a second reference image.
[0253] When training the segmentation neural network, the training device can expand the training data by adjusting image contrast, etc., to increase the richness of the training data, so that the segmentation neural network can better segment the specified image content in the input image under drastic changes in the shooting environment of the input image. The segmentation neural network can be a convolutional neural network, such as an SSD network, a Faster-RCNN network, etc.
[0254] In the embodiment of the present application, the image content removal system 1300 can be applied on the terminal 100 .
[0255] In one possible implementation, the image content removal system 1300 can be applied on a server, and the terminal 100 can send the first target image and the first reference image to the server. The server can remove the specified image content (such as a selfie stick) in the first target image based on the first target image and the first reference image to obtain a fourth target image, and send the fourth target image to the terminal 100.
[0256] In one possible implementation, the image content removal system 1300 can be applied on a server and a terminal 100. Some functional modules of the image content removal system 1300 may reside on the server, while the remaining functional modules may reside on the terminal 100. For example, the terminal 100 may include an image segmentation module 1301, and the server may include a coarse restoration module 1302, a mask map generation module 1303, and a fine restoration module 1304. After acquiring a first target image and a first reference image, the terminal 100 may use the image segmentation module 1301 to segment a first region containing specified image content in the first target image to obtain a second target image, and segment a second region containing specified image content in the first reference image to obtain a second reference image. The terminal 100 then sends the second target image and the second reference image to the server. The server may process the second target image and the second reference image using the coarse restoration module 1302, the mask map generation module 1303, and the fine restoration module 1304 to obtain a fourth target image, and then send the fourth target image to the terminal 100. The examples are only used to explain the present application and should not be construed as limiting. In specific implementations, the functional modules included in the image content removal system 1300 may also have other distribution methods on the server and the terminal 100, which will not be described in detail here.
[0257] The following describes the optical flow rough restoration process in an embodiment of the present application.
[0258] Figure 15 A schematic structural diagram of a rough repair module 1302 in an embodiment of the present application is shown.
[0259] like Figure 15 As shown, the above-mentioned rough repair module 1302 may include an optical flow network 1501, an optical flow completion model 1502 and a filling module 1503. Among them:
[0260] The optical flow network 1501 can be used to calculate the missing optical flow information between the second target image and the second reference image. The optical flow can be used to indicate the instantaneous speed of the pixel motion of the moving object in the two images on the observation imaging plane.
[0261] The optical flow completion model 1502 can be used to complete the missing optical flow information between the second target image and the second target image based on the second reference image, so as to obtain complete optical flow information between the second target image and the second reference image.
[0262] The filling module 1503 can be used to determine the filling pixel information in the second reference image that needs to be filled in the first area of the second target image through the complete optical flow information, and fill the pixels in the first area of the second target image through the filling pixel information to obtain the third target image.
[0263] In the embodiment of the present application, the optical flow network 1501 can adopt an optical flow network such as flownet or flownet2.
[0264] The following describes the multi-frame feature coarse repair process in an embodiment of the present application.
[0265] Figure 16 A schematic structural diagram of another rough repair module 1302 in an embodiment of the present application is shown.
[0266] like Figure 16 As shown, the rough repair module 1302 may include an encoder 1601, an attention mechanism module 1602, a feature filling module 1603, and a decoder 1604.
[0267] The encoder 1601 can be used to encode the second target image into a first target feature map and encode the second reference image into a first reference feature map. Figure 17A As shown, the second target feature map can be referred to Figure 17B The examples shown are only used to explain the present application and should not be construed as limiting.
[0268] The attention mechanism module 1602 may be configured to find feature information similar to features surrounding the first region in the first target feature map from the first reference feature map based on the first target feature map and the first reference feature map, wherein the feature information may include texture, color, shape, and the like.
[0269] The feature filling module 1603 can be used to fill the first area of the first target feature map with feature information in the first reference feature map that is similar to features around the first area in the first target feature map to obtain a second target feature map.
[0270] The decoder 1604 may be used to decode the second target feature map into a third target image.
[0271] The following describes the single-frame feature coarse repair process in the embodiment of the present application.
[0272] Figure 18 A schematic structural diagram of another rough repair module 1302 in an embodiment of the present application is shown.
[0273] like Figure 18 As shown, the above-mentioned rough repair module 1302 may include an encoder 1801, an attention mechanism module 1802, a feature filling module 1803 and a decoder 1804. Among them:
[0274] The encoder 1801 can be used to encode the second target image into the first target feature map. Figure 17A The examples shown are only used to explain the present application and should not be construed as limiting.
[0275] The attention mechanism module 1802 can be used to find feature information similar to features around the first region from the first target feature map, wherein the feature information includes texture, color, shape, etc.
[0276] The feature filling module 1803 can be used to fill the first region of the first target feature map with feature information similar to features around the first region in the first target feature map to obtain a second target feature map.
[0277] The decoder 1604 may be used to decode the second target feature map into a third target image.
[0278] In an embodiment of the present application, when the first target image is a target frame captured by the camera of the terminal 100 and the first reference image is an adjacent frame of the target frame, the above-mentioned image content removal system 1300 may further include a motion detection module 1305.
[0279] like Figure 19 As shown, the motion detection module 1305 can be used to determine whether the shooting picture of the terminal 100 is moving significantly based on the motion data obtained from the inertial measurement unit (IMU) of the terminal 100. If so, the rough repair module 1302 can use the above Figure 16 The structure shown in FIG. 1 is used to perform multi-frame coarse restoration on the second target image according to the second target image and the second reference image. If not, the coarse restoration module 1302 can adopt the above Figure 15 The structure shown in FIG1 is used to perform optical flow coarse restoration on the second target image based on the second target image and the second reference image. The motion data includes angular velocity data and acceleration data of the terminal 100. For example, when the angular velocity of any of the three directions of the terminal 100 is greater than a specified angular velocity value, or when the acceleration of any of the three directions of the terminal 100 is greater than a specified acceleration value, the motion detection module 1305 can determine that the captured image of the terminal 100 is in large-scale motion. When the motion data is otherwise, the motion detection module 1305 can determine that the captured image of the terminal 100 is in small-scale motion.
[0280] In a possible implementation, the motion detection module 1305 can be used to determine whether the captured image of the terminal 100 has a large movement based on the intersection over union (IoU) between the mask image of the second target image and the mask image of the second reference image. If so, the coarse repair module 1302 can use the above Figure 16 The structure shown in FIG. 1 is used to perform multi-frame coarse restoration on the second target image according to the second target image and the second reference image. If not, the coarse restoration module 1302 can adopt the above Figure 15 The structure shown in FIG1 is used to perform coarse optical flow restoration on the second target image based on the second target image and the second reference image. For example, when the intersection-and-union ratio between the mask image of the second target image and the mask image of the second reference image is less than a specified intersection-and-union ratio, the motion detection module 1305 can determine that the captured image of the terminal 100 is in large-scale motion. When the intersection-and-union ratio between the mask image of the second target image and the mask image of the second reference image is greater than or equal to the specified intersection-and-union ratio, the motion detection module 1305 can determine that the captured image of the terminal 100 is in small-scale motion.
[0281] In a possible implementation, the motion detection module 1305 can be used to determine whether the shooting picture of the terminal 100 has a large movement based on the similarity between the first target feature map and the first reference feature map. If so, the coarse repair module 1302 can use the above Figure 16 The structure shown in FIG. 1 is used to perform multi-frame coarse restoration on the second target image according to the second target image and the second reference image. If not, the coarse restoration module 1302 can adopt the above Figure 15 The structure shown in FIG1 performs coarse optical flow restoration on the second target image based on the second target image and the second reference image. For example, when the similarity between the first target feature map and the first reference feature map is less than a specified similarity value, the motion detection module 1305 can determine that the image captured by the terminal 100 is in significant motion. When the similarity between the first target feature map and the first reference feature map is greater than or equal to the specified similarity value, the motion detection module 1305 can determine that the image captured by the terminal 100 is in minor motion.
[0282] In some embodiments, the motion detection module 1305 can also determine whether the shooting picture of the terminal 100 has a large movement based on the motion data obtained from the IMU of the terminal 100, the intersection-over-union ratio between the mask map of the second target image and the mask map of the second reference image, and the similarity between the first target feature map and the first reference feature map. If so, the coarse repair module 1302 can use the above Figure 16 The structure shown in FIG. 1 is used to perform multi-frame coarse restoration on the second target image according to the second target image and the second reference image. If not, the coarse restoration module 1302 can adopt the above Figure 15The structure shown performs optical flow coarse restoration on the second target image based on the second target image and the second reference image.
[0283] The following describes an image content removal method provided in an embodiment of the present application.
[0284] Figure 20 A flow chart of an image content removal method provided in an embodiment of the present application is shown.
[0285] like Figure 20 As shown, the method includes:
[0286] S2001. Terminal 100 obtains a first target image and a first reference image.
[0287] The first target image may be a first preview image captured by a camera of terminal 100, and the first reference image may be a first reference frame image captured by the camera before or after the first preview image. The first preview image and the first reference frame image both include image content of a first object and image content of a second object, and in the first preview image, the image content of the first object partially obscures an image of the second object.
[0288] For example, the first preview screen can be the above Figure 3B The preview screen 324 shown, or the above Figure 8A The preview screen 324 shown, or the above Figures 8B to 8C Each frame of the picture captured by the camera of the terminal 100 during the video recording process, etc. For another example, the first preview picture can be the above Figure 7C When the terminal 100 receives the user's input for the shooting control 322, the preview image 324 captured by the camera of the terminal 100.
[0289] In some embodiments, the first target image may also be a picture saved in the gallery application of the terminal 100, and the first reference image may be a continuous shot of the saved picture. Figure 9C The picture corresponding to the thumbnail 921 shown above Figure 9C The picture corresponding to the thumbnail 922 is shown, and so on.
[0290] In some embodiments, the first target image may also be any frame during the recording process of the terminal 100, and the first reference image may be an adjacent frame of the frame during the recording process. Figures 8B to 8C Any frame captured by the camera of the terminal 100 during the video recording process, etc.
[0291] In some embodiments, the first target image may also be any frame in a video stored on the terminal 100, and the first reference image may be an adjacent frame of the frame in the video. Figure 12C Any frame in the video 1221 shown, and so on.
[0292] For specific content, please refer to the above Figure 13 The embodiment shown.
[0293] S2002: The terminal 100 segments a first region where a first object is located in a first target image to obtain a second target image.
[0294] S2003 : The terminal 100 segments a second region where the first object is located in the first reference image to obtain a second reference image.
[0295] The first object selected as the object to be removed may be a system default on terminal 100 or a user-selected input. The first object may include one or more of the following image contents: a selfie stick, a person in the background, glasses, etc. The first object is the designated image content in the above embodiment. For details, please refer to the above embodiment and will not be repeated here.
[0296] The terminal 100 performs feature matching with the first target image based on the pre-acquired feature information of the first object (e.g., a selfie stick), determines the area of the first object in the first target image from the first target image, and segments the area of the first object from the first target image to obtain a second target image. The terminal 100 can perform feature matching with the first reference image based on the pre-acquired feature information of the first object, determines the area of the first object in the first reference image from the first reference image, and segments the area of the first object from the first target image to obtain a second reference image.
[0297] In one possible implementation, the terminal 100 may use a trained segmentation neural network to identify, based on the RGB information of the first target image, a first region where a first object (e.g., a selfie stick) is located in the first target image, and segment the first region from the first target image to obtain a second target image. The terminal 100 may use a trained segmentation neural network to identify, based on the RGB information of the first reference image, a second region where the first object (e.g., a selfie stick) is located in the first reference image, and segment the second region from the first reference image to obtain a second reference image.
[0298] In one possible implementation, the terminal 100 may use a trained segmentation neural network to identify a first region where a first object (e.g., a selfie stick) is located in the first target image based on the RGB information, depth of field information, and confidence information of the first target image, and segment the region where the first object is located from the first target image to obtain a second target image. The terminal 100 may use a trained segmentation neural network to identify a second region where the first object (e.g., a selfie stick) is located in the first reference image based on the RGB information, depth of field information, and confidence information of the first reference image, and segment the second region from the first reference image to obtain a second reference image.
[0299] In one possible implementation, the terminal 100 can use a trained segmentation neural network to identify a first region where a first object (e.g., a selfie stick) is located in the first target image based on the RGB information and thermal imaging information of the first target image, and segment the region where the first object is located from the first target image to obtain a second target image. The terminal 100 can use a trained segmentation neural network to identify a second region where the first object (e.g., a selfie stick) is located in the first reference image based on the RGB information and thermal imaging information of the first reference image, and segment the second region from the first reference image to obtain a second reference image.
[0300] For specific content, please refer to the above Figure 13 The embodiment shown.
[0301] S2004: The terminal 100 finds content similar to features around the first area in the second target image from the second reference image, fills the first area in the second target image, and obtains a third target image.
[0302] In a possible implementation, the terminal 100 may perform coarse optical flow restoration on the second target image based on the second target image and the second reference image.
[0303] Specifically, the terminal 100 can calculate the missing optical flow information in the second target image and the second reference image through the optical flow network. Then, the terminal 100 can use the optical flow completion model to complete the missing optical flow information in the second target image based on the second reference image, and obtain complete optical flow information in the second target image and the second reference image. Then, the terminal 100 can use the complete optical flow information to determine the fill pixel information in the second reference image that needs to be filled in the first area of the second target image, and use the fill pixel information to fill the pixels in the first area of the second target image to obtain the third target image.
[0304] In a possible implementation, the terminal 100 may perform multi-frame feature coarse restoration on the second target image based on the second target image and the second reference image.
[0305] Specifically, the terminal 100 can encode the second target image into a first target feature map and encode the second reference image into a first reference feature map. Then, the terminal 100 can find feature information similar to the features around the first area in the first target feature map from the first reference feature map based on the first target feature map and the first reference feature map. The feature information includes texture, color, shape, etc. Then, the terminal 100 can fill the feature information in the first reference feature map that is similar to the features around the first area in the first target feature map to the first area of the first target feature map to obtain the second target feature map. Then, the terminal 100 can decode the second target feature map into a third target image.
[0306] In a possible implementation, the terminal 100 may perform a single-frame feature coarse restoration on the second target image based on the second target image.
[0307] Specifically, terminal 100 may encode the second target image into a first target feature map. Terminal 100 may then find feature information similar to features surrounding the first region from the first target feature map. Terminal 100 may then fill the first region of the first target feature map with feature information similar to features surrounding the first region from the first target feature map to obtain a second target feature map. Terminal 100 may then decode the second target feature map into a third target image.
[0308] For specific content, please refer to the above Figure 13 、 Figure 15 、 Figure 16 、 Figure 18 The embodiment shown.
[0309] S2005. The terminal 100 generates a mask image of the second target image according to the second target image.
[0310] For specific content, please refer to the above Figure 13 The embodiments shown will not be described in detail here.
[0311] S2006: The terminal 100 optimizes and generates the texture in the first area of the third target image according to the mask image of the second target image and the third target image to obtain a fourth target image.
[0312] After the terminal 100 obtains the fourth target image, the terminal 100 may use the fourth target image as the first repaired image and display the first repaired image. For example, when the first target image is the first preview image captured by the camera, the terminal 100 may display the fourth target image as the preview image after removing the first object in the camera application interface. For another example, when the first target image is a saved picture, the terminal 100 may display the fourth target image in the picture preview interface, and so on.
[0313] In some embodiments, the terminal 100 may not execute the above step S2006, and directly use the above third target image as the first restoration picture and display the first restoration picture.
[0314] For the specific content, please refer to the above embodiments and will not be described again here.
[0315] In some embodiments, the terminal 100 can remove the first object in the continuous frame images. For example, after opening the camera application, the terminal 100 can remove the first object (such as a selfie stick) in each frame image captured by the camera. The terminal 100 can remove the first object (such as a selfie stick) in the first two frames through the above Figure 13 The image content removal process in the embodiment shown in the figure removes the first object in the first two frames. When the terminal 100 removes the first object in the third frame and the subsequent frames, the terminal 100 can combine the movement speed of the terminal 100 and the rotation angle of the terminal 100, as well as the position of the first object in the first frame, to infer the position of the first object in the third frame or the subsequent frame. Then, the terminal 100 determines the filling content of the position of the first object in the third frame or the subsequent frame from the first frame based on the position of the first object in the third frame or the subsequent frame. Then, the terminal 100 can replace the first object in the third frame or the subsequent frame with the determined filling content and fill it into the position of the first object in the third frame or the subsequent frame. In this way, when removing the first object in consecutive frames, the processing time can be reduced.
[0316] In one possible implementation, when the terminal 100 determines that the position of the first object in the third frame or subsequent frame has not changed based on the movement speed of the terminal 100 and the rotation angle of the terminal 100, the terminal 100 can directly replace the first object in the third frame or subsequent frame with the filling content in the first frame, and fill it into the position of the first object in the third frame or subsequent frame.
[0317] In some embodiments, the terminal 100 can remove the first object in consecutive frames. For example, after opening the camera application, the terminal 100 can remove the first object (such as a selfie stick) in each frame captured by the camera. For another example, the terminal 100 can remove the first object in each frame of a saved video. Among them, the terminal 100 can skip frames to remove the first object in the frame, and then copy the frame after removing the first object and insert it between two frames from which the first object has been removed. In this way, the processing time can be reduced when removing the first object in consecutive frames.
[0318] For example, a video lasting 1 second may include 60 frames. All 60 frames may include the selfie stick. When removing the selfie stick from these 60 frames, the terminal 100 may skip frames and remove the selfie stick from the 1st, 11th, 21st, 31st, 41st, and 51st frames. The terminal 100 may then copy the 1st frame after the selfie stick is removed into 10 frames, which serve as frames 1-10 of the video after the selfie stick is removed. The terminal 100 may also copy the 11th frame after the selfie stick is removed into 10 frames, which serve as frames 11-20 of the video after the selfie stick is removed. The terminal 100 may also copy the 21st frame after the selfie stick is removed into 10 frames, which serve as frames 21-30 of the video after the selfie stick is removed. Terminal 100 can copy the 31st frame after the selfie stick is removed into 10 frames, which serve as frames 31-40 of the video after the selfie stick is removed. Terminal 100 can copy the 41st frame after the selfie stick is removed into 10 frames, which serve as frames 41-50 of the video after the selfie stick is removed. Terminal 100 can copy the 51st frame after the selfie stick is removed into 10 frames, which serve as frames 51-60 of the video after the selfie stick is removed.
[0319] An image content removal method provided in an embodiment of the present application can be used to remove unwanted image content (such as a selfie stick) in pictures or videos taken by users on terminals without special cameras, thereby improving the display effect of the image content that the user wants in pictures or videos and enhancing the user experience.
[0320] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for removing image content, characterized in that: include: The terminal starts the camera application; The terminal displays a photo preview interface of the camera application; The terminal displays a third preview screen on the photo preview interface of the camera application; When the terminal recognizes that the third preview screen includes the first object, the terminal displays a movement operation prompt, wherein the movement operation prompt is used to prompt the user to move the terminal in a specified direction; The terminal obtains a first preview image and a first reference frame image captured by a camera, wherein the first preview image and the first reference frame image both include image content of the first object and image content of the second object, and in the first preview image, the image content of the first object obscures a portion of the image of the second object; determining, by the terminal, that the first object in the first preview image is an object to be removed; The terminal determines, based on the first reference frame, content to be filled in the first preview image, wherein the content to be filled is image content of the second object in the first preview image that is blocked by the first object; The terminal generates a first repairing picture according to the content to be filled and the first preview picture, wherein the image content of the first object in the first repairing picture is replaced by the image content of the obscured second object; The terminal displays the first repairing screen on the photo preview interface.
2. The method according to claim 1, characterized in that After the terminal displays the first repairing picture on the photo preview interface, the method further includes: The terminal displays a removal close control on the photo preview interface; The terminal receives a first input from a user for removing the closing control; In response to the first input, the terminal obtains a second preview image captured by a camera; The terminal displays the second preview screen on the photo preview interface.
3. The method according to claim 1, characterized in that The method further comprises: After the terminal recognizes that the third preview image includes the object to be removed, displaying a removal confirmation control; The terminal receives a second input from the user regarding the removal confirmation control; The terminal obtains a first preview image and a first reference frame image captured by a camera, specifically including: In response to the second input, the terminal obtains the first preview image and the first reference frame image captured by the camera.
4. The method according to claim 3, characterized in that The method further comprises: In response to the second input, the terminal displays a countdown of a specified duration on the photo preview interface.
5. The method according to claim 1, wherein The method further comprises: The terminal receives a click operation on the third preview screen by the user; The terminal determining that the first object in the first preview screen is an object to be removed specifically includes: In response to the click operation, the terminal identifies a click position of the click operation in the third preview picture; The terminal determines, based on image content at the clicked position in the third preview image, that the first object is the object to be removed.
6. The method according to claim 1, characterized in that The method further comprises: The terminal identifies image contents of one or more removable objects in the third preview screen and displays removal controls corresponding to the removable objects; The terminal receives a fourth input from the user for a first removal control among the one or more removal controls; The terminal determining that the first object in the first preview screen is an object to be removed specifically includes: In response to the fourth input, the terminal determines the first object corresponding to the first removal control as the object to be removed.
7. The method according to claim 1, characterized in that The method further comprises: The terminal displays a first shooting mode control on the photo preview interface; The terminal receives a fifth input from the user for the first shooting mode control; The terminal obtains a first preview image and a first reference frame image captured by a camera, specifically including: In response to the fifth input, the terminal obtains the first preview image and the first reference frame image captured by the camera.
8. The method according to claim 1, characterized in that Before the terminal obtains the first preview image and the first reference frame image captured by the camera, the method further includes: When the terminal determines that the captured image of the terminal moves significantly, the terminal displays a picture shaking prompt, where the picture shaking prompt is used to prompt a user that the captured image of the terminal moves significantly.
9. The method according to claim 8, characterized in that The terminal determines that a captured image of the terminal has a large movement, specifically including: The terminal obtains angular velocity data and acceleration data of the terminal through an inertial measurement unit; When the angular velocity in any direction of the angular velocity data is greater than a specified angular velocity value, or the acceleration in any direction of the acceleration data is greater than a specified acceleration value, the terminal determines that the captured image of the terminal has moved significantly.
10. The method according to claim 1, characterized in that The terminal determines that a motion amplitude between the first preview picture and the first reference frame picture exceeds a specified threshold, specifically including: generating, by the terminal, a first mask image after segmenting the first object in the first preview picture; The terminal generates a second mask image after segmenting the first object in the first reference frame; The terminal calculates an intersection-and-union (IoU) ratio between the first mask image and the second mask image. When the IoU ratio between the first mask image and the second mask image is less than a specified IoU ratio value, the terminal determines that a picture motion amplitude between the first preview picture and the first reference frame picture exceeds the specified threshold.
11. The method according to claim 1, wherein The terminal determines that a motion amplitude between the first preview picture and the first reference frame picture exceeds a specified threshold, specifically including: The terminal recognizes the first object in the first preview picture and segments the first object in the first preview picture; The terminal identifies the first object in the first reference frame, and segments the first object in the first reference frame to obtain the second reference frame; The terminal encodes the first preview image after segmenting the first object into a first target feature map; The terminal encodes the second reference frame picture into a first reference feature map; The terminal calculates the similarity between the first target feature map and the first reference feature map. When the similarity between the first target feature map and the first reference feature map is less than a specified similarity value, the terminal determines that the amplitude of the picture motion between the first preview picture and the first reference frame picture exceeds the specified threshold.
12. The method according to claim 1, characterized in that The method further comprises: The terminal receives a fifth input from the user; In response to the fifth input, the terminal locally saves the first repairing screen.
13. The method according to claim 1, wherein The terminal determines, based on the first reference frame, content to be filled in the first preview picture, specifically including: The terminal recognizes the first object in the first preview picture and segments the first object in the first preview picture; The terminal identifies the first object in the first reference frame, and segments the first object in the first reference frame to obtain the second reference frame; Calculating, by the terminal, missing optical flow information between the first preview picture and the second reference frame picture after segmenting the first object; The terminal completes the missing optical flow information by using an optical flow completion model according to the second reference frame, to obtain complete optical flow information between the first preview frame after segmenting the first object and the second reference frame; The terminal determines the content to be filled in the first preview picture from the second reference frame picture by using the complete optical flow information.
14. The method according to claim 1, wherein The terminal determines, based on the first reference frame, content to be filled in the first preview picture, specifically including: The terminal recognizes the first object in the first preview picture and segments the first object in the first preview picture; The terminal identifies the first object in the first reference frame, and segments the first object in the first reference frame to obtain the second reference frame; The terminal encodes the first preview image after segmenting the first object into a first target feature map; The terminal encodes the second reference frame picture into a first reference feature map; The terminal determines, from the first reference feature map, features to be filled that are similar to features surrounding the first area in the first target feature map; The terminal generates a first repairing picture according to the content to be filled and the first preview picture, specifically including: The terminal replaces the feature to be filled into the area where the first object is located in the first target feature map to obtain a second target feature map; The terminal decodes the second target feature map to obtain the first repaired image.
15. The method according to claim 1, wherein The terminal generates a first repairing picture according to the content to be filled and the first preview picture, specifically including: The terminal fills the area where the first object is located in the first preview image with the content to be filled, to obtain a roughly repaired image; The terminal generates a detail texture of the filled area in the coarse repaired image to obtain the first repaired image.
16. The method according to claim 1, wherein After the terminal determines the content to be filled in the first preview picture according to the first reference frame, the method further includes: The terminal obtains a fourth preview image captured by the camera; The terminal obtains, by the terminal, a movement angle and a rotation angle of the terminal between the first preview picture and the fourth preview picture being captured by the camera; determining, by the terminal, a region where the first object is located in the fourth preview screen according to a motion angle and a rotation angle of the terminal and a region where the first object is located in the first preview screen; Splitting, by the terminal, the first object in the fourth preview screen; The terminal determines, from the first preview picture, a fill-in content for the fourth preview picture according to an area where the first object is located in the fourth preview picture; The terminal fills the area of the fourth preview image where the first object is located with the filling content of the fourth preview image to obtain a second repaired image; The terminal displays the second repairing screen on the photo preview interface.
17. The method according to claim 1, wherein The first object includes a selfie stick or a background person.
18. A terminal, characterized in that: include: A camera, one or more processors and one or more memories; the one or more processors are coupled to the camera and the one or more memories, the one or more memories are used to store computer program code, the computer program code includes computer instructions, and when the one or more processors execute the computer instructions, the terminal executes the method according to any one of claims 1 to 17.
19. A computer-readable storage medium comprising instructions, characterized in that: When the instruction is executed on a terminal, the terminal is caused to execute the method according to any one of claims 1 to 17.
Citation Information
Patent Citations
Image processing method and device, memory medium and mobile terminal
CN108566516A