Image processing method and electronic device

By utilizing similar image combination technology in electronic devices, based on the similarity of multiple images and shooting conditions, it is possible to efficiently and realistically eliminate moving objects in photos in the default shooting mode, thereby improving the user experience.

CN122134549APending Publication Date: 2026-06-02HUAWEI TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2020-12-15
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies are inconvenient for users to operate when removing moving objects from photos, and the resulting images have low realism, especially for single photos taken in the default shooting mode.

Method used

By determining the similarity, shooting time, and location of multiple images, and selecting appropriate images for combination, the moving object is eliminated by covering or replacing it with real image content.

Benefits of technology

It improves the ease of user operation and the realism of the removed image, thus enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122134549A_ABST
    Figure CN122134549A_ABST
Patent Text Reader

Abstract

This application provides an image processing method applied to an electronic device. The method includes: determining multiple images that meet a first condition, the first condition including: the similarity between any two images is greater than or equal to a first threshold, and the multiple images including at least two images; determining a first image from the multiple images that meets a second condition, the first image including a first object, the first object being an object to be eliminated in the first image; determining a second image from the multiple images, the second image including a second object, the position of the second object in the second image corresponding to the position of the first object in the first image; and using the second object to cover or replace the first object to obtain a target image. This application embodiment can eliminate moving objects in images in general scenarios, is convenient for users, and yields a highly realistic target image.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application. The original application has the application number 2020114831241 and the original application date is December 15, 2020. The entire contents of the original application are incorporated herein by reference. Technical Field

[0002] This application relates to the field of image processing technology, and in particular to an image processing method and electronic device. Background Technology

[0003] When users take photos with electronic devices such as mobile phones and tablets, they often capture moving objects such as pedestrians and vehicles passing through the shooting area in the photos. However, in most cases, users do not want to keep these moving objects in the photos. To meet users' needs, electronic devices can provide the function of removing moving objects from photos.

[0004] Currently, users can activate a specific shooting mode on their electronic device and capture multiple frames (e.g., a 1.7-second video) within that mode. The device can then use these frames to remove passersby and fill in the area after the passersby are removed. However, this method is inconvenient, as the device cannot process single photos taken using the default shooting mode. Alternatively, after taking a single photo using the default mode, users can select a specific person in the photo to remove. In this case, the image content used to fill the area after the person is removed is guessed by the device using artificial intelligence (AI) and other technologies. This guessed image content can easily differ significantly from the surrounding environment, resulting in low realism. Summary of the Invention

[0005] This application discloses an image processing method and electronic device that can remove moving objects from images in general scenarios. It is convenient for users to use, and the images after removing moving objects have a high degree of realism.

[0006] In a first aspect, embodiments of this application provide an image processing method applied to an electronic device. The method includes: determining a plurality of images that meet a first condition, the first condition including: the similarity between any two images in the plurality of images is greater than or equal to a first threshold, the plurality of images including at least two images; determining a first image from the plurality of images that meets a second condition, the first image including a first object, the first object being an object to be eliminated in the first image; determining a second image from the plurality of images, the second image including a second object, the position of the second object in the second image corresponding to the position of the first object in the first image; and using the second object to cover or replace the first object to obtain a target image.

[0007] In this embodiment, the electronic device can eliminate a first object based on multiple images, where any two images have a high degree of similarity. Any one of the images can be captured using the electronic device's default shooting mode, without requiring a specific shooting mode, making it more convenient for users and applicable to a wide range of scenarios. Furthermore, the second object used to cover or replace the first object is obtained based on a real second image, resulting in higher consistency between the target image and the real world, better display effects, and thus improved user experience.

[0008] In one possible implementation, the first condition further includes at least one of the following: the shooting time of any one of the multiple images is within a first range, and the shooting location of any one of the multiple images is within a second range.

[0009] In this embodiment, the electronic device can determine multiple images for first object elimination not only based on image similarity but also based on the image's shooting time and / or location, further ensuring that the multiple images belong to the same shooting scene. Eliminating the first object based on such multiple images improves the realism of the target image and enhances the user experience.

[0010] In one possible implementation, before determining the plurality of images that meet the first condition, the method further includes: receiving a first operation, the first operation being used to select the plurality of images.

[0011] In this embodiment, the user can choose multiple images to achieve the first object elimination, which is more flexible and the resulting target image is more in line with the user's needs, resulting in a better user experience.

[0012] In one possible implementation, the second condition includes at least one of the following: the sharpness of the subject in the first image is greater than or equal to a second threshold, the number of other objects in the first image besides the subject is less than a third threshold, and a second operation is received, the second operation being used to select the first image.

[0013] In this embodiment, the first image used to obtain the target image must meet a second condition, such as high clarity of the subject being photographed and a small number of other objects besides the subject being photographed. Therefore, the display effect of the subject being photographed in the obtained target image is also better, resulting in a better user experience.

[0014] In one possible implementation, the second image is one of the plurality of images other than the first image, and the similarity between the second image and the first image is greater than or equal to a fourth threshold.

[0015] In this embodiment, the second image can be one of multiple images that has a high degree of similarity to the first image. When an electronic device uses a second object in such a second image to cover or replace a first object in the first image, the resulting target image will have a better display effect and a better user experience.

[0016] In one possible implementation, the method further includes: determining the subject being photographed from the plurality of images that satisfies a third condition; the third condition includes at least one of the following: the sharpness of the subject being photographed in any one of the plurality of images is greater than or equal to a fifth threshold; the focus point of any one of the plurality of images is located in the area where the subject being photographed is located; the area of ​​the subject being photographed in any one of the plurality of images is greater than or equal to a sixth threshold; the subject being photographed belongs to a preset category; and a third operation is received, the third operation being used to select the subject being photographed.

[0017] In this application embodiment, the electronic device can determine the subject of the photograph in various ways. The electronic device can flexibly choose the method to determine the subject based on its own capabilities, and the application scenarios are wide-ranging. For example, when the processing power is strong, the electronic device can combine multiple methods to determine the subject of the photograph, so that the determined subject of the photograph better meets the user's needs.

[0018] In one possible implementation, the method further includes receiving a fourth operation for selecting the first object.

[0019] In this embodiment, the user can select the object to be eliminated, which is more flexible and the resulting target image is more in line with the user's needs, resulting in a better user experience.

[0020] In one possible implementation, the method further includes: determining the center point of a third object in the first image, the center point of a fourth object in the first image, the center point of a fifth object in the third image, and the center point of a sixth object in the third image; wherein the third image is any image other than the first image among the plurality of images, the third object and the fifth object have the same attributes, and the fourth object and the sixth object have the same attributes; setting the center point of the third object and the center point of the fifth object to the same coordinate origin, and establishing a first coordinate system based on the coordinate origin; determining a first distance between the center point of the fourth object and the center point of the sixth object based on the first coordinate system; and determining the fourth object as the first object when the first distance is greater than or equal to a seventh threshold.

[0021] In this embodiment of the application, before the electronic device determines the first object, it can first set the two images to be determined to the same first coordinate system, thereby eliminating the influence of factors such as displacement and rotation, so that the display effect of the target image is better and the user experience is better.

[0022] In one possible implementation, the objects represented by the third object and the fifth object are in the same position at any point in time.

[0023] In this embodiment, the third and fifth objects used to determine the first coordinate system can be the same fixed object (e.g., a building, a tree, or a flower), so that the obtained first coordinate system and the world coordinate system are as consistent as possible, the obtained target image has higher consistency with the real world, and the display effect is better.

[0024] In one possible implementation, the method further includes: receiving a fifth operation; and in response to the fifth operation, displaying a first interface, wherein the first interface displays the plurality of images and the target image.

[0025] In this embodiment, the electronic device can eliminate the first object without the user noticing and recommend the resulting target image to the user for viewing, without requiring the user to manually trigger the elimination function, making it more convenient to use.

[0026] Secondly, embodiments of this application provide an electronic device, which includes at least one memory and at least one processor. The at least one memory is coupled to the at least one processor. The at least one memory is used to store a computer program, and the at least one processor is used to call the computer program. The computer program includes instructions. When the instructions are executed by the at least one processor, the electronic device performs the image processing method provided by the first aspect or any implementation of the first aspect in the embodiments of this application.

[0027] Thirdly, embodiments of this application provide a computer storage medium including computer instructions, which, when executed on an electronic device, cause the electronic device to perform the image processing method provided by the first aspect or any implementation thereof in the embodiments of this application.

[0028] Fourthly, embodiments of this application provide a computer program product that, when run on an electronic device, causes the electronic device to execute the image processing method provided by the first aspect or any implementation thereof in the embodiments of this application.

[0029] Fifthly, embodiments of this application provide a chip, which includes at least one processor, an interface circuit, and a memory. The memory, the interface circuit, and the at least one processor are interconnected via circuits. The memory stores a computer program, and when the computer program is executed by the at least one processor, it implements the image processing method provided by the first aspect or any implementation of the first aspect in the embodiments of this application.

[0030] Understandably, the electronic device provided in the second aspect, the computer storage medium provided in the third aspect, the computer program product provided in the fourth aspect, and the chip provided in the fifth aspect are all used to execute the image processing method provided in the first aspect or any implementation thereof. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the image processing method provided in the first aspect, and will not be repeated here. Attached Figure Description

[0031] The accompanying drawings used in the embodiments of this application are described below.

[0032] Figure 1 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application; Figure 2 This is a schematic diagram of the software architecture of another electronic device provided in the embodiments of this application; Figure 3 This is a schematic diagram of a user interface embodiment provided in this application; Figure 4 This is a schematic diagram of a grouping method provided in an embodiment of this application; Figures 5-6 These are image groups obtained through some grouping processes provided in the embodiments of this application; Figure 7 This is a schematic flowchart of an image processing method provided in an embodiment of this application; Figures 8-9 , Figures 10A-10D , Figure 11These are schematic diagrams illustrating some image processing procedures provided in the embodiments of this application; Figure 12 , Figures 13A-13B , Figures 14-18 These are schematic diagrams of some other user interface embodiments provided in this application; Figure 19 This is a flowchart illustrating another image processing method provided in the embodiments of this application. Detailed Implementation

[0033] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to and includes any or all possible combinations of one or more of the listed items.

[0034] This application provides an image processing method applicable to electronic devices. The electronic device can group multiple images based on the similarity between any two images to obtain at least one image group to be processed. Any image group to be processed may include a first image and at least one second image. The electronic device can eliminate a first object in the first image based on an image group to be processed; that is, it first obtains a second object based on at least one second image, and then uses the second object to cover or replace the first object in the first image. The multiple images can be captured using the electronic device's default shooting mode, without requiring a specific shooting mode, thus broadening the application scenarios and making it more convenient for users. Furthermore, the second object is a reliable image content obtained from at least one second image, resulting in better display effects and improved user experience.

[0035] Understandably, when users take pictures using electronic devices, they usually want to retain one or more real objects in the image; these real objects can be called the subject. Any image in the target image set can include the subject, such as people, mountains, trees, water, sky, animals, etc. However, when users take pictures using electronic devices, other real objects besides the subject (i.e., the aforementioned first object) often pass through the shooting area, resulting in the presence of the aforementioned first object in the captured image. For example, a pedestrian may appear in the first image, and pedestrians and passing vehicles may appear in multiple second images. Users usually do not want to retain the aforementioned first object in the captured image; therefore, the first object can also be understood as an object to be removed. Since the first object passes through the shooting area, it can also be understood as a moving object relative to the subject (referred to as a moving object).

[0036] The electronic devices involved in the embodiments of this application may be smart screens, smart TVs, mobile phones, tablets, desktops, laptops, ultra-mobile personal computers (UMPCs), handheld computers, netbooks, personal digital assistants (PDAs), wearable electronic devices (such as smart bracelets, smart glasses), etc.

[0037] The following describes an exemplary electronic device provided in the embodiments of this application.

[0038] Please see Figure 1 , Figure 1 A schematic diagram of the structure of an electronic device 100 is shown.

[0039] Electronic device 100 may include processor 110, external memory interface 120, internal memory 121, universal serial bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, button 190, motor 191, indicator 192, camera 193, display screen 194, and subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0040] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0041] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors.

[0042] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0043] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0044] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0045] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.

[0046] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.

[0047] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through an audio device (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.

[0048] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0049] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling electronic device 100 to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).

[0050] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0051] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1.

[0052] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.

[0053] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and color. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0054] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0055] In some embodiments, the electronic device 100 may be configured with a plurality of cameras 193, which may include front-facing cameras and rear-facing cameras. Optionally, there may be multiple front-facing cameras, which may be located, for example, at the top of the front of the electronic device 100. Optionally, there may also be multiple rear-facing cameras, such as a rear-facing wide-angle camera, a rear-facing ultra-wide-angle camera, or a rear-facing telephoto camera. The rear-facing cameras may be located, for example, on the back of the electronic device 100. In some embodiments of this application, the plurality of cameras 193 may also be pop-up cameras, detachable cameras, etc. The embodiments of this application do not limit the connection method and mechanical mechanism of the plurality of cameras 193 and the electronic device 100.

[0056] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.

[0057] Internal memory 121 can be used to store computer executable program code, which includes instructions. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of electronic device 100 by running instructions stored in internal memory 121 and / or instructions stored in memory located in the processor.

[0058] In this embodiment, the electronic device 100 can acquire and save multiple images via the camera 193. These images can be stored in the internal memory 121 or in an external memory card connected to the external memory interface 120. Then, the processor 110 of the electronic device 100 can group these images according to their similarity to obtain at least one group of images to be processed. For example, but not limited to, the similarity between two images can be characterized by the difference in shooting time, the difference in distance between shooting locations, the distance between image feature vectors, angles, etc. Based on the obtained group of images to be processed, the processor 110 can perform the identification and removal of moving objects, for example, by covering the area where the moving object is located with a real image, which can be obtained from at least one image in the group of images to be processed.

[0059] Pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A can be disposed on display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive material. When force is applied to pressure sensor 180A, the capacitance between the electrodes changes. Electronic device 100 determines the pressure intensity based on the change in capacitance. When a touch operation is applied to display screen 194, electronic device 100 detects the intensity of the touch operation based on pressure sensor 180A. Electronic device 100 can also calculate the touch position based on the detection signal from pressure sensor 180A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation commands. For example, when a touch operation with an intensity less than a first pressure threshold is applied to the SMS application icon, a command to view an SMS is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the SMS application icon, a command to create a new SMS is executed.

[0060] Touch sensor 180K, also known as a "touch device," can be located on display screen 194. The touch sensor 180K and display screen 194 together form a touchscreen, also known as a "touchscreen." Touch sensor 180K detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K may also be located on the surface of electronic device 100, in a different position than display screen 194.

[0061] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Electronic device 100 can receive button input and generate key signal inputs related to user settings and function control of electronic device 100.

[0062] The software system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses the layered architecture Android system as an example to exemplify the software structure of electronic device 100.

[0063] Figure 2 This is a software structure block diagram of the electronic device 100 according to an embodiment of the present invention.

[0064] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android® system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer. In this application, Figure 2 The software framework shown is just an example. The system of electronic device 100 can also be other operating systems, such as iOS®, Windows®, Huawei Mobile Services (HMS), etc.

[0065] The application layer can include a series of applications.

[0066] like Figure 2 As shown, the applications can include camera, gallery, map, music, SMS, calendar, call, navigation, Bluetooth, file management, and other applications.

[0067] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0068] like Figure 2 As shown, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.

[0069] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.

[0070] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.

[0071] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.

[0072] The phone manager is used to provide communication functions for electronic device 100. For example, it manages call status (including connection and disconnection).

[0073] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.

[0074] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.

[0075] The Android Runtime consists of core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.

[0076] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.

[0077] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0078] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.

[0079] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.

[0080] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.

[0081] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0082] A 2D graphics engine is a graphics engine for 2D drawing.

[0083] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.

[0084] Understandably, applications such as camera or gallery apps may include a moving object removal function, which users can use to obtain an image after the moving object has been removed. However, this is not a limitation; the moving object removal function can also be an application installed on the electronic device 100, or an online application, such as a web application or a mini-program application. This application embodiment does not impose such limitations.

[0085] For ease of description, the following examples use the feature of moving object elimination, which is included in a gallery application, as an example.

[0086] The following example, using a photography scenario, illustrates the workflow of the software and hardware of electronic device 100.

[0087] When the pressure sensor 180A and / or the touch sensor 180K receive a touch operation, a corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including touch coordinates, touch operation timestamp, etc.). The raw input event is stored in the kernel layer. The application framework layer retrieves the raw input event from the kernel layer and identifies the control corresponding to the input event. Taking a single touch operation as an example, where the corresponding control is the camera application's photo capture control, the camera application calls the application framework layer's interface, which in turn calls the kernel layer to start the camera driver, capturing one or more photos through the camera 193. These one or more photos can be saved as multiple images (also known as pictures) in the gallery (also known as an album).

[0088] Please see Figure 3 , Figure 3 An exemplary user interface 30 for a camera application on an electronic device such as a smartphone (the electronic device here may correspond to the aforementioned electronic device 100) is shown. The electronic device can detect touch operations (e.g., click operations) performed on the icon of the camera application, which can be displayed on the desktop of the electronic device, which may include icons for multiple applications. In response to the touch operation, the electronic device can display... Figure 3 The user interface 30 shown is a default shooting mode interface for a camera application, which can be used by the user to take photos using the default rear camera of the electronic device. That is, the user can tap the camera application icon to open the camera application's user interface 30. However, users can also open the user interface 30 in other applications, such as tapping the shooting control in a social media application.

[0089] like Figure 3As shown, the user interface 30 may include: area 301, shooting function list 302, shooting mode list 303, controls 304, controls 305, and controls 306. Among them: Area 301 can be referred to as preview frame 301 or viewfinder 301. Preview frame 301 can be used to display images captured in real time by camera 193. The electronic device can refresh the displayed content in real time so that the user can preview the image currently captured by camera 193.

[0090] The shooting function list 302 may display at least one shooting function option: Smart Vision option 302A, Flash option 302B, Live Photo option 302C, Color Mode option 302D, and Camera Settings option 302E.

[0091] For example, the electronic device can detect user actions (such as clicks) on the Live Photo option 302C to enable or disable the Live Photo shooting function. When the Live Photo shooting function is enabled, the electronic device can detect the user action that triggers the photo taking. In response to the user action, the electronic device can capture multiple frames of images and encode these frames into a video, which is the Live Photo. For example, the electronic device can capture 40 frames of images and encode these 40 frames into a video (i.e., a Live Photo) at 24 frames per second (fps), with a duration of 1.7 seconds.

[0092] The shooting mode list 303 can display at least one shooting mode option: aperture mode option 303A, night scene mode option 303B, portrait mode option 303C, photo mode option 303D, video mode option 303E, professional mode option 303F, and more mode option 303G. Figure 3 In the image, the shooting mode option 303D is selected, indicating that the electronic device is currently in shooting mode, which can be the default shooting mode of the electronic device. The electronic device can detect user actions (such as clicks) applied to other shooting mode options in the shooting mode list 303, and in response to the user action, the electronic device can switch shooting modes.

[0093] Control 304 can be used to listen for user actions that trigger shooting (taking a photo or recording a video). The electronic device can detect user actions (such as clicks) applied to control 304, and in response to this action, the electronic device can save the image in preview box 301 as a picture or video in the gallery application. In other words, the user can click control 304 to trigger shooting. The gallery application can support various operations performed by the user on the pictures or videos stored on the electronic device, such as browsing, editing, deleting, and selecting. Furthermore, the electronic device can also display thumbnails of the saved images in control 305.

[0094] Control 306 can be used to listen for user actions that trigger the camera to flip. The electronic device can detect user actions (such as clicks) applied to control 306, and in response to such actions, the electronic device can switch the camera used to acquire images, for example, switching the camera used to acquire images from the rear camera to the front camera.

[0095] In this application, the electronic device can capture and store multiple images (also referred to as pictures) using a camera 193. The electronic device can group these multiple images according to the similarity between any two images to obtain at least one image group to be processed (referred to as a first group). This at least one first group is used by the electronic device to eliminate moving objects. Here, the electronic device can correspond to the aforementioned electronic device 100.

[0096] When grouping images, electronic devices can first extract image features for each image. Traditional computer vision (CV) algorithms can be used for feature extraction, such as scale-invariant feature transform (SIFT) and accelerated uprobust features (SURF) for corner detection and feature representation. Alternatively, deep learning (DL) algorithms, such as convolutional neural networks (CNNs), can be employed for feature extraction. After extracting the feature vectors of the images, the electronic device can calculate parameters such as the distance or angle between the feature vectors of any two images and use these parameters to determine the similarity between the two images. When the similarity is greater than or equal to a first threshold, the electronic device determines that the two images meet the similarity condition.

[0097] For example, the electronic device can calculate the Euclidean distance between the feature vectors of any two images, where a smaller Euclidean distance indicates a greater similarity between the two images, and a larger Euclidean distance indicates a less similarity between the two images. Alternatively, the electronic device can also calculate the cosine distance between the feature vectors of any two images, where a smaller cosine distance indicates a greater similarity between the two images, and a larger cosine distance indicates a less similarity between the two images.

[0098] In some embodiments, the electronic device may process multiple images stored in the device based on their shooting time before calculating the similarity, to obtain a second group that meets a preset time condition. The electronic device then determines the similarity based on the second group. The preset time condition includes: the shooting time of any image in the second group is within a first range, which can also be understood as the difference in shooting time between any two images in the second group being less than or equal to a first time threshold. For example, a photo library application includes 10 ungrouped images. These 10 images were taken on the same day at the following times: 9:10 AM, 9:11 AM, 9:11 AM, 9:15 AM, 10:10 AM, 10:15 AM, 11:21 AM, 2:35 PM, 2:35 PM, and 2:36 PM. The electronic device groups these 10 images based on their shooting time, with a first time threshold of 5 minutes. Therefore, the four images taken at 9:10, 9:11, 9:11, and 9:15 belong to group 1; the two images taken at 10:10 and 10:15 belong to group 2; the image taken at 11:21 belongs to group 3; and the three images taken at 14:35, 14:35, and 14:36 ​​belong to group 4. In other words, the electronic device obtained four second groups: group 1, group 2, group 3, and group 4.

[0099] In some embodiments, the electronic device may also process multiple images stored in the electronic device according to the shooting location before calculating the similarity, so as to obtain a third group that meets the preset location conditions. Specific examples are as follows. Figure 4 As shown. The electronic device then determines the similarity based on the third group. The preset location conditions include: the shooting location of any image in the second group is within a second range; this can also be understood as the distance difference between the shooting locations of any two images in the second group being less than or equal to a first distance threshold. In this application, the shooting location of any image can be obtained by the electronic device through technologies such as GPS. The shooting location can be represented by longitude and latitude. For example, if the shooting location is the Wuhan Municipal Government, the longitude is 114.305215 and the latitude is 0.592935, then the shooting location can be represented as (114.305215, 30.592935).

[0100] Please see Figure 4 , Figure 4 An example is shown of a map area 40 obtained by an electronic device using technologies such as GPS. The map area 40 contains four locations: location 401, location 402, location 403, and location 404. Location 401 is the location where image A was taken, location 402 is the location where image B was taken, location 403 is the location where image C was taken, and location 404 is the location where image D was taken. Images A, B, C, and D are multiple images stored in the electronic device. It should be noted that any one of images A, B, C, and D may include one or more images.

[0101] like Figure 4 As shown, a circle centered at position 401 with a radius equal to the first distance threshold is defined as range 411, meaning the distance difference between any position within range 411 and position 401 is less than or equal to the first distance threshold. Similarly, a circle centered at position 402 with a radius equal to the first distance threshold is defined as range 412, meaning the distance difference between any position within range 412 and position 402 is less than or equal to the first distance threshold. A circle centered at position 403 with a radius equal to the first distance threshold is defined as range 413, meaning the distance difference between any position within range 413 and position 403 is less than or equal to the first distance threshold. A circle centered at position 404 with a radius equal to the first distance threshold is defined as range 414, meaning the distance difference between any position within range 414 and position 404 is less than or equal to the first distance threshold.

[0102] Therefore, from Figure 4 It can be seen that the distance difference between positions 401 and 402, and between 403 is less than the first distance threshold; the distance difference between positions 402 and 403 is less than the first distance threshold; and the distance difference between positions 403 and 404 is less than the first distance threshold. Therefore, the electronic device can divide images A, B, and C into one third group, and divide images C and D into another third group.

[0103] Not limited to the situations listed above, in specific implementations, electronic devices can also measure the similarity between two images by shooting time and / or shooting location. That is, the smaller the difference in shooting time, the greater the similarity; the greater the difference in shooting time, the less similar the image. Similarly, the smaller the difference in shooting location, the greater the similarity; the greater the difference in shooting location, the less similar the image. Alternatively, the electronic device can first process the images stored in the electronic device based on the shooting time and shooting location to obtain a fourth group that meets preset time and location conditions. Then, the electronic device performs a similarity judgment based on the fourth group to obtain at least one first group. This application does not limit the specific method of obtaining the first group.

[0104] In some embodiments, the electronic device may first perform focal length registration on the stored multiple images, that is, set the focal length of each image to the same focal length. Then, the electronic device groups the focal length-registered multiple images to obtain at least one first group, thereby reducing processing errors. Not limited to this, the electronic device may also first perform angle registration on the stored multiple images, for example, controlling the angle between any two subjects in the images to be less than or equal to 5 degrees, and then group the angle-registered multiple images.

[0105] Understandably, the electronic device can process multiple images stored in the device according to the above grouping process to obtain a first group that satisfies the first condition. The first condition includes: the similarity between any two images in the first group is greater than or equal to a first threshold, meaning that any two images in the first group satisfy the similarity condition. An example of the first group can be found here. Figure 5 Image group A and shown Figure 6 Image group B shown, in which, Figure 5 The image group A shown may include four images: image 501, image 502, image 503, and image 504. Figure 6 The image group B shown may include four images: image 601, image 602, image 603 and image 604.

[0106] Not limited to the cases listed above, in specific implementations, any image processed by the electronic device can also be a frame extracted from a saved video. For example, the electronic device can use AI technology to extract at least one frame from the video whose similarity to the first image is greater than or equal to a first threshold. This at least one frame and the first image belong to a first group.

[0107] In this application, the electronic device can eliminate moving objects in each of the first groups obtained above, and the specific process is as follows: Figure 7 As shown. Among them, Figure 7 by Figure 6 The image group B shown is the first group for explanation.

[0108] Please see Figure 7 , Figure 7 An exemplary flowchart of an image processing method is shown. This method can be applied to... Figure 1 The electronic device 100 shown. This method can also be applied to... Figure 2 The electronic device 100 shown. The method includes, but is not limited to, the following steps: S701: The electronic device performs semantic segmentation on the images in the first group.

[0109] Specifically, electronic devices can identify objects in an image through semantic segmentation, as shown in the following examples. Figure 8 As shown.

[0110] Please see Figure 8 , Figure 8 An example is shown to compare semantic segmentation before and after. Figure 8 To Figure 6 The semantic segmentation of the image 601 shown is used as an example for illustration. Figure 8 Image (A) shown is image 601 before semantic segmentation. Figure 8 Image (B) shown is image 601 after semantic segmentation. Figure 8 As shown in (B), after semantic segmentation, image 601 can include person A, person B, person C, buildings, trees, and cars. Electronic devices... Figure 6 The process of semantic segmentation of other images in image group B shown and Figure 8 Similarly, image 602 also includes a balloon, image 603 is identical to image 601, and image 604 does not include a car.

[0111] Not limited to Figure 8 In the example implementation, after semantic segmentation, the objects included in the image are such as: people, buildings, cars, green plants (including grass, trees, and flowers), food, pets, water, beaches, and mountains. This application does not limit the specific type of the objects.

[0112] S702: The electronic device identifies the subject in the first group.

[0113] Specifically, any image in the first group includes the subject being photographed. This application may refer to the object included in each image of the first group as the first test object. For each first test object, the electronic device may first acquire at least one of the following: sharpness, area occupied, whether it belongs to a first preset category, and whether the area where the focus point is located is within the area where the first test object is located. Then, based on the acquired information, the first test object satisfying the first preset condition is determined as the subject being photographed. The first preset condition may include at least one of the following: sharpness greater than or equal to a first preset threshold, area occupied greater than or equal to a second preset threshold, belonging to a first preset category, and the area where the focus point is located is within the area where the first test object is located. The sharpness can be characterized, but is not limited to, by the grayscale difference or gradient between adjacent pixels in the image. For example, the sharpness of the image can be characterized by values ​​calculated using algorithms such as the Brenner gradient function, Tenengrad gradient function, and Laplacian gradient function; the larger the value, the higher the image sharpness, and the smaller the value, the lower the image sharpness.

[0114] The first preset category can be a classification of an object obtained in advance by the electronic device based on information such as historical images, such as people, pets, buildings, and landscapes. For example, suppose the electronic device can obtain the facial feature vector of a first person from historical images and identify the first person as a child (e.g., the user directly labels the person's identity). When the facial feature vector of the first test object matches the facial feature vector of the first person (e.g., the similarity is greater than a third preset threshold), the electronic device can identify the object as a child, meaning the first test object belongs to the first preset category. The facial feature vector represents the user's facial information, and may include features such as facial structure, face size, and shape.

[0115] It should be noted that the aforementioned sharpness and area occupied can be the sharpness and area occupied by the first object under test in any image of the first group, and the aforementioned area where the focus point is located can also be the area where the focus point is located in any image of the first group. However, this application does not limit this to the aforementioned sharpness and area occupied, which can also be the sharpness and area occupied by the first object under test in a predetermined number of images of the first group, and the aforementioned area where the focus point is located can also be the area where the focus point is located in a predetermined number of images of the first group.

[0116] For example, suppose the first preset conditions include: largest area occupied and belonging to a first preset category. Here, suppose the first preset category includes people and pets. And... Figure 6 In the image group B shown, each image includes objects such as person A, person B, person C, a building, and a tree. Among them, person A occupies the largest area and belongs to the first preset category. Therefore, the electronic device can identify person A as the subject of the image group A.

[0117] In some embodiments, any one of the first test objects in the first group may have a corresponding priority, and the electronic device may determine the first test object with the highest priority as the subject of the first group. The priority of an object may be determined by at least one of the following: sharpness, area occupied, whether it belongs to a first preset category, and whether the area where the focus point is located is located within the area of ​​the first test object.

[0118] For example, the priority of the first test object is determined by its sharpness, area occupied, and whether it belongs to a first preset category. The priority of the first test object can be represented as W. Sharpness is represented as qa in the calculation of W, with a weight of wa. Area occupied is represented as qb in the calculation of W, with a weight of wb. Whether it belongs to the first preset category is represented as qc in the calculation of W, with a weight of wc. Therefore, the expression for W can be as follows:

[0119] in, .For example, It is not limited to this, it can also be This application does not limit the specific value of the weight.

[0120] The value of qc can be 0 or 1. When qc is 0, it indicates that the first test object does not belong to the first preset category; when qc is 1, it indicates that the first test object belongs to the first preset category. However, it is not limited to this; different values ​​of qc can also represent the specific category to which the first test object belongs. This application does not limit the way qa, qb, and qc are valued.

[0121] For example, the priority of the first test object is determined solely by the sharpness of the first test object. The higher the sharpness of the first test object, the higher its priority.

[0122] For example, the priority of the first object under test is determined solely by its area. The larger the area occupied, the higher the priority of the first object under test.

[0123] For example, the priority of the first test object is determined solely by whether it belongs to a first preset category; alternatively, it may also be determined by its specific category. When the first test object belongs to the first preset category, its priority can be increased. Alternatively, assuming the first preset category includes people and buildings, the priority of the first test object increases more significantly when it belongs to a person and less significantly when it belongs to a building. Or, assuming the first preset category includes people, and includes the specific identity of the people: relatives or friends. The priority of the first test object increases more significantly when it belongs to a relative and less significantly when it belongs to a friend; and the priority increase is minimal when it belongs only to a person but not to a relative or friend.

[0124] S703: The electronic device determines the first image in the first group.

[0125] Specifically, the electronic device can determine a first image from the first group that satisfies a second condition, which includes at least one of the following: the sharpness of the subject in the first image is greater than or equal to a second threshold; the number of other objects in the first image besides the subject is less than a third threshold; the subject in the first image is not occluded; and the state of the subject in the first image is a preset state. The electronic device can determine the number of other objects based on the result of semantic segmentation. However, it is not limited to this; the electronic device can also use technologies such as AI to identify objects included in the image, thereby determining the number of other objects.

[0126] The electronic device can determine whether the subject being photographed is occluded by checking whether the similarity of the subject changes within a first group. For example, based on the first group, the electronic device can obtain the similarity of the subject in any image and other images. When the similarity is less than a fourth preset threshold, the electronic device can determine that there is a significant feature change in the subject in the current image, that is, determine that the subject in the current image is occluded.

[0127] Electronic devices can use technologies such as artificial intelligence (AI) to recognize the state of the subject being photographed. For example, when the subject is a person, their state can include facial expressions such as smiling or crying, and postures such as standing, leaning, or crouching. When the subject is a pet such as a cat or dog, their state can include postures such as lying down, standing, or running. When the subject is an object such as a car or bicycle, their state can include stopping or moving.

[0128] For example, Figure 6 In image group B shown, the subject is person A. Assume the second condition includes: the number of objects other than the subject is minimized, and the subject is not obscured. Person A in image 602 is obscured, therefore image 602 is disregarded. Images 601 and 603 contain five objects besides person A: person B, person C, a car, a building, and a tree. Image 604 contains four objects besides person A: person B, person C, a building, and a tree, but does not include the car. Therefore, the electronic device can identify image 604 as the first image in image group B.

[0129] In some embodiments, any image in the first group may have a corresponding priority, and the electronic device may determine the image with the highest priority as the first image of the first group. The priority of an image may be determined by at least one of the following: the sharpness of the subject being photographed, the number of other objects besides the subject being photographed, whether the subject being photographed is obscured, and the state of the subject being photographed.

[0130] For example, the priority of an image is determined by the sharpness of the subject in the image, the number of other objects besides the subject, and whether the subject is occluded. The priority of an image can be represented as U. The sharpness of the subject is represented as qd in calculating U, with a weight of wd. Whether the subject is occluded is represented as qe in calculating U, with a weight of we. The number of other objects besides the subject is represented as qf in calculating U, with a weight of wf. Therefore, the expression for U can be as follows:

[0131] in, .For example, It is not limited to this, it can also be This application does not limit the specific value of the weight.

[0132] The value of qf can be less than or equal to 0. When qf is 0, it means that the number of objects other than the subject being photographed is 0. When qf is less than 0, the smaller the qf, the more objects other than the subject being photographed; the larger the qf, the fewer objects other than the subject being photographed. However, qf can also be greater than or equal to 0. This application does not limit the way qd, qe, and qf are valued.

[0133] For example, the priority of an image is determined solely by the sharpness of the subject in the image. The higher the sharpness of the subject in the image, the higher the priority of the image.

[0134] For example, the priority of an image is determined solely by whether the subject in the image is occluded. When the subject in the image is not occluded, the priority of the image can be increased; when the subject in the image is occluded, the priority of the image can be decreased.

[0135] For example, the priority of an image is determined solely by the state of the subject being photographed. Assuming the subject in the image is a person, the person's facial features, expression, etc., represent the state of the subject. When the person's eyes are open, the image's priority can be increased; when the person's expression is smiling, the image's priority can be increased.

[0136] For example, the priority of an image is determined solely by the number and area of ​​objects other than the subject in the image. The fewer the number of objects other than the subject in the image, the higher the priority of the image; and the smaller the area of ​​objects other than the subject in the image, the higher the priority of the image.

[0137] S704: The electronic device identifies the first object in the first group.

[0138] Specifically, this application can designate any object in the first group other than the subject being photographed as a third object. The electronic device can obtain the distance between the third object in the first image and the third objects in other images of the first group. When the distance is greater than or equal to a fifth preset threshold, the electronic device can identify the third object as the first object. The first object is the moving object to be eliminated.

[0139] In some embodiments, when an electronic device determines a first object in a first group, it first performs coordinate registration, that is, arranging the images in the first group in the same coordinate system. This coordinate system can be a two-dimensional coordinate system or a three-dimensional coordinate system; this application uses a two-dimensional coordinate system as an example. This application can refer to any two images in the first group as the first image to be tested and the second image to be tested. The coordinate registration process of the first image to be tested and the second image to be tested will be described exemplarily below.

[0140] The electronic device can first acquire at least one first key point in a first image to be tested and at least one second key point in a second image to be tested, wherein the number of first key points and second key points is the same. One first key point corresponds to one second key point, meaning the similarity between the first key point and the second key point is greater than or equal to a sixth preset threshold. For example, if the first image to be tested is image 604, and the first key point is the center point of the left eye of person A in image 604, and the second image to be tested is image 603, and the second key point is the center point of the left eye of person A in image 603, then the number of first key points and second key points can be represented as n, and the first key points included in the first image to be tested can be represented as a sequence. The second keypoints included in the second image to be tested can be represented as a sequence. Assuming the first image to be tested is in the standard coordinate system, the second image to be tested needs to be rotated and translated to be in the standard coordinate system. The rotation value can be expressed as... The translation value can be expressed as . Satisfy the following formula:

[0141] The electronic device can obtain the rotation value R and translation value N by matrix inversion, and then rotate the second image under test by R and translate it by N. At this time, the first and second images under test are in the same standard coordinate system, and the first key point and the corresponding second key point coincide.

[0142] In some embodiments, in order to make the established standard coordinate system and the world coordinate system as consistent as possible, the selected first key point and second key point can be located on the subject being photographed, or on one or more objects whose positions do not change at any point in time, such as buildings, green plants (including grass, trees, and flowers), beaches, mountains, etc.

[0143] The electronic device can acquire the center point of the third object in different images under the aforementioned standard coordinate system and calculate the distance between these center points. When any distance is greater than or equal to a fifth preset threshold, the electronic device can determine that the third object is a moving object to be eliminated (i.e., the first object). A specific example is as follows. Figure 9 As shown.

[0144] Please see Figure 9 , Figure 9 An exemplary schematic diagram is shown for determining a first object. Figure 9 The third object to be confirmed Figure 6 Let's take person B in image group A as an example for illustration.

[0145] like Figure 9 As shown, assume that the two-dimensional coordinate system established with point O as the origin is the standard coordinate system obtained by the above coordinate registration. The image with a gray background is the first image (i.e., image 604), region 900 is the region where person A (i.e., the subject) is located in image 604, and point O is the center point of person A in image 604. In the other images of image group A, the center point of person A coincides with point O, and the region where person A is located also coincides with region 900. Among them, the region where image 601 is located completely coincides with the region where image 604 is located.

[0146] like Figure 9 As shown, region 6010 is the region where person B is located in image 601, and region 6020 is the region where person B is located in image 602. The overlap between regions 6010 and 6020 is greater than a sixth preset threshold, therefore the electronic device can consider the displacement of person B in images 601 and 602 to be 0. Region 6030 is the region where person B is located in image 603, and the distance between the center point of region 6010 or region 6020 and the center point of region 6030 is a first distance. Region 6040 is the region where person B is located in image 604. The distance between the center point of region 6030 and the center point of region 6040 is the second distance. Accordingly, the distance between the center point of region 6010 or region 6020 and the center point of region 6040 is... .when , , If any one of the following is greater than the second distance threshold, the electronic device can identify person B as a moving object to be eliminated (i.e., the first object).

[0147] Understandably, electronic devices can also follow the S704 method. Figure 6 The person C and the car in image group B shown are also identified as moving objects to be eliminated (i.e., the first object). The specific process and... Figure 9 The embodiments shown are similar and will not be described again.

[0148] In some embodiments, if the first image includes a third object, but other images in the first group do not include the third object, the electronic device may also determine that the third object is a moving object to be eliminated (i.e., the first object).

[0149] In this application, the center point of any object can be the center point of the rectangle when the object is converted into a rectangle. Specifically, when converting the object into a rectangle, the widest line segment of the object can be used as one pair of opposite sides of the rectangle, and the highest line segment as the other pair of opposite sides. However, this is not a limitation; the center point can also be the centroid of the irregular object itself.

[0150] In some embodiments, the electronic device may first set the focal length of each image in the first group to the same focal length before performing coordinate registration, thereby reducing processing errors.

[0151] S705: The electronic device removes the first object from the first image of the first group to obtain the target image.

[0152] Specifically, the electronic device can first determine a second image, including a second object, from the first group. The second image is any image in the first group other than the first image, and there is at least one second image. The second object is used to cover or replace the first object in the first image. The position of the second object in the standard coordinate system obtained by the coordinate registration is the same as the position of the first object in the same standard coordinate system.

[0153] For example, such as Figure 9 As shown, the first image (i.e., image 604) includes a moving object to be eliminated: person B, and the area where person B is located is region 6040. The electronic device can convert the irregular region 6040 into... Figure 10A The rectangle 1040 shown can be represented as follows: At this point, the first object can be equated to rectangle 1040. The electronic device can determine the position of the first group of images within the aforementioned standard coordinate system. The area in question, for example, is as follows: Figures 10B-10D As shown. Among them, Figure 10B The area 1010 shown is the location in image 601. The area where it is located Figure 10C The area 1020 shown is the location in image 602. The area where it is located Figure 10D The area 1030 shown is the location in image 603. The region in which it is located. Assuming that the similarity between region 1010 and region 1020 is greater than or equal to the seventh preset threshold, and the similarity between region 1030 and regions 1010 and 1020 is less than the seventh preset threshold, then either region 1010 or region 1020 can be used as the second object.

[0154] Assume image 601 is the second image determined by the electronic device, and region 1010 is the second object determined by the electronic device. Then, as follows... Figure 11 As shown, electronic devices can use Figure 11 Region 1010 (i.e., the second object) in image 601 (i.e., the second image) shown in (A) is overwritten or replaced. Figure 11 Region 1040 (i.e., the first object) in image 604 (i.e., the first image) shown in (B) is used to obtain Figure 11 The target image 1100 shown in (C) does not include person B (i.e., the moving object).

[0155] In addition to the cases listed above, in specific implementations, there can be multiple second images, and the second object can be the image content obtained by splicing multiple second images.

[0156] In some embodiments, after the electronic device covers or replaces the first object with the second object, the electronic device can also process the edges of the second object in the target image to make the second object and other image content of the target image more coordinated, the transition more natural, the realism stronger, and the user experience better.

[0157] In this application, the electronic device can be in accordance with Figures 4-9 , Figures 10A-10D , Figure 11 The illustrated embodiment acquires a target image and recommends the target image to the user. Specific examples are as follows: Figure 12 , Figures 13A-13B As shown, users do not need to manually trigger the process of eliminating moving objects, making it more convenient for users.

[0158] Please see Figure 12 , Figure 12 An exemplary user interface 120 of a gallery application on an electronic device such as a smartphone is shown. The electronic device can detect touch operations (e.g., clicks) applied to the icon of the gallery application, which is displayed on the desktop (also known as the main interface) of the electronic device. In response to the touch operation, the electronic device can display... Figure 12 The user interface 120 shown is the main interface of a gallery application. That is, a user can click the gallery application icon to open the gallery application's user interface 120. However, users can also open the user interface 120 from other applications; for example, a user can click the album control in a social media application to open the user interface 120, or a user can click the control 305 in the camera application's user interface 30 to open the user interface 120.

[0159] like Figure 12As shown, the user interface 120 may include controls 121, a photo album list 122, and a gallery function list 123. Wherein: Control 121 can be referred to as search bar 121. Search bar 121 can be used to receive information input by the user. The electronic device can search for images or videos stored in the electronic device based on the information input by the user, thereby obtaining images or videos that match the information input by the user, and displaying the matching images or videos to the user.

[0160] The photo album list 122 may include one or more image categories, such as camera category 122A, all image category 122B, similar image category 122C, etc. Each image category may include one or more images or videos. The electronic device can classify these images or videos into one or more of the above image categories according to their source, content, etc. For example, images and videos captured by the electronic device through camera 193 belong to camera category 122A. Images captured by the electronic device through camera 193, obtained from other devices, or downloaded from the Internet belong to all image category 122B. Similar image category 122C may include: the electronic device grouping multiple stored images to obtain at least one image in a first group. The following embodiment uses similar image category 122C to include two first groups: group 1 and group 2, where group 1 is... Figure 5 The following explanation uses image group A as an example.

[0161] The gallery function list 123 may include one or more function options, such as photo function option 123A, gallery function option 123B, timeline function option 123C, and discovery function option 123D. The electronic device can detect a user's touch operation (e.g., a click) on the photo function option 123A, and in response to this touch operation, the electronic device can display images and videos captured by the camera 193. Not limited to this, the electronic device can also detect a user's touch operation (e.g., a click) on the camera category 122A, and in response to this touch operation, the electronic device can also display images and videos captured by the camera 193. When the gallery function option 123B is selected, the electronic device can display... Figure 12 The user interface 120 shown.

[0162] Please see Figure 13A , Figure 13A This example illustrates yet another user interface 130 for a gallery application on an electronic device such as a smartphone. The electronic device can detect actions performed on... Figure 12 Touch operations (e.g., click operations) on the similar image classification 122C in the user interface 120 shown, in response to which the electronic device can display Figure 13AThe user interface 130 shown.

[0163] like Figure 13A As shown, the user interface 130 may include controls 131, a list of similar images 132, and an area 133. Wherein: Control 131 can display the text information: Similar Images (Group 1). The electronic device can detect user touch operations (e.g., click operations) on control 131, and in response to the touch operation, the electronic device can display... Figure 13B The user interface 130 shown is included. Figure 13B The user interface 130 shown includes an option list 134, which includes similar pictures (all) 134A, similar pictures (group 1) 134B, and similar pictures (group 2) 134C. When similar pictures (group 1) 134B is selected, the electronic device can display... Figure 13A and Figure 13B The user interface 130 is shown. The similar image list 132 of the user interface 130 includes images from a first group obtained by the electronic device. The electronic device can detect a touch operation (e.g., a click operation) applied to the similar image (all) 134A, and in response to this touch operation, the electronic device can display images from all the first groups (i.e., group 1 and group 2) obtained by grouping. The electronic device can also detect a touch operation (e.g., a click operation) applied to the similar image (group 2) 134C, and in response to this touch operation, the electronic device can display another first group (e.g., group 3) obtained by grouping. Figure 6 The image shown is from image group B.

[0164] The list of similar images 132 may include images 132A, 132B, 132C, and 132D, which are four images that form a first grouping obtained by grouping electronic devices. Figure 5 The images in image group A shown.

[0165] Region 133 may include a title 133A and region 133B. Title 133A is used to display the text information: Smart Recommendation. Region 133B may display the text information: Recommendation "Moving Object Elimination". Electronic devices can display in region 133B: Electronic devices are based on Group 1 (i.e. Figure 5 The image group A shown is a thumbnail of the target image obtained by eliminating moving objects. The electronic device can detect touch operations (e.g., click operations) applied to area 133B, and in response to the touch operation, the electronic device can display the target image. When the electronic device displays the target image, the user can perform various operations on the target image, such as editing, deleting, and selecting.

[0166] In some embodiments, the electronic device may also receive a user operation to select multiple images and trigger a function to eliminate moving objects. In response to this user operation, the electronic device can recognize the multiple images selected by the user as a group of images to be processed (i.e., a first group), and then perform moving object elimination based on this first group. That is, the electronic device can eliminate moving objects based on a group of images manually selected by the user, as illustrated in specific examples. Figures 14-15 As shown.

[0167] Please see Figure 14 , Figure 14 An exemplary diagram of human-computer interaction is shown. Figure 14 The user interface 141 shown in (A) is a user interface for selecting multiple images before the user clicks control 1413D. Figure 14 The user interface 142 shown in (B) is the user interface after the user clicks control 1413D.

[0168] like Figure 14 As shown in (A), the user interface 141 may include a title 1411, an image list 1412, and image function options 1413. Wherein: The title 1411 may include a control 1411A and text information 1411B. The text information 1411B may be determined by the number of images selected by the user. Figure 14 In the user interface 141 shown in (A), the text information 1411B is: 4 items have been selected, which indicates that the number of pictures selected by the user is 4.

[0169] Image list 1412 may include one or more images, such as images 1412A, 1412B, 1412C, 1412D, 1412E, and 1412F. Image 1412A may include selection box 1412A-1. Figure 14 The selection box 1412A-1 shown in (A) is in a selected state, indicating that the user has selected image 1412A. Similarly, image 1412B may include selection box 1412B-1, image 1412D may include selection box 1412D-1, and image 1412E may include selection box 1412E-1. Figure 14 In (A), selection boxes 1412B-1, 1412D-1, and 1412E-1 are all in a selected state, indicating that the user has selected images 1412B, 1412D, and 1412E. Image 1412C may include selection box 1412C-1, and image 1412F may include selection box 1412F-1. Figure 14The selection boxes 1412C-1 and 1412F-1 shown in (A) are both in an unselected state, indicating that the user has not selected images 1412C and 1412F.

[0170] Image function option 1413 may include one or more function options, such as sharing function option 1413A, deletion function option 1413B, select all function option 1413C, move object elimination function option 1413D, and more function options 1413E.

[0171] The electronic device can detect touch operations (e.g., click operations) applied to the moving object removal function option 1413D. In response to this touch operation, the electronic device determines that the user has selected images 1412A, 1412B, 1412D, and 1412E. Then, the electronic device can obtain the similarity between any two images among these four images. An explanation of obtaining the similarity is provided in the description of the grouping process above, and will not be repeated here. Assuming that the similarity between any two images is greater than or equal to a first threshold, the electronic device determines that these four images form a group of images to be processed (i.e., the first group). The electronic device can then perform moving object removal based on this first group to obtain the target image after removing the moving object. The specific process is described in the above... Figure 7 The illustrated embodiment will not be described in detail again. That is, the user can manually select multiple images to be processed (i.e., the first group) and trigger the moving object removal function by clicking the moving object removal option 1413D. After obtaining the target image, the electronic device can display... Figure 14 User interface 142 shown in (B).

[0172] like Figure 14 As shown in (B), the user interface 142 may include an image list 1421 and a prompt box 1422. The prompt box 1422 may include a prompt message and an area 1422A. Area 1422A can be used to display a thumbnail of the target image after the moving object has been eliminated. The prompt box 1422 may include the text message: "The image obtained after 'Moving Object Elimination' has been stored in 'All Images'." The text message in the prompt box 1422 indicates that the electronic device stores the target image in the gallery application, and the target image belongs to the All Images category 122B. The image list 1421 may include area 1421A and... Figure 14 The images in image list 1412 shown in (A) are used to display a thumbnail of the target image after the moving object has been removed. The areas of the target image displayed in area 1421A and the target image displayed in area 1422A may be different.

[0173] Understandably, the electronic device can detect touch operations (e.g., click operations) applied to control 1411A. In response to this touch operation, the electronic device can deselect selection boxes 1412A-1, 1412B-1, 1412C-1, 1412D-1, 1412E-1, 1412F-1, and the picture function option 1413. Furthermore, in response to this touch operation, the electronic device can change the text information 1411B to all pictures. At this time, the user interface displayed by the electronic device can be: user click... Figure 12 After all the images in the user interface 120 are categorized 122B, the user interface displayed by the electronic device in response to the click operation.

[0174] Understandably, if the multiple images selected by the user do not meet the requirements—for example, if two images have a similarity level less than a first threshold—the electronic device cannot eliminate the moving object based on the user's selections. Therefore, the electronic device can prompt the user to reselect images. See the example below for details. Figure 15 The illustrated embodiment.

[0175] Please see Figure 15 , Figure 15 Another example of human-computer interaction is illustrated below. Figure 15 The user interface 141 shown in (A) is a user interface for selecting multiple images before the user clicks control 1413D. Figure 15 The user interface 150 shown in (B) is the user interface after the user clicks control 1413D.

[0176] like Figure 15 As shown in (A), user interface 141 and Figure 14 The user interface 141 shown in (A) is similar, except that the images selected by the user are changed to: image 1412A, image 1412C, and image 1412F, and the text information 1411B is also changed to "3 items selected".

[0177] The electronic device can detect touch operations (e.g., click operations) applied to the moving object elimination function option 1413D. In response to this touch operation, the electronic device determines that the user has selected images 1412A, 1412C, and 1412F. The electronic device then obtains the similarity between any two of these three images; the explanation of obtaining the similarity is detailed in the above description of the grouping process and will not be repeated here. If the similarity between any two of the images is less than a first threshold, the electronic device determines that these three images cannot be processed as a first group. At this point, the electronic device will display... Figure 15 User interface 150 shown in (B).

[0178] like Figure 15As shown in (B), the user interface 150 may include an image list 151 and a prompt box 152. The image list 151 may include... Figure 15 The images in image list 1412 shown in (A). Prompt box 152 may include the text message: "Moving object elimination failed, please select an image from the same scene."

[0179] In some embodiments, the electronic device may also receive a user operation for selecting an image and triggering a function to eliminate moving objects. In response to the user operation, the electronic device can recognize the user-selected image as a first image, then obtain a first group including the first image, and eliminate moving objects based on the first group. That is, the electronic device can eliminate moving objects based on a first image manually selected by the user, as illustrated in specific examples. Figure 16 As shown.

[0180] Please see Figure 16 , Figure 16 Another example of human-computer interaction is illustrated below. Figure 16 The user interface 160 shown in (A) is the user interface before the user clicks control 162C. Figure 16 The user interface 160 shown in (B) is the user interface after the user clicks control 162C.

[0181] like Figure 16 As shown in (A), the user interface 160 may include an image 161 and image function options 162. The electronic device can detect a touch operation (e.g., a click) applied to a thumbnail of the image 161, and in response to the touch operation, the electronic device can display... Figure 16 User interface 160 is shown in (A). A thumbnail of image 161 can be displayed... Figure 3 In control 305 shown. Not limited to this, the thumbnail of image 161 can also be displayed in any image list, for example... Figure 15 The image list shown in (B) is in 151.

[0182] Image function option 162 may include one or more function options for image 161, such as sharing function option 162A, deletion function option 162B, moving object removal function option 162C, and more function option 162D.

[0183] The electronic device can detect a touch operation (e.g., a click) applied to the moving object removal function option 162C. In response to this touch operation, the electronic device can identify image 161 as the first image. Then, the electronic device can acquire multiple images whose similarity to image 161 is greater than or equal to a first threshold. Image 161 and the acquired multiple images are identified as a first group. The method for calculating the similarity is described in the grouping process above and will not be repeated here. The electronic device can then remove moving objects based on this first group to obtain the target image after removing the moving objects. The specific process is described in the above... Figure 7 The illustrated embodiment will not be described in detail again. After obtaining the target image, the electronic device can display it. Figure 16 User interface 160 is shown in (B). That is, the user can trigger the function to eliminate moving objects by clicking the moving object elimination function option 162C.

[0184] compared to Figure 16 User interface 160 shown in (A) Figure 16 The user interface 160 shown in (B) also includes a prompt box 163, which may include the text information: "The image obtained by 'Moving Object Elimination' has been stored in 'All Pictures'." The text information in the prompt box 163 indicates that the electronic device has stored the target image in the gallery application, and that the target image belongs to the All Pictures category 122B. The user can click... Figure 12 The user interface 120 shown contains all image categories 122B for viewing the target image. That is, the user can manually select the first image and trigger the function to eliminate moving objects.

[0185] Understandably, if the electronic device cannot obtain an image with a similarity greater than or equal to the first threshold to image 161, the electronic device confirms that image 161 cannot be used as the first image for eliminating the moving object. Therefore, the electronic device can prompt the user to reselect an image, as shown in the specific examples. Figure 15 The embodiments shown are similar and will not be described again.

[0186] Not limited to Figure 16 The example shown can also be implemented in practice by the user. Figure 14 (A) and Figure 15 The user interface 141 shown in (A) selects an image and then clicks the moving object elimination function option 1413D. In response to this click operation, the electronic device can recognize the image selected by the user as a first image and eliminate moving objects based on the first image, which is not limited in this embodiment.

[0187] In some embodiments, the electronic device may also receive a user operation on an image, whereby the user operation selects the subject to be retained in the image and triggers a function to eliminate moving objects. In response to the user operation, the electronic device can acquire a first group including the images and identify the images as the first image of the first group. Then, the electronic device can eliminate moving objects based on the first group, the first image, and the user-selected subject. That is, the electronic device can eliminate moving objects based on the subject manually selected by the user, as illustrated in the specific example below. Figure 17 As shown.

[0188] Please see Figure 17 , Figure 17 Another example of human-computer interaction is illustrated below. Figure 17 The user interface 171 shown in (A) is the user interface before the user clicks control 1711B. Figure 17 The user interface 172 shown in (B) is the user interface after the user clicks control 1711B.

[0189] like Figure 17 As shown in (A), the user interface 171 may include a picture 161 and an elimination function option 1711. The electronic device can detect the action performed on... Figure 16 (A) The touch operation (e.g., click operation) of the moving object elimination function option 162C in the user interface 160 shown in (A) responds to the touch operation, and the electronic device can display Figure 17 The user interface 171 shown in (A) is shown.

[0190] The elimination function option 1711 may include a smart elimination function option 1711A and a manual elimination function option 1711B. The electronic device can detect a touch operation (e.g., a click) applied to the smart elimination function option 1711A. In response to this touch operation, the electronic device can identify image 161 as the first image. Then, the electronic device can acquire multiple images whose similarity to image 161 is greater than or equal to a first threshold. Image 161 and the acquired multiple images are identified as a first group. The method for calculating the similarity is described in the above grouping process description and will not be repeated here. Then, the electronic device can eliminate moving objects based on this first group to obtain the target image after eliminating the moving objects. The specific process is described in the above... Figure 7 The illustrated embodiment will not be described in detail again. After obtaining the target image, the electronic device can display it. Figure 16 The user interface 160 is shown in (B). That is, the user can trigger the function to eliminate moving objects by clicking the smart elimination function option 1711A.

[0191] The electronic device can also detect touch operations (such as clicks) performed on the manual clear function option 1711B, and in response to the touch operation, the electronic device can display... Figure 17 User interface 172 is shown in (B). User interface 172 may include picture 161 and function options 1721. Wherein: The electronic device can detect touch operations (such as click operations) applied to any object in image 161, and in response to the touch operation, the electronic device can identify the object as the subject currently selected by the user. Figure 17 In the user interface 172 shown in (B), the user has selected a character object in area 161A.

[0192] Function option 1721 may include a confirm option 1721A and a cancel option 1721B. When the user has selected an object, the electronic device can detect a touch operation (e.g., a click operation) applied to the confirm option 1721A. In response to this touch operation, the electronic device can acquire multiple images whose similarity to image 161 is greater than or equal to a first threshold. Image 161 and the acquired multiple images are identified as a first group. The method for calculating the similarity can be found in the description of the grouping process above, and will not be repeated here. Then, the electronic device can identify image 161 as the first image in the first group, the object selected by the user as the subject to be photographed in the first group, and eliminate moving objects based on the first group. The specific elimination process can be found in the above description. Figure 7 The illustrated embodiment will not be described in detail again. Finally, the electronic device can obtain the target image after the moving object has been eliminated, at which point it can be displayed. Figure 16 User interface 160 is shown in (B). That is, the user can manually select the subject to be photographed and trigger the function to eliminate moving objects by confirming option 1721A.

[0193] It should be noted that if the user does not select an object, the function of eliminating the moving object will not be triggered even if the user clicks the OK option 1721A.

[0194] The electronic device can also detect touch operations (such as clicks) applied to the cancel option 1721B, and in response to such touch operations, the electronic device can display... Figure 17 User interface 171 shown in (A) or Figure 16 The user interface 160 shown in (A) is shown.

[0195] In some embodiments, the electronic device may also receive a user operation on an image, wherein the user operation is used to select a moving object to be eliminated in the image and trigger a function to eliminate the moving object. In response to the user operation, the electronic device may acquire a first group including the images and identify the images as the first image of the first group. Then, the electronic device may eliminate the moving object based on the first group, the first image, and the moving object selected by the user. That is, the electronic device can eliminate moving objects manually selected by the user, as illustrated in the specific example below. Figure 18 As shown.

[0196] Please see Figure 18 , Figure 18 An example is shown of a user interface 180 of a gallery application on an electronic device such as a smartphone. The electronic device can detect actions performed on... Figure 17 The touch operation (e.g., a click operation) of the manual cancellation function option 1711B in the user interface 171 shown in (A) responds to the touch operation, and the electronic device can display Figure 18 The user interface shown is 180.

[0197] like Figure 18 As shown, the user interface 180 may include an image 161 and function options 181. The electronic device can detect a touch operation (e.g., a click operation) applied to any object in the image 161, and in response to the touch operation, the electronic device can identify that object as the currently selected moving object to be eliminated by the user. Figure 18 In the user interface 180 shown, the user has selected a character object in area 161B.

[0198] Function option 181 may include a confirm option 181A and a cancel option 181B. When the user has selected an object, the electronic device can detect a touch operation (e.g., a click operation) applied to the confirm option 181A. In response to this touch operation, the electronic device can acquire multiple images whose similarity to image 161 is greater than or equal to a first threshold. Image 161 and the acquired multiple images are identified as a first group. The method for calculating the similarity can be found in the description of the grouping process above, and will not be repeated here. Then, the electronic device can identify image 161 as the first image in the first group, and the object selected by the user as the moving object to be eliminated in the first group, and eliminate the moving object based on the first group. The specific elimination process can be found in the above description. Figure 7 The illustrated embodiment will not be described in detail again. Finally, the electronic device obtains the target image after the moving object has been eliminated, at which point it can be displayed. Figure 16 User interface 160 is shown in (B). That is, the user can manually select the moving object to be eliminated and trigger the function to eliminate the moving object by confirming option 181A.

[0199] It should be noted that if the user does not select an object, the function of eliminating the moving object will not be triggered even if the user clicks the OK option 181A.

[0200] The electronic device can also detect touch operations (such as clicks) applied to the cancel option 181B, and in response to such touch operations, the electronic device can display... Figure 17 User interface 171 shown in (A) or Figure 16 The user interface 160 shown in (A) is shown.

[0201] In some embodiments, after determining the first group, before determining that the object selected by the user is a moving object to be eliminated in the first group, the electronic device may first obtain the distance between any two images of the user-selected object in the first group. Only when this distance is greater than or equal to a fifth preset threshold will the electronic device determine that the user-selected object is a moving object to be eliminated in the first group, and then eliminate the moving object based on the first group. When there is a distance less than the fifth preset threshold, the electronic device determines that the user-selected object cannot be used as the first object for elimination. Therefore, the electronic device may prompt the user to reselect the moving object. Specific examples and... Figure 15 The illustrated embodiment is similar and will not be described again. For details on the process of determining whether the user-selected object is the first object, please refer to [link to documentation]. Figure 7 The description of S704 will not be repeated here.

[0202] Beyond the examples above, in practical implementations, users can select both the subject being photographed and the moving object. For instance, after selecting the subject, the user can click... Figure 17 The "Confirm" option 1721A in the user interface 172 shown in (B). In response to this click, the electronic device can display... Figure 18 The user interface 180 is shown. After selecting a moving object in the user interface 180, the user can click the "OK" option 181A in the user interface 180. In response to this click operation, the electronic device can eliminate the moving object based on the subject and the moving object selected by the user. This embodiment of the application does not limit this process.

[0203] Understandably, Figure 12 , Figures 13A-13B , Figures 14-18 In the illustrated embodiment, the description of obtaining the first group can be found in the above description of the grouping process, and the description of determining the subject being photographed, determining the moving object, and determining the first image can be found in the above description. Figure 7 The process shown is not repeated here.

[0204] Based on the above Figures 1-9 , Figures 10A-10D , Figures 11-12, Figures 13A-13B , Figures 14-18 The following describes the image processing method provided in this application, based on some embodiments shown.

[0205] Please see Figure 19 , Figure 19 This application provides an image processing method. This method can be applied to... Figure 1 The electronic device shown. This method can also be applied to... Figure 2 The electronic device shown. The method includes, but is not limited to, the following steps: S101: The electronic device determines multiple images that meet the first condition.

[0206] Specifically, the first condition may include: the similarity between any two images in the plurality of images is greater than or equal to a first threshold, and these plurality of images may be referred to as a first group. In some embodiments, the first condition may further include at least one of the following: receiving a first operation, the shooting time of any one of the plurality of images being within a first range, and the shooting location of any one of the plurality of images being within a second range, wherein the first operation is used to select the plurality of images. An example of determining the plurality of images through the first operation can be found in [reference needed]. Figures 14-15 The illustrated embodiment will not be described again. The method for calculating similarity and determining multiple images (i.e., the first group) that meet the first condition can be found in the description of the grouping process above, and will not be repeated here.

[0207] The multiple images determined by the electronic device can be obtained by the electronic device through the default shooting mode, so that the embodiments of this application can be implemented in general scenarios and the application is more extensive.

[0208] S102: The electronic device determines a first image that satisfies the second condition from a plurality of images, the first image including the first object.

[0209] Specifically, the first object is the moving object to be eliminated in the first image. The second condition includes at least one of the following: the sharpness of the subject in the first image is greater than or equal to a second threshold, and the number of other objects in the first image besides the subject is less than a third threshold. The electronic device can process and judge each image in the first group to obtain a first image in the first group that meets the second condition, as detailed in [reference needed]. Figure 7 The description of S703 will not be repeated here.

[0210] In some embodiments, the second condition may further include: receiving a user action for selecting a first image. An example of determining the first image through this user action can be found in [link to relevant documentation]. Figure 16 The embodiments shown are not described in detail here.

[0211] In some embodiments, prior to S102, the electronic device may first perform semantic segmentation on each image in the aforementioned plurality of images (i.e., the first group) to identify the objects included in the image (e.g., people, buildings, cars, etc.), as detailed in [reference needed]. Figure 7 S701 Figure 8 The illustrated embodiment.

[0212] In some embodiments, prior to S102, the method further includes: the electronic device determining a subject that satisfies a third condition from the plurality of images (i.e., the first group). The third condition includes at least one of the following: receiving a second operation; the sharpness of the subject in any image of the first group being greater than or equal to a fourth threshold; the focus point of any image of the first group being located in the area where the subject is located; the area of ​​the subject in any image of the first group being greater than or equal to a fifth threshold; and the subject belonging to a preset category. The second operation is used to select the subject. The electronic device can process and judge each object in the first group to obtain the subject in the first group that satisfies the third condition. See [link to relevant documentation] for details. Figure 7 The description of S702 will not be repeated here. An example of determining the subject through the second operation can be found in [link to documentation]. Figure 17 The embodiments shown are not described in detail here.

[0213] In some embodiments, after determining the first image, the electronic device can determine the first object in the first group. The electronic device can first arrange the images in the first group in the same coordinate system (i.e., Figure 7 (Coordinate registration in S704), and then based on this coordinate system, the distance between any two images of each object in the first group is obtained. When this distance is greater than or equal to a fifth preset threshold, the electronic device can identify the object as the first object. See [link to documentation] for details. Figure 7 The description of S704 will not be repeated here.

[0214] S103: The electronic device determines a second image from a plurality of images, the second image including a second object.

[0215] Specifically, the second image can be at least one image from a set of images, excluding the first image. For example, the second image can be the image with the highest similarity to the first image from a set of images, or the second image can be at least one image from a set of images whose similarity to the first image is greater than a preset threshold. The second object can be obtained from a single second image, or it can be obtained by stitching together at least one second image.

[0216] The position of the second object in the second image corresponds to the position of the first object in the first image. That is, when the first and second images are in the same first coordinate system, the position of the second object in the second image is the same as the position of the first object in the first image. Here, the first coordinate system is the coordinate system obtained after coordinate registration of the electronic device, for example... Figure 7 The standard coordinate system shown in S704. An example of an electronic device determining a second object, and of a second image including the second object, can be found in [reference needed]. Figure 9 , Figures 10A-10D The embodiments shown are not described in detail here.

[0217] S104: The electronic device uses a second object to cover or replace the first object to obtain the target image.

[0218] Specifically, an example of an electronic device using a second object to cover or replace a first object can be found here. Figure 7 S705 Figure 9 , Figures 10A-10D , Figure 11 The embodiment shown. An example of the obtained target image can be found in [reference needed]. Figure 11 Image 1100 is shown in (C). Compared to the first image, the target image not only eliminates the moving object, but also displays realistic image content in the area where the moving object was eliminated, resulting in a better user experience.

[0219] Understandably, Figure 19 The process shown can be executed in the background by the electronic device, without the user's awareness. After obtaining the target image, the electronic device can recommend the target image to the user without requiring the user to manually trigger the function to eliminate moving objects. See the specific example below. Figures 13A-13B The illustrated embodiment.

[0220] exist Figure 19 In the method shown, the electronic device can determine a first group within the same shooting scene. The images in the first group can be images obtained by the electronic device using a default shooting mode. The electronic device can then eliminate moving objects (i.e., the first object) based on the first group. Therefore, users do not need to use a specific shooting mode to eliminate moving objects, making it more convenient and applicable to a wider range of scenarios. Furthermore, the image content used to fill or cover the moving object is obtained based on a real second image, resulting in better display effects and a better user experience.

[0221] Furthermore, the electronic device can eliminate moving objects without the user noticing and recommend the resulting target image to the user for viewing, without requiring the user to manually trigger the elimination function, making it more convenient to use. Users can also manually select the first group, the first image, the subject being photographed, or the moving object, offering greater flexibility.

[0222] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, the processes or functions described in this application are generated, in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state drive).

[0223] In summary, the above description is merely an embodiment of the technical solution of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made according to the disclosure of the present invention should be included within the scope of protection of the present invention.

[0224] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. An image processing method, characterized in that, Applied to electronic devices, the method includes: A second image similar to a first image is determined based on the degree of similarity of image features, wherein the first image includes a first object, which is an object to be eliminated in the first image, and the second image includes a second object; Use the second object to cover or replace the first object to obtain the target image.

2. The method as described in claim 1, characterized in that, The step of determining a second image similar to the first image based on the similarity of image features includes: Based on the difference in shooting time and / or the distance between shooting locations, as well as the similarity of image features, a second image similar to the first image is determined.

3. The method as described in claim 1 or 2, characterized in that, The difference between the shooting time of the second image and the shooting time of the first image is less than or equal to a first time threshold; and / or, the distance between the shooting location of the second image and the shooting location of the first image is less than or equal to a first distance threshold.

4. The method according to any one of claims 1-3, characterized in that, The position of the second object in the second image corresponds to the position of the first object in the first image.

5. The method according to any one of claims 1-4, characterized in that, Before using the second object to overwrite or replace the first object, the method further includes: Display a first interface, the first interface including the first image and the first control; The system receives operations on the first control, which is used to perform image processing on the first image.

6. The method according to any one of claims 1-5, characterized in that, The step of using the second object to cover or replace the first object includes: Receive the user's operation to select the first object, and use the second object to overwrite or replace the first object.

7. The method according to any one of claims 1-6, characterized in that, Before determining the second image similar to the first image based on the similarity of image features, the method further includes: Receive a first operation, which is used to select the first image.

8. The method according to any one of claims 1-7, characterized in that, The sharpness of the subject in the first image is greater than or equal to a second threshold, and / or the number of other objects in the first image besides the subject is less than a third threshold.

9. The method as described in claim 8, characterized in that, The method further includes: The subject being photographed is determined from the first image and the second image to satisfy a first condition; the first condition includes at least one of the following: the sharpness of the subject being photographed in the first image and the second image is greater than or equal to a fifth threshold; the focus point in the first image and the second image is located in the area where the subject being photographed is located; the area of ​​the subject being photographed in the first image and the second image is greater than or equal to a sixth threshold; the subject being photographed belongs to a preset category; and a third operation is received, the third operation being used to select the subject being photographed.

10. The method as described in claim 1, characterized in that, Before using the second object to cover or replace the first object, the method further includes: Determine the center point of the third object in the first image, the center point of the fourth object in the first image, the center point of the fifth object in the second image, and the center point of the sixth object in the second image; the third object and the fifth object have the same attributes, and the fourth object and the sixth object have the same attributes; Place the center point of the third object and the center point of the fifth object in the same coordinate system, set them as the same coordinate origin, and establish a first coordinate system based on the coordinate origin; Based on the first coordinate system, determine the first distance between the center point of the fourth object and the center point of the sixth object; When the first distance is greater than or equal to the seventh threshold, the fourth object is determined to be the first object.

11. The method as described in claim 10, characterized in that, The third object and the fifth object represent objects in the same position.

12. An electronic device, characterized in that, The electronic device includes at least one memory and at least one processor, the at least one memory being coupled to the at least one processor, the at least one memory being used to store a computer program, the at least one processor being used to invoke the computer program, the computer program including instructions that, when executed by the at least one processor, cause the electronic device to perform the method as described in any one of claims 1-11.

13. A computer program product, when run on an electronic device, causes the electronic device to perform the method of any one of claims 1-11.

14. A computer storage medium, characterized in that, Includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1-11.