Image processing method and related apparatus
By displaying markers on the image display page and utilizing image reconstruction technology from different cameras, the image composition is automatically adjusted, solving the problems of users needing to take continuous shots and tedious image editing, achieving stable acquisition of high-quality images and simplifying the image editing process.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2026-01-16
- Publication Date
- 2026-07-30
AI Technical Summary
Users need to take multiple photos in a row to get a satisfactory image, and the post-processing is cumbersome, resulting in a poor photo-taking experience. Furthermore, existing photo-editing methods require complicated manual operations.
By displaying a first marker on the image display page, and utilizing image reconstruction technology from different cameras, the image composition is automatically adjusted to generate a high-quality third image, eliminating the need for complicated user operations.
It enables the stable acquisition of high-quality images without manual operation, improving the photo-taking experience and simplifying the image editing process.
Smart Images

Figure CN2026073067_30072026_PF_FP_ABST
Abstract
Description
An image processing method and related apparatus
[0001] This application claims priority to Chinese Patent Application No. 202510108496.2, filed on January 22, 2025, entitled “An Image Processing Method and Related Apparatus”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of image processing technology, and in particular to an image processing method and related apparatus. Background Technology
[0003] Electronic devices have become an integral part of everyday life for most people, and taking photos has become a fundamental life experience. However, taking good photos remains a challenge for many. To get an ideal photo, people often take multiple shots in succession, wasting a lot of storage space. Furthermore, they have to select the best photo later, resulting in a poor photography experience.
[0004] When you don't get a satisfactory photo, another way is to crop and recompose the existing image. This way, even if the original photo is not satisfactory, you can get a satisfactory image by recomposing it. However, when editing the photo, the user needs to manually perform image cropping and other composition operations, which requires the user to perform complicated manual operations. Summary of the Invention
[0005] This application provides an image processing method and related apparatus that can reliably obtain high-quality images without requiring users to perform complicated manual operations.
[0006] In a first aspect, embodiments of this application provide an image processing method, the method comprising: displaying a first image and a first marker on an image display page; the first marker being used to identify a first object in the first image; when the distance between the first marker and a preset position on the image display page is less than a preset value, obtaining a third image based on the first image or a second image; the second image and the first image being images of the same scene captured by different cameras, and the third image being an image obtained by reconstructing the first image or the second image with the first object as the main subject of the composition; and displaying the third image on the image display page.
[0007] The preset value can be 1mm, 2mm, 3mm, 4mm, 5mm, 1cm, etc. The preset value can be related to the size of the display screen. The larger the display screen, the larger the preset value. In terms of display effect, when there is an overlap between the first mark and the second mark displayed at the preset position, it is considered that the distance between the first mark and the preset position is less than the preset value.
[0008] Since the images are taken by different cameras, the content of the first object in the first image and the second image may not be exactly the same. The fact that both the first image and the second image include the first object can be understood as including the same target, such as the same person or object.
[0009] Reconstruction can be image editing based on existing photos, such as image cropping, image generation, and image style transformation. When cropping an image, it can involve determining which portion of the first or second image contains the first object and has a more aesthetically pleasing composition, and then cropping out that portion of the image.
[0010] In this embodiment, a first mark indicating a first object is displayed on the first image. If the user needs to recompose the first image around the first object, a positional association can be established between the first mark and a preset position on the image display interface, thereby triggering the electronic device to recompose the first image around the first object. In a photography scenario, the image can be reconstructed based on the first mark selected by the user. Thus, when the user takes a picture, the captured image is the reconstructed image (i.e., the third image). The reconstructed image is a graphic with high aesthetic standards. Therefore, when the user uses the assisted composition method of this embodiment, they can stably obtain a high-quality image without complicated manual operations. Similarly, in image editing, rapid image editing can be achieved, reducing the time required for image editing.
[0011] The statement that the distance between the first marker and the preset position is less than a preset value can also be extended to "there is a positional association between the first marker and the preset position of the image display page".
[0012] In one possible implementation, the second image and the first image are images of the same scene captured by different cameras.
[0013] For example, different cameras can be electronic devices with cameras that have different magnification.
[0014] In one possible implementation, a second marker is displayed at a preset position on the image display page.
[0015] In one possible implementation, the reconstruction includes: performing image cropping; or, generating content expansion of the first image and cropping the expanded first image.
[0016] In one possible implementation, the third image is specifically an image obtained by reconstructing either the first image or the second image based on the aesthetics of the composition, using the first object as the main subject.
[0017] When users use the assisted composition method of this embodiment, they can reliably obtain high-quality images without the need for complicated manual operations or high aesthetic requirements.
[0018] In one possible implementation, the existence of location association includes: the distance between the first marker and the preset location is less than a preset value.
[0019] In one possible implementation, the method further includes: in response to a received first operation, moving the first marker until a positional association exists between the first marker and the preset position; the first operation is a movement of the first image on the image display page or a movement of the camera that acquires the first image.
[0020] In one possible implementation, the second marker is displayed automatically upon entering the image display page, or the second marker is displayed in response to an operation on the image display page.
[0021] In one possible implementation, displaying the first image and the first marker on the image display page includes:
[0022] A first image and a third marker are displayed on the image display page, the third marker being used to identify a second object in the first image; the first object and the second object are different.
[0023] In response to the received second operation, the display of the third marker is deleted and the display of the first marker is added, the second operation indicating an update of the markers for objects in the first image.
[0024] In one possible implementation, the method further includes:
[0025] Upon receiving a third operation, the third operation instructs the fixing of the first mark;
[0026] Upon receiving a fourth operation, the display of the first marker is maintained; the fourth operation is either movement of the first image on the image display page or movement of the camera that captures the first image.
[0027] In one possible implementation, the reconstruction includes image cropping, the cropping ratio being consistent with the first image, or determined based on ratio information input by the user.
[0028] In one possible implementation, the method further includes:
[0029] When the distance between the first marker and the preset position on the image display page is less than a preset value, a fusion effect between the first marker and the second marker is displayed.
[0030] In one possible implementation, the method further includes:
[0031] Upon receiving the fifth operation, the fifth operation indicates switching to displaying the image before reconstruction;
[0032] In response to the fifth operation, the third image on the image display page is switched to the first image or the second image.
[0033] In one possible implementation, the image display interface displays a return control, and the fifth operation is an operation performed on the return control. After performing a preset operation on the return control, the image reconstruction is deactivated, the display of the third image is exited, and the first image or the second image is re-displayed.
[0034] In one possible implementation, the image display interface displays a switching control, the fifth operation is a preset operation performed on the switching control, and the method further includes:
[0035] Upon receiving the sixth operation, the image display interface no longer displays the first or second image, but instead displays the reconstructed third image; wherein, the sixth operation includes: performing a preset operation on the switching control and releasing the fifth operation.
[0036] In one possible implementation, the third image is obtained by cropping the first image, and displaying the third image on the image display page includes:
[0037] The first image and an indicator box showing the cropping area of the third image within the first image are displayed on the image display page.
[0038] In one possible implementation, the preset position is located in the central region of the first image.
[0039] In one possible implementation, obtaining the third image based on the first image or the second image includes:
[0040] Using the first object as the main subject of the composition, and based on the aesthetics of the composition, the first image or the second image is reconstructed using image mapping rules or a large model to obtain the third image.
[0041] Secondly, embodiments of this application provide an image processing apparatus, the apparatus comprising:
[0042] The display module is used to display a first image and a first marker on an image display page; the first marker is used to identify a first object in the first image;
[0043] The image reconstruction module is used to obtain a third image based on the first image or the second image when the distance between the first mark and the preset position of the image display page is less than a preset value; the second image and the first image are images captured by different cameras for the same scene, and the third image is an image obtained by reconstructing the first image or the second image with the first object as the main subject of the composition;
[0044] The display module is used to display the third image on the image display page.
[0045] In one possible implementation, a second marker is displayed at a preset position on the image display page.
[0046] In one possible implementation, the refactoring includes:
[0047] Perform image cropping; or,
[0048] The content of the first image is expanded and generated, and the expanded first image is then cropped.
[0049] In one possible implementation, the third image specifically refers to:
[0050] The image is obtained by reconstructing the first image or the second image based on the aesthetics of the composition, using the first object as the main subject.
[0051] In one possible implementation, the existence of location association includes:
[0052] The distance between the first marker and the preset position is less than a preset value.
[0053] In one possible implementation, the display module is further configured to:
[0054] In response to a received first operation, the first marker is moved until a positional association exists between the first marker and the preset position; the first operation is a movement of the first image on the image display page or a movement of the camera that captures the first image.
[0055] In one possible implementation, the second marker is displayed automatically upon entering the image display page, or the second marker is displayed in response to an operation on the image display page.
[0056] In one possible implementation, the display module is specifically used for:
[0057] A first image and a third marker are displayed on the image display page, the third marker being used to identify a second object in the first image; the first object and the second object are different.
[0058] In response to the received second operation, the display of the third marker is deleted and the display of the first marker is added, the second operation indicating an update of the markers for objects in the first image.
[0059] In one possible implementation, the display module is further configured to:
[0060] Upon receiving a third operation, the third operation instructs the fixing of the first mark;
[0061] Upon receiving a fourth operation, the display of the first marker is maintained; the fourth operation is either movement of the first image on the image display page or movement of the camera that captures the first image.
[0062] In one possible implementation, the reconstruction includes image cropping, the cropping ratio being consistent with the first image, or determined based on ratio information input by the user.
[0063] In one possible implementation, the display module is further configured to:
[0064] When the distance between the first marker and the preset position on the image display page is less than a preset value, a fusion effect between the first marker and the second marker is displayed.
[0065] In one possible implementation, the display module is further configured to:
[0066] Upon receiving the fifth operation, the fifth operation indicates switching to displaying the image before reconstruction;
[0067] In response to the fifth operation, the third image on the image display page is switched to the first image or the second image.
[0068] In one possible implementation, the image display interface displays a return control, and the fifth operation is an operation performed on the return control. After performing a preset operation on the return control, the image reconstruction is deactivated, the display of the third image is exited, and the first image or the second image is re-displayed.
[0069] In one possible implementation, the image display interface displays a switching control, the fifth operation is a preset operation performed on the switching control, and the display module is further configured to:
[0070] Upon receiving the sixth operation, the image display interface no longer displays the first or second image, but instead displays the reconstructed third image; wherein, the sixth operation includes: performing a preset operation on the switching control and releasing the fifth operation.
[0071] In one possible implementation, the third image is obtained by cropping the first image, and the display module is specifically used for:
[0072] The first image and an indicator box showing the cropping area of the third image within the first image are displayed on the image display page.
[0073] In one possible implementation, the preset position is located in the central region of the first image.
[0074] In one possible implementation, the image reconstruction module is specifically used for:
[0075] Using the first object as the main subject of the composition, and based on the aesthetics of the composition, the first image or the second image is reconstructed using image mapping rules or a large model to obtain the third image.
[0076] Thirdly, this application provides an image processing device, including: a processor, a memory, a camera, and a bus, wherein: the processor, the memory, and the camera are connected via the bus;
[0077] The camera is used to capture video in real time;
[0078] The memory is used to store computer programs or instructions;
[0079] The processor is used to call or execute programs or instructions stored in the memory, and is also used to call the camera to implement the steps described in the first aspect and any possible implementation of the first aspect.
[0080] Fourthly, this application provides a computer storage medium including computer instructions that, when executed on an electronic device or server, perform the steps described in the first aspect and any possible implementation thereof.
[0081] Fifthly, this application provides a computer program product that, when run on an electronic device or server, performs the steps described in the first aspect and any possible implementation thereof.
[0082] Sixthly, this application provides a chip system including a processor for supporting a device in implementing the functions involved in the foregoing aspects, such as transmitting or processing data involved in the foregoing methods; or, information. In one possible design, the chip system further includes a memory for storing program instructions and data necessary for executing or training the device. This chip system may be composed of chips or may include chips and other discrete devices. Attached Figure Description
[0083] To more clearly illustrate the technical methods of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly described below.
[0084] Figure 1 is a schematic diagram of an application architecture provided in an embodiment of this application;
[0085] Figure 2 is a schematic diagram of an application architecture provided in an embodiment of this application;
[0086] Figure 3 is a schematic diagram of an image processing method provided in an embodiment of this application;
[0087] Figures 4 to 15 are schematic diagrams of an interface provided in an embodiment of this application;
[0088] Figure 16 is a schematic diagram of a device provided in an embodiment of this application;
[0089] Figure 17 is a schematic diagram of an apparatus provided in an embodiment of this application. Detailed Implementation
[0090] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.
[0091] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.
[0092] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0093] The terms “substantially,” “about,” and similar terms used herein are used as approximations, not as terms of degree, and are intended to take into account the inherent biases of measurements or calculations known to those skilled in the art. Furthermore, the term “may” used in describing embodiments of this application means “one or more possible embodiments.” The terms “use,” “using,” and “used” used herein are to be considered synonymous with the terms “utilize,” “utilizing,” and “utilized,” respectively. Additionally, the term “exemplary” is intended to refer to an instance or illustration.
[0094] For ease of understanding, the structure of the terminal 100 provided in the embodiments of this application will be illustrated below. Refer to Figure 1, which is a schematic diagram of the structure of the terminal device provided in the embodiments of this application.
[0095] As shown in Figure 1, the terminal 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0096] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the terminal 100. In other embodiments of this application, the terminal 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0097] Processor 110 may include one or more processing units, such as application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.
[0098] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.
[0099] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0100] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0101] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple I2C buses. The processor 110 can couple to the touch sensor 180K, charger, flash, camera 193, etc., through different I2C bus interfaces. For example, the processor 110 can couple to the touch sensor 180K through the I2C interface, enabling the processor 110 and the touch sensor 180K to communicate through the I2C bus interface, thereby realizing the touch function of the terminal 100.
[0102] The I2S interface can be used for audio communication. In some embodiments, the processor 110 may include multiple I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to enable communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the I2S interface to enable the function of answering phone calls through a Bluetooth headset.
[0103] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled via the PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 via the PCM interface, enabling the function of answering phone calls through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0104] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the UART interface to enable music playback through Bluetooth headphones.
[0105] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI) and a display serial interface (DSI). In some embodiments, the processor 110 and the camera 193 communicate via the CSI interface to enable the shooting function of the terminal 100. The processor 110 and the display screen 194 communicate via the DSI interface to enable the display function of the terminal 100.
[0106] The GPIO interface can be configured via software. It can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to a camera 193, a display screen 194, a wireless communication module 160, an audio module 170, a sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0107] Specifically, the video captured by the camera 193 (including image frame sequences, such as the first and second images in this application) can be, but is not limited to, transmitted to the processor 110 through the interface (such as the CSI interface or GPIO interface) described above for connecting the camera 193 and the processor 110.
[0108] The processor 110 can retrieve instructions from the memory and perform video processing on the video captured by the camera 193 based on the retrieved instructions (such as image stabilization processing, perspective distortion correction processing, etc. in this application) to obtain processed images (such as the third image and the fourth image in this application).
[0109] The processor 110 can, but is not limited to, transmit the processed image to the display screen 194 through the interface (e.g., DSI interface or GPIO interface) described above for connecting the display screen 194 and the processor 110, so that the display screen 194 can display video.
[0110] USB port 130 is a USB standard compliant interface, specifically a Mini USB port, Micro USB port, USB Type-C port, etc. USB port 130 can be used to connect a charger to charge terminal 100, and can also be used for data transfer between terminal 100 and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other electronic devices, such as AR devices.
[0111] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the terminal 100. In other embodiments of this application, the terminal 100 may also adopt different interface connection methods or a combination of multiple interface connection methods as described in the above embodiments.
[0112] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the terminal 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device via the power management module 141.
[0113] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.
[0114] The wireless communication function of terminal 100 can be implemented through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor.
[0115] Antennas 1 and 2 are used to transmit and receive electromagnetic wave signals. Each antenna in terminal 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.
[0116] The mobile communication module 150 can provide solutions for wireless communication applications including 2G / 3G / 4G / 5G on the terminal 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.
[0117] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through an audio device (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.
[0118] The wireless communication module 160 can provide solutions for wireless communication applications on the terminal 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.
[0119] In some embodiments, antenna 1 of terminal 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling terminal 100 to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).
[0120] Terminal 100 implements display functions through a GPU, display screen 194, and application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information. Specifically, one or more GPUs in processor 110 can perform image rendering tasks (such as the rendering tasks related to the image to be displayed in this application), and pass the rendering results to the application processor or other display driver, which triggers the display screen 194 to display video.
[0121] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled LED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, terminal 100 may include one or N display screens 194, where N is a positive integer greater than 1. The display screen 194 can display the target video in the embodiments of this application. In one implementation, terminal 100 can run a camera-related application. When the camera-related application is opened on the terminal, display screen 194 can display a shooting interface, which may include a viewfinder, within which the target video can be displayed.
[0122] Terminal 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.
[0123] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, converting it into an image visible to the naked eye. The ISP can also perform algorithmic optimization on image noise and brightness. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.
[0124] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, terminal 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0125] The DSP converts the digital image signal into a standard RGB, YUV format image signal to obtain the original image (e.g., the first image and the second image in this embodiment). The processor 110 can perform further image processing on the original image. Image processing includes, but is not limited to, image stabilization, perspective distortion correction, optical distortion correction, and cropping to fit the size of the display screen 194. The processed image (e.g., the third image and the fourth image in this embodiment) can be displayed in the viewfinder of the shooting interface displayed on the display screen 194.
[0126] In this embodiment, the terminal 100 may have at least two cameras 193. For example, with two cameras, one is a front-facing camera and the other is a rear-facing camera; with three cameras, one is a front-facing camera and the other two are rear-facing cameras; with four cameras, one is a front-facing camera and the other three are rear-facing cameras. It should be noted that the camera 193 may be one or more of a wide-angle camera, a main camera, or a telephoto camera.
[0127] For example, taking two cameras, the front camera can be a wide-angle camera, and the rear camera can be the main camera. In this case, the image captured by the rear camera has a larger field of view and richer image information.
[0128] For example, taking three cameras as an example, the front camera can be a wide-angle camera, and the rear camera can be a wide-angle camera and a main camera.
[0129] For example, taking four cameras as an example, the front camera can be a wide-angle camera, and the rear camera can be a wide-angle camera, a main camera, and a telephoto camera.
[0130] A digital signal processor (DSP) is used to process digital signals. Besides digital image signals, it can also process other digital signals. For example, when terminal 100 selects a frequency point, the DSP can perform Fourier transforms on the frequency energy.
[0131] Video codecs are used to compress or decompress digital video. Terminal 100 may support one or more video codecs. Thus, terminal 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, etc.
[0132] NPU stands for Neural Network (NN) Computing Processor. By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs can enable intelligent cognitive applications in terminals, such as image recognition, facial recognition, speech recognition, and text understanding.
[0133] The external storage interface 120 can be used to connect an external storage card, such as a Micro SD card, to expand the storage capacity of the terminal 100. The external storage card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external storage card.
[0134] Internal memory 121 can be used to store computer executable program code, which includes instructions. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of terminal 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of terminal 100 by running instructions stored in internal memory 121 and / or instructions stored in memory located in the processor.
[0135] Terminal 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.
[0136] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.
[0137] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The terminal 100 can listen to music or make hands-free calls through the speaker 170A.
[0138] The receiver 170B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the terminal 100 receives a phone call or voice message, the receiver 170B can be brought close to the listener's ear to hear the voice.
[0139] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Terminal 100 may have at least one microphone 170C. In some embodiments, terminal 100 may have two microphones 170C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, terminal 100 may have three, four, or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.
[0140] The 170D headphone jack is used to connect wired headphones. The 170D headphone jack can be a USB 130 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.
[0141] Pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A can be disposed on display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive material. When force is applied to pressure sensor 180A, the capacitance between the electrodes changes. Terminal 100 determines the pressure intensity based on the change in capacitance. When a touch operation is applied to display screen 194, terminal 100 detects the intensity of the touch operation based on pressure sensor 180A. Terminal 100 can also calculate the touch position based on the detection signal from pressure sensor 180A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation commands. For example: when a touch operation with an intensity less than a first pressure threshold is applied to the SMS application icon, a command to view an SMS is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the SMS application icon, a command to create a new SMS is executed.
[0142] The shooting interface displayed on the screen 194 may include a first control and a second control. The first control is used to enable or disable the image stabilization function, and the second control is used to enable or disable the perspective distortion correction function. For example, a user can enable the image stabilization function by clicking the first control. The terminal 100 can determine the location of the first control based on the detection signal from the pressure sensor 180A, and then generate an operation command to enable the image stabilization function. Similarly, a user can enable the perspective distortion correction function by clicking the second control. The terminal 100 can determine the location of the second control based on the detection signal from the pressure sensor 180A, and then generate an operation command to enable the perspective distortion correction function.
[0143] The gyroscope sensor 180B can be used to determine the motion attitude of the terminal 100. In some embodiments, the gyroscope sensor 180B can determine the angular velocity of the terminal 100 around three axes (i.e., the x, y, and z axes). The gyroscope sensor 180B can be used for image stabilization. For example, when the shutter is pressed, the gyroscope sensor 180B detects the angle of the terminal 100's shake, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to counteract the shake of the terminal 100 through reverse movement, thus achieving image stabilization. The gyroscope sensor 180B can also be used in navigation and motion-sensing game scenarios.
[0144] The barometric pressure sensor 180C is used to measure air pressure. In some embodiments, the terminal 100 calculates altitude using the air pressure value measured by the barometric pressure sensor 180C to assist in positioning and navigation.
[0145] The magnetic sensor 180D includes a Hall sensor. The terminal 100 can use the magnetic sensor 180D to detect the opening and closing of the flip cover. In some embodiments, when the terminal 100 is a flip phone, the terminal 100 can detect the opening and closing of the flip cover using the magnetic sensor 180D. Then, based on the detected opening and closing state of the cover or the flip cover, features such as automatic flip unlocking can be set.
[0146] The 180E accelerometer can detect the magnitude of acceleration of terminal 100 in various directions (typically three axes). When terminal 100 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the posture of electronic devices, and is applied to applications such as screen orientation switching and pedometers.
[0147] A distance sensor 180F is used to measure distance. The terminal 100 can measure distance via infrared or laser. In some embodiments, during a shooting scene, the terminal 100 can utilize the distance sensor 180F to measure distance for rapid focusing.
[0148] The proximity sensor 180G may include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The LED may be an infrared LED. The terminal 100 emits infrared light outward through the LED. The terminal 100 uses the photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the terminal 100. When insufficient reflected light is detected, the terminal 100 can determine that there is no object near the terminal 100. The terminal 100 may use the proximity sensor 180G to detect when a user holds the terminal 100 close to their ear for a call, so as to automatically turn off the screen to save power. The proximity sensor 180G can also be used in holster mode and pocket mode for automatic unlocking and screen locking.
[0149] The ambient light sensor 180L is used to sense the ambient light intensity. The terminal 100 can adaptively adjust the brightness of the display screen 194 based on the sensed ambient light intensity. The ambient light sensor 180L can also be used to automatically adjust the white balance when taking pictures. The ambient light sensor 180L can also work with the proximity sensor 180G to detect whether the terminal 100 is in a pocket to prevent accidental touches.
[0150] The fingerprint sensor 180H is used to collect fingerprints. The terminal 100 can use the characteristics of the collected fingerprints to unlock the device, access application locks, take photos with fingerprints, and answer calls with fingerprints.
[0151] Temperature sensor 180J is used to detect temperature. In some embodiments, terminal 100 uses the temperature detected by temperature sensor 180J to execute a temperature processing strategy. For example, when the temperature reported by temperature sensor 180J exceeds a threshold, terminal 100 reduces the performance of the processor located near temperature sensor 180J to reduce power consumption and implement thermal protection. In other embodiments, when the temperature is below another threshold, terminal 100 heats battery 142 to prevent abnormal shutdown of terminal 100 due to low temperature. In still other embodiments, when the temperature is below yet another threshold, terminal 100 boosts the output voltage of battery 142 to prevent abnormal shutdown due to low temperature.
[0152] Touch sensor 180K, also known as a "touch device," can be located on display screen 194. The touch sensor 180K and display screen 194 together form a touchscreen, also known as a "touchscreen." Touch sensor 180K detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K may also be located on the surface of terminal 100, in a different position than display screen 194.
[0153] The bone conduction sensor 180M can acquire vibration signals. In some embodiments, the bone conduction sensor 180M can acquire vibration signals from the vibrating bone segments of the human vocal cords. The bone conduction sensor 180M can also contact the human pulse to receive blood pressure signals. In some embodiments, the bone conduction sensor 180M can also be incorporated into headphones to form bone conduction headphones. The audio module 170 can parse the voice signals from the vibrating bone segments of the vocal cords acquired by the bone conduction sensor 180M to realize voice functionality. The application processor can parse heart rate information from the blood pressure signals acquired by the bone conduction sensor 180M to realize heart rate detection functionality.
[0154] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Terminal 100 can receive button input and generate key signal inputs related to user settings and function control of terminal 100.
[0155] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. For example, different vibration feedback effects can correspond to touch operations performed on different applications (such as taking photos, playing audio, etc.). Motor 191 can also correspond to different vibration feedback effects for touch operations performed on different areas of the display screen 194. Different application scenarios (such as time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also be customized.
[0156] Indicator 192 can be an indicator light, used to indicate charging status, power changes, or to indicate messages, missed calls, notifications, etc.
[0157] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to make contact with and separate from the terminal 100. The terminal 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 simultaneously. The multiple cards can be of the same or different types. The SIM card interface 195 is also compatible with different types of SIM cards. The SIM card interface 195 is also compatible with external memory cards. The terminal 100 interacts with the network through the SIM card to realize functions such as calls and data communication. In some embodiments, the terminal 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the terminal 100 and cannot be separated from the terminal 100.
[0158] The software system of terminal 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses a layered Android system as an example to exemplify the software structure of terminal 100.
[0159] Figure 2 is a software structure block diagram of a terminal 100 according to an embodiment of this disclosure.
[0160] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.
[0161] The application layer can include a series of application packages.
[0162] As shown in Figure 2, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and SMS.
[0163] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.
[0164] As shown in Figure 2, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.
[0165] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.
[0166] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.
[0167] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.
[0168] The phone manager is used to provide communication functions for terminal 100. For example, it manages call status (including connection, hang-up, etc.).
[0169] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.
[0170] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.
[0171] The Android Runtime consists of core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.
[0172] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.
[0173] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0174] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.
[0175] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.
[0176] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.
[0177] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0178] A 2D graphics engine is a graphics engine for 2D drawing.
[0179] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.
[0180] The following example, using a photography scenario, illustrates the workflow of the terminal 100's software and hardware.
[0181] When touch sensor 180K receives a touch operation, a corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including touch coordinates, touch operation timestamp, etc.). The raw input event is stored in the kernel layer. The application framework layer obtains the raw input event from the kernel layer and identifies the control corresponding to the input event. Taking a touch click operation as an example, where the control corresponding to the click operation is the camera application icon, the camera application calls the interface of the application framework layer to start the camera application, and then calls the kernel layer to start the camera driver, capturing a still image or video through camera 193. The captured video can be the first image and the second image in this embodiment.
[0182] To facilitate understanding, an image processing method provided in this application embodiment will be specifically described in conjunction with the accompanying drawings and application scenarios.
[0183] Referring to Figure 3, which is a flowchart illustrating an image processing method according to an embodiment of this application, the method includes:
[0184] 301. Display a first image and a first marker on the image display page; the first marker is used to identify a first object in the first image.
[0185] The embodiments of this application can be applied to scenarios such as real-time photography and image post-processing, which will be described in detail below:
[0186] Example 1: Instant photo capture
[0187] In step 301, the first image can be a preview screen displayed on the image display interface. The image display page can be a camera capture page, and the first image can be displayed within the viewfinder.
[0188] In this embodiment of the application, in scenarios where the terminal takes photos or videos in real time, the terminal's camera can capture video streams in real time and display a preview image generated based on the video stream captured by the camera on the shooting interface. Specifically, the video captured by the terminal in real time is a sequence of original images arranged sequentially in the time domain. For example, the video captured by the terminal in real time includes frame 0, frame 1, frame 2, ..., frame X, arranged in the time domain, and the first image can be a preview image of frame n.
[0189] When a terminal takes a picture, it can open the shutter, allowing light to pass through the lens to the camera's photosensitive element. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to the image signal processor (ISP), digital signal processor (DSP), and image enhancement processing (e.g., noise reduction, image stabilization, super-resolution), resulting in an image. The first image described in this embodiment (and the second image described later) can be a preview image displayed on an image display page. It should be understood that the first image can also be a cropped image obtained after processing by the ISP and DSP, cropped to fit the size of the terminal's display screen or the viewfinder of the shooting interface.
[0190] The image processing method of this application embodiment can reconstruct a first image (or an image related to the first image, such as a second image) after the user triggers the shutter, so that the captured image is the reconstructed image.
[0191] Example 2: Image Post-processing
[0192] The first image can be an image stored locally on the electronic device, selected by the user as the first image. When viewing an image in the photo album, the user enters the photo editing interface by operating a specified editing control (which can be an in-interface control or a second / third-level hidden control). The previously viewed image becomes the first image, and the area of the photo editing interface displaying this image is the preview area. At this time, the image display page in this embodiment is the photo album's photo editing page.
[0193] In addition, in an image editing application, a picture (e.g., the first image) can be selected and opened. The selected picture is the first image, and the area of the picture displayed after the image editing application opens the picture is the preview area. In this case, the image display page in this embodiment is the page of the image editing application.
[0194] The image processing method of this application embodiment can reconstruct the first image after the user triggers the image reconstruction selection operation to obtain the reconstructed image.
[0195] In one possible implementation, a first marker may be displayed on the first image, the first marker being used to mark a first object in the first image, wherein the first object may be an object in the image (or a partial region of an object).
[0196] An image often needs to be composed around a main subject. In this embodiment, a first marker can be used to indicate which objects in the image can be used as the main subject of the composition. In addition, a preset position on the image display page (e.g., a second marker displayed at the preset position) can be used as a trigger option for selecting the main subject of the composition. Specifically, when the user moves the first marker on the image display page, so that the first marker and the preset position on the image display page are associated (e.g., the distance is less than a preset value, or the user has performed an association operation), it can be determined that the user's intention is to use the object indicated by the first marker as the main subject to assist in the composition of the current image.
[0197] The preset position can be the location of the center area. Specifically, the preset position can be a preset position on the image display page, or a preset position of the image display area in the image display page (e.g., the preset position of the first image). In addition, the preset position can be a fixed position, or a position that can be determined in real time based on preset rules. For example, when the display position of the first image changes, the preset position can move with the movement of the first image, for example, always remaining within the center area of the first image.
[0198] It should be understood that using the first object as the main subject of the composition can be interpreted as: composing a picture with the first object as the most eye-catching object in the picture. The first object is included in the image after composition. The composition can be based on the aesthetics of the image. For example, symmetry, color, structural integrity, and the aesthetics of the positional relationship between the first object and other objects will all affect the aesthetics of the image.
[0199] In this embodiment of the application, the composition may be achieved by cropping the image or by expanding the image content before cropping.
[0200] The first marker in the embodiments of this application will be introduced next.
[0201] In one possible implementation, the first image can be identified to identify objects contained within it. These objects are then filtered to identify potential compositional subjects. Finally, markers corresponding to these potential compositional subjects (including the first marker) are displayed on the image. Optionally, the number of first markers should not be excessive; for example, in this embodiment, the number of first markers is ≤5.
[0202] In this embodiment, the label recognition of the first image can be achieved, but is not limited to, through a large AI model. By pre-training the large model, it can identify several labels from objects contained in an image. Thus, the first image can be used as input to the large model to obtain at least one label, and the identified label can be displayed. The large AI model can be a locally configured large model, a cloud-configured large model, or a combination of an edge-side large model and a cloud-based large model to achieve label recognition.
[0203] Of course, other methods can also be used to recognize the markers, such as image recognition based on preset rules.
[0204] Generally, one object (i.e., a complete object in an image) corresponds to one label. However, there are exceptions. For example, if an object occupies a large area of the first image, then different reconstructions based on the object's location will result in significantly different images. Therefore, an object can also be labeled with multiple labels.
[0205] For example, if a building occupies most of the image in the first image, a marker can be displayed for the upper half of the building (e.g., a top corner) and a marker can be displayed for the lower half of the building (e.g., a bottom corner). When the upper half of the building (e.g., a top corner) is used as the main subject for recomposition, the resulting image includes the sky and the upper half of the building. When the lower half of the building (e.g., a bottom corner) is used as the main subject for recomposition, the resulting image includes the street and the lower half of the building.
[0206] In some embodiments, whether an object is tagged with multiple markers can be determined based on the distance between the markers. For example, if a large model identifies multiple markers on the same object, meaning that multiple markers can yield different reconstructed images with high aesthetic standards, then during marker generation, multiple markers are pre-generated, and then the distance between the multiple markers is used to determine whether to generate multiple markers. If the distance between two markers is less than a preset value, one of the markers is deleted; otherwise, if the distance between two markers is greater than or equal to the preset value, both markers are retained.
[0207] When the distance between the tags of two different objects is less than a preset value, the tags of the different objects are still retained.
[0208] Once a marker is identified, it is displayed on the first image. For example, an anchor point is added to the location where the marker is displayed on the first image, allowing the user to indicate that objects near that area are identified markers.
[0209] Figure 4 shows the markers (401 and 402) identified in a photo-taking scenario. The electronic device transmits the image data acquired by the graphics sensor to a large model (of course, in this embodiment, the object recognition process may not require processing by a large model; other models, rules, algorithms, etc., can be used for object recognition. The large model or other types of models / algorithms can be deployed on a cloud server or built into the device / application, etc.). The large model can then output multiple objects within the image, and the electronic device marks the object based on the object image. Alternatively, the large model can directly output the coordinates of the marked object or the coordinates of the displayed marker, and the electronic device directly displays the marker at the coordinates. In Figure 4, two markers (401 and 402) are identified, representing the sun and a leopard, respectively.
[0210] The aforementioned markers can be automatically recognized by electronic devices, and users can update the markers if they do not meet their requirements.
[0211] In one possible implementation, a third marker may be displayed on the image display page to identify a second object in the first image; the first object and the second object are different; a user may perform a second operation on the electronic device, and the electronic device may, in response to the received second operation, delete the display of the third marker and add the display of the first marker, the second operation indicating an update of the markers for the objects in the first image.
[0212] For example, users can refresh the markers through preset operations, that is, re-identify the markers through preset operations. When identifying markers a second time, the markers identified the first time will be discarded and will no longer be used as markers. Taking Figure 4 as an example, if the preset marker refresh operation is performed, the sun and the leopard will no longer be used as markers, but multiple trees and two of the grass on the ground will be used as markers.
[0213] For example, users can manually select markers: users can manually add markers through specified interactions. For instance, a user can long-press near the display area of an object in the first image and then add a marker. Another example is that a user can directly drag an existing marker to change its position on the first image; after releasing the drag gesture, the new marker will be in its current location. Alternatively, after releasing the gesture, the marker will automatically move to the object closest to the released position, and that object will become the marker.
[0214] Taking Figure 4 as an example, if the user drags the marker indicating the location of the sun onto one of the trees, that tree will become the new marker, and the sun will no longer be a marker. As shown in Figure 5, the moved markers become the nearest tree (403) and the leopard (401).
[0215] Optionally, after identifying a marker, preset visual effects can be added to it to alert the user that the marker has been recognized. For example, a halo can be added around the marker's anchor point to prompt the user. The halo flashes at a preset frequency to reduce interference from the reminder.
[0216] The preset positions in the embodiments of this application will be described next.
[0217] In this embodiment of the application, a preset position on the image display page (e.g., a second mark displayed at the preset position) can be used as a trigger option for selecting the composition subject.
[0218] For example, the preset position can be, but is not limited to, the center of the image display area.
[0219] In one possible implementation, a second marker is displayed at a preset position on the image display page.
[0220] In one possible implementation, the second marker is displayed automatically upon entering the image display page, or the second marker is displayed in response to an operation on the image display page.
[0221] The second marker is a visual identifier displayed on the screen to assist in composition. It is generally used to indicate the (geometric) center position of the preview area. Initially, the second marker is displayed in the center position of the first image.
[0222] Optionally, there are two ways to trigger the second marker's interaction:
[0223] 1. Upon entering the first interface, the second marker is displayed (that is, it is automatically displayed after entering the image display page); 2. After entering the first interface, the second marker is invoked through a specified interaction method (that is, it is invoked by an operation on the image display page).
[0224] For the first interaction method, the second marker is regarded as a display element of the first interface, and the interaction of entering the first interface is reused as the interaction of calling up the second marker.
[0225] For example, in a photo-taking scenario, the second marker is displayed once the photo / video interface is entered; or, the second marker is only displayed when switching to a specific photo-taking mode, and not in normal photo-taking mode. In a photo-editing scenario, the second marker is simultaneously displayed when the space is clicked to enter the photo-editing interface. For photo-editing applications, the second marker is displayed in the preview area before the first image is selected.
[0226] Figure 6 illustrates the second marker (601) in three different scenarios, from left to right: shooting scenario, album editing scenario, and album editing scenario. In the shooting scenario, when switching to a specific shooting mode, the second marker, i.e., the "+" symbol in Figure 6, is added to the image preview area. In the album application, the second marker is displayed when entering the editing interface by clicking on spaces such as "edit" to edit the image. In the editing application, the second marker is displayed in the preview area even before the first image is selected.
[0227] As shown in Figure 6, since the preview area is used to display the first image, the initial display position of the second mark is generally located at the geometric center of the first image.
[0228] For the second interaction method, the auxiliary composition function of this application embodiment can be activated by a preset operation to display the second marker. The preset operation can be, for example, interaction with a physical button / pressure-sensitive button / touch button on the side frame, voice wake-up, preset gesture, or clicking a preset control, or a combination thereof. For example, pressing / double-pressing the power button activates the auxiliary composition function and displays the second marker on the interface. As another example, a pressure-sensitive button is provided on the side frame; after interacting with the pressure-sensitive button, several controls are displayed on the screen, one of which is an auxiliary composition control (or another name); clicking this control displays the second marker.
[0229] In this embodiment, a first mark indicating a first object is displayed on the first image. If the user needs to recompose the first image around the first object, a positional association can be established between the first mark and a preset position on the image display interface, thereby triggering the electronic device to recompose the first image around the first object. In a photography scenario, the image can be reconstructed based on the first mark selected by the user. Thus, when the user takes a picture, the captured image is the reconstructed image. The reconstructed image is a graphic with high aesthetic standards; therefore, when the user uses the assisted imaging method of this embodiment, they can consistently obtain high-quality photos. Similarly, during image editing, rapid image editing can be achieved, reducing the time required for editing.
[0230] In one possible implementation, a user can perform a first operation, and the electronic device can respond to the received first operation by moving the first marker until a positional association exists between the first marker and the preset position. The first operation can be a movement of the first image on the image display page or a movement of the camera that captures the first image.
[0231] In this step, different scenarios require different processing methods. In the shooting scenario, the electronic device acquires the first image in real time, so the first image is constantly changing, while the position of the first marker remains basically unchanged (this is because the first marker indicates an object in the image; if the object's position relative to the camera in the physical world remains unchanged, then the object's position in the image remains unchanged, and consequently, the position of the first marker in the image also remains unchanged; conversely, if the object's position relative to the camera in the physical world changes, then the object's position in the image changes, and consequently, the position of the first marker in the image also changes). Only when the position or posture of the electronic device changes can the relative position between the preset position and the first marker on the image display page be changed. In the retouching scenario, since the first image is a fixed image, only by changing the display position of the first image on the screen can the relative position between the preset position and the first marker on the image display page be changed. The shooting scenario and the retouching scenario are described below respectively.
[0232] 1) Shooting location:
[0233] Figure 7 shows the change in the relative position between the preset position and the first marker after the phone is moved. After the phone moves to the right, the image captured by the image sensor changes; the first marker remains the sun and the leopard, while the preset position remains in the center of the area.
[0234] In Figure 7, since the preview area displays an image captured in real time by the image sensor, which is constantly being updated, the content displayed in the preview area can be considered a video stream. However, at a specific moment, the preview area displays a specific frame of graphics, such as the static image in Figure 7. In this image, the relative positions of the preset position (601) and the first markers (401 and 402) remain unchanged. Therefore, in the shooting scenario, changing the relative position of the first marker and the preset position essentially means changing the position of the first marker and the first image in the next frame by moving the camera. Since the image captured by the electronic device does not change much during continuous movement, the first marker usually remains unchanged. Therefore, the preset position can be moved closer to the first marker by moving the electronic device.
[0235] Of course, when the movement distance / posture changes to a certain extent, causing the newly captured image to differ significantly from the image captured in the previous period, the first marker for recognition may also change.
[0236] Figure 8 illustrates the recognition of the first marker after continuous changes in the first image. One approach is to keep the number of first markers unchanged, assign the new object as the first marker, and make the original first markers disappear (402 disappears, and marker 404 is added). Another approach is to increase the number of first markers (402 does not disappear, and marker 404 is added).
[0237] When the first marker changes, the following scenario may occur: In the previous moment, an object was identified as the first marker in the first image at that moment, but in the first image at the next moment, the object is no longer identified as the first marker. This is the disappearance of an existing first marker. If the disappeared first marker happens to be the first marker that the user expected to select, the above-mentioned first marker update operation may need to be performed, thereby degrading the user experience.
[0238] In one possible implementation, a user can input a third operation to an electronic device, which can receive the third operation indicating the fixing of the first mark; upon receiving a fourth operation, the display of the first mark is maintained; the fourth operation is either movement of the first image on the image display page or movement of the camera capturing the first image.
[0239] Figure 9 illustrates the first marker recognition before and after the device is moved. Before moving, long-pressing the first marker (401) where the leopard is located will bring up a fixed control (901). Clicking the pin (901) indicates that the object corresponding to the first marker is fixedly recognized as the first marker (401). Therefore, after moving, the object (leopard) is still recognized as the first marker (401).
[0240] It should be noted that fixing the first marker does not mean fixing its position on the screen, but rather that the object corresponding to the first marker is fixedly identified as the first marker. Therefore, in Figure 6, the position of the first marker corresponding to the leopard object has changed so that the position of the first marker on the screen matches the position of the leopard in the first image.
[0241] 2) Image editing scenarios:
[0242] In image editing scenarios, the first image is often fixed, and users can directly drag it within the screen to change the position of the first marker. Alternatively, the first image can be zoomed in first, and then dragged to move it.
[0243] Figure 10 shows the screen before and after moving the image. By moving the image, the first marker is aligned with a preset position. In image editing scenarios, the first image is sized to fit the preview area, preventing it from moving. In this case, the first image can be moved by manually zooming in. Often, part of the first image is off-screen, as shown in the right-hand view of Figure 10.
[0244] In the foregoing description, the second marker displayed at the preset position is often immovable and fixed in the center of the preview area. This is because, in the shooting scenario (the main scenario of this embodiment), when the first marker is located in the middle of the preview area, more content near the object corresponding to the first marker can be captured, which makes it easier to recompose the image around the object corresponding to the first marker.
[0245] Optionally, the second marker displayed at the preset position can also be designed to be movable. By moving the second marker displayed at the preset position, the relative position of the second marker displayed at the preset position and the first marker can be changed. This facilitates recomposition in photo editing scenarios and is also applicable to shooting scenarios.
[0246] 302. When the distance between the first mark and the preset position of the image display page is less than a preset value, a third image is obtained based on the first image or the second image; the second image and the first image are images captured by different cameras (e.g., cameras with different magnification) of the same scene, and the third image is an image obtained by reconstructing the first image or the second image with the first object as the main subject of the composition;
[0247] In one possible implementation, the existence of location association includes: the distance between the first marker and the preset location is less than a preset value.
[0248] For example, when the preset position overlaps with the first mark, a composition calculation can be performed based on the first mark to obtain a reconstructed image (third image), which is then displayed in the preview area; wherein, the reconstructed image is obtained by the first image or the second image through preset processing.
[0249] By moving the first marker, a preset position can be made to overlap with (or be closer to) the first marker, indicating that the first marker is selected, and the composition can be made around the object corresponding to the first marker as a reference object.
[0250] 1) Overlap between the preset position and the first mark
[0251] The overlap mentioned here can mean that the center of the first mark and the center of the preset position (optionally, the second mark displayed at the preset position) are the same point, or it can mean that the distance between the geometric center of the first mark and the geometric center of the preset position is less than a preset value.
[0252] When the distance between the geometric center of the first mark and the geometric center of the second mark displayed at the preset position is less than a preset value, the first mark is first enlarged until it at least partially surrounds the second mark displayed at the preset position, until it completely surrounds the second mark displayed at the preset position. After the second mark is surrounded, it disappears.
[0253] Figure 11 illustrates the fusion process of the first marker (401) and the second marker 901 displayed at a preset position. Initially, the second marker (901) displayed at the preset position does not intersect with the hot zone of the first marker (the dashed circle 1101 outside the first marker). When the center point of the second marker displayed at the preset position, i.e., the center point of the cross, enters the hot zone, the circle of the first marker enlarges (1102), and the enlargement direction is towards the second marker displayed at the preset position, causing the center of the first marker to move closer to the center of the preset position. When the distance from the center of the first marker to the center of the second marker displayed at the preset position is less than a preset value, the second marker displayed at the preset position shrinks and is completely contained within the first marker. At this time, the second marker displayed at the preset position disappears, and only the circle of the first marker is displayed on the interface. When the position of the second marker displayed at the preset position relative to the first marker changes (by moving the phone / changing the phone's posture), if the change in movement distance / posture is small, the enlarged circle of the first marker moves around the area defined by the hot zone. If the movement distance / attitude changes too much, the first marker returns to its original size, and the second marker detaches from the circle of the first marker.
[0254] Figure 12 shows the interface after the first mark overlaps with the preset position. In Figure 12, the user selects the leopard as the target object (601 and 401 overlap) and recomposes the image based on the leopard.
[0255] In one possible implementation, a third image can be obtained based on the first image or the second image; the second image and the first image are images of the same scene captured by different cameras, and the third image is an image obtained by reconstructing the first image or the second image with the first object as the main subject of the composition.
[0256] Next, we will introduce the second image:
[0257] In this embodiment, when recomposing the image around the first object, the image captured by other cameras on the electronic device can be processed as a (second image) instead of the image displayed in the current preview stream (first image). In this case, multiple cameras on the electronic device are simultaneously operational. Of course, to ensure processing accuracy, the second image and the first image can be images captured by cameras with different magnifications on the electronic device at the same timestamp for the same scene.
[0258] If an electronic device integrates multiple cameras, each camera has its own inherent magnification. The inherent magnification of a camera can be understood as the magnification factor of the camera lens, which is also the imaging angle. This parameter determines the field of view of the camera when shooting, that is, the imaging range.
[0259] In some embodiments, the user adjusts the magnification of the preview image within the viewfinder in the shooting interface. Subsequently, the electronic device can perform digital zoom on the image captured by the camera. That is, the ISP or other processor of the electronic device enlarges the area of each pixel of the image captured by the camera at its inherent magnification and correspondingly reduces the framing range of the image, so that the processed image presents an image equivalent to an image captured by the camera at other shooting magnifications. The shooting magnification of the electronic device in this application can be understood as the shooting magnification of the image presented by the above-mentioned processed image.
[0260] In one implementation, the electronic device may integrate one or more of a front-facing short-focus (wide-angle) camera, a rear-facing short-focus (wide-angle) camera, a rear-facing medium-focus camera, and a rear-facing telephoto camera.
[0261] Typically, users utilize the mid-range telephoto camera most frequently; therefore, it is generally set as the main camera. The main camera's focal length is set to the reference focal length, and its inherent magnification is usually "1×". In some embodiments, digital zoom (or digital zoom) can be applied to the image captured by the main camera. That is, the ISP or other processor in the phone enlarges the area of each pixel in the "1×" image captured by the main camera and correspondingly reduces the framing of the image, so that the processed image is equivalent to an image captured by the main camera at other shooting magnifications (e.g., "2×"). In other words, images captured using the main camera can correspond to a shooting magnification range, such as "1×" to "5×". It should be understood that when an electronic device also integrates a rear telephoto camera, the magnification range corresponding to the image captured by the main camera can be 1 to b. The maximum magnification value b in the magnification range can be the inherent magnification of the rear telephoto camera. The value of b can be 3 to 15. For example, b can be, but is not limited to, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15.
[0262] The first image can be an image captured by the rear short-focus (wide-angle) camera, and the second image can be an image captured by the rear medium-focus camera.
[0263] The following section explains how to obtain the third image based on the first or second image.
[0264] In one possible implementation, the first object can be used as the main subject of the composition, and the first image or the second image can be reconstructed based on the aesthetics of the composition.
[0265] Once a positional association is established between the preset position and the first marker, it means that the user expects to use the object corresponding to the first marker for composition. The object corresponding to the selected first marker in the first image is called the first object.
[0266] After identifying the first object, the electronic device analyzes the first image or the second image (for ease of description, the first image is used as an example in this embodiment) according to a preset algorithm to calculate a reconstructed image (third image). That is, the reconstructed image (third image) is obtained based on the first image. In this embodiment, the reconstructed image (third image) can be obtained by cropping the first image; that is, the third image is a part of the first image.
[0267] The analysis of the first image can be performed by a large AI model or according to preset rules. The large AI model can be an edge-side model, a cloud-based model, or a collaborative effort between large models on both the edge and cloud sides. Specifically, the first image is used as input, and aesthetic element analysis is performed around the first object. The final output is a reconstructed image that includes the first object. This reconstructed image has a high aesthetic score or conforms to certain aesthetic elements in the aesthetic analysis dimension.
[0268] In one possible implementation, the reconstruction includes image cropping, wherein the cropped image ratio (e.g., aspect ratio) is consistent with the first image, or determined based on ratio information input by the user.
[0269] The aspect ratio of the reconstructed image (third image) boundary can be consistent with the image specifications set by the user. For example, in a camera application, if the user sets the image specifications to 4:3, then the aspect ratio of the reconstructed image boundary is also 4:3. When the user adjusts the image specifications (the user inputs the aspect ratio information), the aspect ratio of the reconstructed image boundary changes accordingly. Typically, the aspect ratio of the preview area is consistent with the image specifications set by the user; therefore, it can also be said that the aspect ratio of the reconstructed image boundary can be consistent with the aspect ratio of the preview area.
[0270] In one possible implementation, the third image is obtained by cropping the first image, and the first image and an indicator box of the cropped area of the third image in the first image can be displayed on the image display page.
[0271] The indicator box can represent the boundary of the image. Figure 13 shows the subsequent process of Figure 12. After the preset position overlaps with the first mark corresponding to the leopard, the boundary of the reconstructed image is determined by analyzing the first image (1301), and the boundary of the reconstructed image is displayed in the preview area. During the process of displaying the boundary of the reconstructed image, the circle used to identify the first mark gradually disappears. In the shooting scenario, the first image is updated in real time; therefore, the reconstructed image is also updated in real time. Similarly, the boundary of the reconstructed image is also updated in real time.
[0272] As can be seen from Figure 13, in the reconstructed image, the object corresponding to the first marker is not necessarily located at the center of the reconstructed image. Instead, it is referenced to objects (trees, the sun, etc.) near the reconstructed image. After analysis, a reconstructed image that meets aesthetic standards is obtained, and the boundaries of image cropping are determined.
[0273] In Figure 13, if the user moves the electronic device during the shooting scene, the final reconstructed image may be the same or different.
[0274] In one possible implementation, the reconstruction includes: generating content expansion of the first image and cropping the expanded first image.
[0275] In image retouching scenarios, when the first marker is located at the edge of the image, the lack of additional visual information beyond the edge may affect image reconstruction. Therefore, before determining the reconstructed image, the first image can be expanded to generate additional content. For example, AI can be used to expand the image content and supplement the visual information beyond the edge of the first image. Then, based on the AI-redrawn first image, a reconstructed image can be generated, the boundaries of the reconstructed image can be marked, and the final reconstructed image can be obtained.
[0276] The extended generation of image content can have certain triggering conditions. For example, the extended generation of image content is only triggered when the first mark, which is close to a preset position, is located at the edge of the first image. Optionally, when extending the image content, only the edge pattern on one side of the first mark can be extended and drawn.
[0277] 303. Display the third image on the image display page.
[0278] After obtaining the third image, it can be displayed on the image display page. For example, the third image can be displayed in the preview area of the image display page, and the first image will no longer be displayed in the preview area.
[0279] Optionally, the boundary of the third image can be displayed on the first image, and then the boundary of the third image can be expanded outward to enlarge the third image so that the third image is displayed in the preview area.
[0280] The operation of enlarging the third image in the preview area can be done automatically. That is, after recognizing the boundaries of the third image, the first image is processed and automatically enlarged for display in the preview area. Alternatively, the third image can be enlarged for display after detecting a specified user interaction. For example, if the user clicks on the border of the third image, it will be enlarged for display in the preview area. Another example is the operation of other virtual controls or physical buttons.
[0281] Figure 14 shows the shooting interface after the boundary of the third image in Figure 13 is magnified and displayed in the preview area. After magnification, content in the original first image that is outside the boundary of the third image is no longer displayed. It should be noted that in the shooting scenario, the first image is refreshed in real time, but as long as the user selects the first marker (overlapping the preset position with the first marker), subsequent new first images can still recognize the same first marker object, and the first image will still be reconstructed. As long as the user does not click to take a picture, the newly received first image will continue to be reconstructed based on the same first marker object, and the third image will be displayed in the preview area.
[0282] As mentioned earlier, in a photo-taking scenario, the first image is updated in real time, and the third image may also change. When the third image is displayed, the user can move the electronic device, and a vibration alert is given as the preset position changes relative to the initial position. The greater the positional shift, the stronger the vibration alert. When the shift exceeds a preset value (for example, the movement causes the object corresponding to the current first marker to not be fully displayed in the first image), the third image is exited, a new first marker is identified based on the new first image, and a new third image is displayed after the user selects the new first marker.
[0283] After zooming in, the image can also respond to user input and re-display the first image in the preview area. This can be done in two design ways:
[0284] In one possible implementation, a fifth operation may also be received, which instructs switching to display the image before reconstruction; in response to the fifth operation, the third image on the image display page is switched to the first image or the second image.
[0285] In one possible implementation, the image display interface displays a return control, and the fifth operation is an operation performed on the return control. After performing a preset operation on the return control, the image reconstruction is deactivated, the display of the third image is exited, and the first image or the second image is re-displayed.
[0286] The fifth operation can be a return / exit operation, specifically, an operation performed on the return control. For example, in one implementation, the user can exit the display of the third image using the return / exit operation. When exiting the display of the third image, the preview area no longer displays the third image, but instead displays the first or second image (of course, in a shooting scenario, since the first image is refreshed in real time, the first image may not be the same as the image before zooming in). After displaying the first image, the first image is re-identified and marked. For example, in a shooting scenario, a zoom control is usually displayed below the preview area. A return control can be displayed near the zoom control. After the user clicks the return control, the display of the third image is exited, and the preview area re-displays the first image. In Figure 14, a return control (1401) is displayed below the preview area.
[0287] In one possible implementation, the image display interface displays a switching control, and the fifth operation is a preset operation performed on the switching control. In addition, a sixth operation can be received, in which the image display interface no longer displays the first image or the second image, and re-displays the reconstructed third image. The sixth operation includes: performing a preset operation on the switching control and releasing the fifth operation.
[0288] Users can switch between the third and first images using a switching operation. Furthermore, a fifth operation allows users to switch back to the first image from the third image on the display screen. This fifth operation can be performed by clicking or long-pressing the switching control. After performing the fifth operation, the third image on the display screen will switch to the first image. Users can then perform a sixth operation to switch back to the third image. This sixth operation can specifically be a preset operation performed on the switching control (e.g., clicking the switching control again or releasing the switch).
[0289] For example, replacing the return control in Figure 15 with a toggle control, clicking the toggle control or holding down the return control changes the preview area from displaying the third image in Figure 14 to displaying the first image. Clicking the toggle control again or releasing the return control restores the preview area to displaying the third image. The difference between the toggle and return operations is that the toggle operation is for comparison; therefore, clicking the toggle operation does not change the first marker object, and there is no need to reselect the first marker. Thus, the image reconstruction (reconstructing the new first image) can be retried using the toggle control without reselecting the first marker. However, with the return operation, if the third image needs to be redisplayed, the positional association between the first marker and the preset position needs to be re-established.
[0290] Figure 15 shows the preview area before and after the switch, and the different shapes of the switching control indicate the graphics currently displayed in the preview area. The switching control includes an outer component (1501) and an inner component (1502). As shown in Figure 15, the switching control includes an outer circle and an inner square. The outer circle indicates the first image, and the inner square indicates the third image. When the preview area displays the third image, the inner square is highlighted; when the preview area displays the first image, the inner square is no longer highlighted, but maintains the same color / hue as the outer circle. By clicking or long-pressing the switching control, the preview area can be switched from displaying the third image to displaying the first image. Clicking again or releasing the long press restores the display of the third image in the preview area.
[0291] In addition, a third image can be saved.
[0292] For example, in a shooting scenario, the image generation operation is a shooting operation, such as clicking the photo / video control. In an image editing scenario, it could be clicking the save control.
[0293] Taking the aforementioned shooting scenario as an example, after clicking to take a photo, the electronic device will not capture the first image, but will directly capture the third image and save it to the album. In this way, high-quality photos can be obtained without the user performing secondary processing.
[0294] This application also provides an image processing apparatus, which can be a terminal device. Referring to FIG16, FIG16 is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of this application. As shown in FIG16, the image processing apparatus 1600 includes:
[0295] Display module 1601 is used to display a first image and a first mark on an image display page; the first mark is used to identify a first object in the first image;
[0296] For a detailed description of the display module 1601, please refer to the descriptions in steps 301 and 303, which will not be repeated here.
[0297] The image reconstruction module 1602 is used to obtain a third image based on the first image or the second image when the distance between the first mark and the preset position of the image display page is less than a preset value; the third image is an image obtained by reconstructing the first image or the second image with the first object as the composition subject;
[0298] For a detailed description of the image reconstruction module 1602, please refer to the description in step 302, which will not be repeated here.
[0299] The display module is used to display the third image on the image display page.
[0300] In one possible implementation, a second marker is displayed at a preset position on the image display page.
[0301] In one possible implementation, the second image and the first image are images of the same scene captured by different cameras.
[0302] In one possible implementation, the refactoring includes:
[0303] Perform image cropping; or,
[0304] The content of the first image is expanded and generated, and the expanded first image is then cropped.
[0305] In one possible implementation, the third image specifically refers to:
[0306] The image is obtained by reconstructing the first image or the second image based on the aesthetics of the composition, using the first object as the main subject.
[0307] In one possible implementation, the existence of location association includes:
[0308] The distance between the first marker and the preset position is less than a preset value.
[0309] In one possible implementation, the display module 1601 is further configured to:
[0310] In response to a received first operation, the first marker is moved until a positional association exists between the first marker and the preset position; the first operation is a movement of the first image on the image display page or a movement of the camera that captures the first image.
[0311] In one possible implementation, the second marker is displayed automatically upon entering the image display page, or the second marker is displayed in response to an operation on the image display page.
[0312] In one possible implementation, the display module 1601 is specifically used for:
[0313] A first image and a third marker are displayed on the image display page, the third marker being used to identify a second object in the first image; the first object and the second object are different.
[0314] In response to the received second operation, the display of the third marker is deleted and the display of the first marker is added, the second operation indicating an update of the markers for objects in the first image.
[0315] In one possible implementation, the display module 1601 is further configured to:
[0316] Upon receiving a third operation, the third operation instructs the fixing of the first mark;
[0317] Upon receiving a fourth operation, the display of the first marker is maintained; the fourth operation is either movement of the first image on the image display page or movement of the camera that captures the first image.
[0318] In one possible implementation, the reconstruction includes image cropping, the cropping ratio being consistent with the first image, or determined based on ratio information input by the user.
[0319] In one possible implementation, the display module 1601 is further configured to:
[0320] When the distance between the first marker and the preset position on the image display page is less than a preset value, a fusion effect between the first marker and the second marker is displayed.
[0321] In one possible implementation, the display module 1601 is further configured to:
[0322] Upon receiving the fifth operation, the fifth operation indicates switching to displaying the image before reconstruction;
[0323] In response to the fifth operation, the third image on the image display page is switched to the first image or the second image.
[0324] In one possible implementation, the image display interface displays a return control, and the fifth operation is an operation performed on the return control. After performing a preset operation on the return control, the image reconstruction is deactivated, the display of the third image is exited, and the first image or the second image is re-displayed.
[0325] In one possible implementation, the image display interface displays a switching control, the fifth operation is a preset operation performed on the switching control, and the display module 1601 is further configured to:
[0326] Upon receiving the sixth operation, the image display interface no longer displays the first or second image, but instead displays the reconstructed third image; wherein, the sixth operation includes: performing a preset operation on the switching control and releasing the fifth operation.
[0327] In one possible implementation, the third image is obtained by cropping the first image, and the display module 1601 is specifically used for:
[0328] The first image and an indicator box showing the cropping area of the third image within the first image are displayed on the image display page.
[0329] In one possible implementation, the preset position is located in the central region of the first image.
[0330] In one possible implementation, the image reconstruction module 1602 is specifically used for:
[0331] Using the first object as the main subject of the composition, and based on the aesthetics of the composition, the first image or the second image is reconstructed using image mapping rules or a large model to obtain the third image.
[0332] The following describes a terminal device provided in an embodiment of this application. The terminal device can be the image processing device shown in Figure 16. Please refer to Figure 17, which is a structural schematic diagram of a terminal device provided in an embodiment of this application. The terminal device 1700 can specifically be a virtual reality (VR) device, a mobile phone, a tablet, a laptop computer, a smart wearable device, etc., and is not limited here. Specifically, the terminal device 1700 includes: a receiver 1701, a transmitter 1702, a processor 1703, and a memory 1704 (the number of processors 1703 in the terminal device 1700 can be one or more; Figure 17 shows one processor as an example). The processor 1703 may include an application processor 17031 and a communication processor 17032. In some embodiments of this application, the receiver 1701, transmitter 1702, processor 1703, and memory 1704 can be connected via a bus or other means.
[0333] Memory 1704 may include read-only memory and random access memory, and provides instructions and data to processor 1703. A portion of memory 1704 may also include non-volatile random access memory (NVRAM). Memory 1704 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations.
[0334] Processor 1703 controls the operation of the terminal device. In specific applications, the various components of the terminal device are coupled together through a bus system. This bus system includes not only the data bus but also power buses, control buses, and status signal buses. However, for clarity, all buses are referred to as the bus system in the diagram.
[0335] The methods disclosed in the embodiments of this application can be applied to or implemented by processor 1703. Processor 1703 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above methods can be completed by the integrated logic circuits in the hardware of processor 1703 or by instructions in software form. Processor 1703 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and may further include application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. Processor 1703 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 1704. Processor 1703 reads the information in memory 1704 and, in conjunction with its hardware, completes the steps of the above method. Specifically, processor 1703 can read the information in memory 1704 and, in conjunction with its hardware, complete the data processing-related steps 301 to 303 in the above embodiments.
[0336] Receiver 1701 can be used to receive input digital or character information, and to generate signal inputs related to the settings and function control of the terminal device. Transmitter 1702 can be used to output digital or character information through the first interface; transmitter 1702 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; transmitter 1702 may also include a display device such as a display screen.
[0337] This application also provides a computer program product that, when run on a computer, causes the computer to perform the steps of the image processing method described in the embodiment corresponding to FIG3 above.
[0338] This application also provides a computer-readable storage medium storing a program for performing signal processing, which, when run on a computer, causes the computer to execute the steps of the image processing method as described in the foregoing embodiments.
[0339] The image display device provided in this application embodiment can specifically be a chip, which includes a processing unit and a communication unit. The processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, pins, or circuits. The processing unit can execute computer execution instructions stored in the storage unit to cause the chip in the execution device to execute the image processing method described in the above embodiments, or to cause the chip in the training device to execute the image processing method described in the above embodiments. Optionally, the storage unit can be a storage unit within the chip, such as a register or cache. Alternatively, the storage unit can be a storage unit located outside the chip within the wireless access device, such as a read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions, such as random access memory (RAM).
[0340] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0341] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0342] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0343] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
Claims
1. An image processing method, characterized in that, The method includes: A first image and a first marker are displayed on the image display page; the first marker is used to identify a first object in the first image. When the distance between the first mark and the preset position of the image display page is less than a preset value, a third image is obtained based on the first image or the second image; the third image is an image obtained by reconstructing the first image or the second image with the first object as the main subject of the composition; The third image is displayed on the image display page.
2. The method according to claim 1, characterized in that, A second mark is displayed at a preset position on the image display page.
3. The method according to claim 1 or 2, characterized in that, The reconstruction includes: Perform image cropping; or, The content of the first image is expanded and generated, and the expanded first image is then cropped.
4. The method according to any one of claims 1 to 3, characterized in that, The third image is specifically: The image is obtained by reconstructing the first image or the second image based on the aesthetics of the composition, using the first object as the main subject.
5. The method according to any one of claims 1 to 4, characterized in that, The second image and the first image are images of the same scene captured by different cameras.
6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: In response to a received first operation, the first marker is moved until a positional association exists between the first marker and the preset position; the first operation is a movement of the first image on the image display page, or a movement of the camera that captures the first image.
7. The method according to claim 2, characterized in that, The second marker is displayed automatically after entering the image display page, or the second marker is activated in response to an operation on the image display page.
8. The method according to any one of claims 1 to 7, characterized in that, The step of displaying the first image and the first marker on the image display page includes: A first image and a third marker are displayed on the image display page, the third marker being used to identify a second object in the first image; the first object and the second object are different. In response to the received second operation, the display of the third marker is deleted and the display of the first marker is added, the second operation indicating an update of the markers for objects in the first image.
9. The method according to any one of claims 1 to 8, characterized in that, The method further includes: Upon receiving a third operation, the third operation instructs the fixing of the first mark; Upon receiving a fourth operation, the display of the first marker is maintained; the fourth operation is either movement of the first image on the image display page or movement of the camera that captures the first image.
10. The method according to any one of claims 1 to 9, characterized in that, The reconstruction includes image cropping, wherein the cropping ratio is consistent with the first image, or determined based on the ratio information input by the user.
11. The method according to claim 2, characterized in that, The method further includes: When the distance between the geometric center of the first mark and the geometric center of the second mark displayed at the preset position is less than a preset value, the first mark is enlarged, and the enlarged first mark at least partially surrounds the second mark within the first mark, until it completely surrounds the second mark.
12. The method according to any one of claims 1 to 11, characterized in that, The method further includes: Upon receiving the fifth operation, the fifth operation indicates switching to displaying the image before reconstruction; In response to the fifth operation, the third image on the image display page is switched to the first image or the second image.
13. The method according to claim 12, characterized in that, The image display interface displays a return control. The fifth operation is an operation performed on the return control. After performing a preset operation on the return control, the image reconstruction is deactivated, the display of the third image is exited, and the first image or the second image is re-displayed.
14. The method according to claim 12, characterized in that, The image display interface displays a switching control, and the fifth operation is a preset operation performed on the switching control. The method further includes: Upon receiving the sixth operation, the image display interface no longer displays the first or second image, but instead displays the reconstructed third image; wherein, the sixth operation includes: performing a preset operation on the switching control and releasing the fifth operation.
15. The method according to any one of claims 1 to 14, characterized in that, The third image is obtained by cropping the first image, and displaying the third image on the image display page includes: The first image and an indicator box showing the cropping area of the third image within the first image are displayed on the image display page.
16. The method according to any one of claims 1 to 15, characterized in that, The preset position is located in the central region of the first image.
17. The method according to any one of claims 1 to 16, characterized in that, The step of obtaining the third image based on the first image or the second image includes: Using the first object as the main subject of the composition, and based on the aesthetics of the composition, the first image or the second image is reconstructed using image mapping rules or a large model to obtain the third image.
18. The method according to any one of claims 1 to 16, characterized in that, Displaying the first image and the first marker on the image display page includes: displaying the first image and at least two markers on the image display page; Each of the markers is used to identify an object in the first image, and the first marker is any one of the at least two markers, which are obtained based on the recognition of the first image.
19. An image processing apparatus, characterized in that, It includes one or more modules for performing the method as described in any one of claims 1 to 18.
20. A computer storage medium, characterized in that, The computer storage medium stores one or more instructions that, when executed by one or more computers, cause the one or more computers to perform the method as described in any one of claims 1 to 18.
21. A computer program product, characterized in that, Includes computer-readable instructions that, when executed on a computer device, cause the computer device to perform the method as described in any one of claims 1 to 18.
22. An electronic device, characterized in that, Includes at least one processor and at least one memory; The at least one memory is used to store code; The at least one processor is configured to execute the code to cause the electronic device to perform the method as described in any one of claims 1 to 18.