A method and electronic device for eliminating shadows
By automatically detecting and eliminating shadows using AI-powered terminal devices, the problem of shadow effects in 3D reconstruction has been solved, achieving intelligent shadow elimination and improved user experience.
Patent Information
- Application Number
- CN202011211244.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-03
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2040-11-03
AI Technical Summary
In the process of 3D reconstruction, the uneven ambient light causes shadows of varying brightness on the surface of objects, which affects the user experience and the mottled appearance of the 3D model. In addition, existing shadow removal solutions require users to manually select the shadow area, which increases the interaction cost.
By employing artificial intelligence terminal devices, multi-frame image processing and convolutional neural networks are used to automatically detect and eliminate shadows, and interface controls are provided for users to adjust the shadow intensity, thus achieving intelligent shadow elimination.
It improves the intelligence of electronic devices, simplifies the shadow removal process, and enhances the user experience and the quality of 3D models.
Smart Images

Figure CN114529663B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminals, and more specifically, to a method and electronic device for eliminating shadows. Background Technology
[0002] 3D reconstruction technology refers to the creation of mathematical models of 3D objects suitable for computer representation and processing. It is a technology for creating virtual reality representations of the objective world in a computer. This technology has begun to be widely used, but one problem is that during the user's scanning of the object, uneven ambient light causes the object's surface to have varying degrees of brightness and darkness. This shadow can be called the object's own soft shadow, which results in a mottled appearance in the final reconstructed 3D model. This problem greatly affects the user's senses and experience.
[0003] Current methods for eliminating image shadows include natural image soft shadow removal. These methods require users to manually select the shadow areas, increasing the interaction cost between the machine and the human. They also place high demands on the user, and the final shadow removal effect depends on the accuracy of the user's selected area. Summary of the Invention
[0004] This application provides a method and electronic device for eliminating shadows. This method can be used in artificial intelligence (AI) terminals, which helps to improve the intelligence level of electronic devices (such as smart terminal devices, such as mobile phones) and can also improve the shadow removal effect of images.
[0005] In a first aspect, a method for eliminating shadows is provided, the method being applied in an electronic device, the method comprising: the electronic device acquiring first video information, the first video information including multiple frames of images, each of the multiple frames including an image of a foreground object; the electronic device displaying a first interface, the first interface including a first 3D model, the first 3D model being a 3D model obtained by three-dimensional reconstruction of the image of the foreground object in each frame of images, the first interface also including a first control; the electronic device detecting a first operation by a user on the first control; and in response to the first operation, the electronic device displaying a second interface, the second interface including a second 3D model, the shadow value of the second 3D model being less than the shadow value of the first 3D model, wherein the shadow value is used to characterize the degree of shadow of the 3D model.
[0006] In this embodiment, users can adjust the shadow level of the 3D model through controls on the display interface to select a 3D model that satisfies them. The user does not need to manually select the shadow area during the shadow removal process, which helps to improve the intelligence of electronic devices and thus enhance the user experience.
[0007] In some possible implementations, the method further includes: detecting a user's action of saving the second 3D model; in response to the user's action, saving the second 3D model in the electronic device; or, detecting a user's action of sharing the 3D model with other users, and in response to the action, sending the second 3D model to the other user's electronic device.
[0008] In some possible implementations, the pose of the first 3D model is the same as the pose of the second 3D model.
[0009] In conjunction with the first aspect, in some implementations of the first aspect, before displaying the first interface, the method further includes: the electronic device displaying a third interface, the third interface including a thumbnail corresponding to the first 3D model; the electronic device detecting a second operation by a user on the thumbnail; and in response to the second operation, the electronic device displaying the first interface.
[0010] In this embodiment of the application, the thumbnails displayed on the third interface make it easy for users to quickly find the 3D model they wish to view.
[0011] In some possible implementations, the thumbnail is a thumbnail of the first 3D model from any angle. For example, the thumbnail could be a thumbnail of the front view of the first 3D model.
[0012] In conjunction with the first aspect, in certain implementations of the first aspect, the electronic device displays a second interface, including: the electronic device displays a first pose of the second 3D model on the second interface; wherein the method further includes: the electronic device detecting a third operation by a user on the second interface; in response to the third operation, the electronic device displays a fourth interface, the fourth interface displaying a second pose of the second 3D model; the electronic device detects a fourth operation by a user on a first control; in response to the fourth operation, the electronic device displays a fifth interface, the fifth interface including a third 3D model, the shadow value of the third 3D model being different from the shadow value of the second 3D model.
[0013] In this embodiment, users can view 3D models in different poses through operations on the interface. When a user wants to further adjust the shadow values of the 3D model in a certain pose, the user can adjust the shadow values of the 3D model by operating the first control. The process of eliminating shadows does not require the user to manually select the shadow area, which helps to improve the intelligence of electronic devices and thus enhances the user experience.
[0014] In some possible implementations, the pose of the third 3D model is the same as the second pose of the second 3D model.
[0015] In conjunction with the first aspect, in some implementations of the first aspect, the electronic device displays a first interface, including: the electronic device extracting an image of a foreground object from each frame of the multi-frame images; the electronic device extracting a multi-scale feature map of the image of each foreground object; the electronic device decoding the multi-scale feature map of the image of each foreground object and removing the shadows of the image of each foreground object to obtain a multi-frame image after shadow removal; the electronic device performing three-dimensional reconstruction of the foreground object based on the multi-frame image after shadow removal to obtain the first 3D model; and the electronic device displaying the first interface.
[0016] In this embodiment, the electronic device extracts multi-scale features from the image of the foreground object, decodes them, and performs shadow removal processing to obtain multiple frames of images after shadow removal, thereby obtaining a 3D model through 3D reconstruction. This process uses an AI-based shadow removal scheme that can accurately remove shadows from the image of the foreground object. Simultaneously, the electronic device can provide the user with a first control, allowing the user to select the degree of shadow removal for the 3D model, thus providing a better user experience.
[0017] In conjunction with the first aspect, in some implementations of the first aspect, the electronic device decodes the multi-scale feature map of the image of each foreground object and removes the shadow of the image of each foreground object to obtain a multi-frame image after shadow removal, including: the electronic device decodes the multi-scale feature map of the image of each foreground object to obtain a shadow effect map of the image of each foreground object; the electronic device removes the shadow of the image of each foreground object according to the shadow effect map of the image of each foreground object to obtain the multi-frame image after shadow removal.
[0018] In conjunction with the first aspect, in some implementations of the first aspect, the electronic device extracts a multi-scale feature map of the image of each foreground object, including: the electronic device downsamples the first image information through a convolutional neural network to obtain a multi-scale feature map of the image of the foreground object in each frame of the image; the electronic device decodes the multi-scale feature map of the image of each foreground object, including: the electronic device upsamples the multi-scale feature map of the image of the foreground object in each frame of the image through a convolutional neural network.
[0019] In conjunction with the first aspect, in some implementations of the first aspect, the electronic device extracts the image of the foreground object of each frame of the multi-frame image, including: the electronic device removes the image of the background object of each frame of the image by using a region of interest (ROI) algorithm to obtain the image of the foreground object of each frame of the image.
[0020] Secondly, a method for eliminating shadows is provided, the method being applied in an electronic device, the method comprising: the electronic device displaying a first interface including a first frame image, the first interface further including a first control; the electronic device detecting a first operation by a user on the first control; and in response to the first operation, the electronic device displaying a second interface including a second frame image, wherein the shadow value of the second frame image is less than the shadow value of the first frame image, the shadow value being used to characterize the degree of shadow in the image.
[0021] In this embodiment, users can adjust the shadow level of an image through controls on the display interface to select a satisfactory image. The process of eliminating shadows does not require users to manually select the shadow area, which helps to improve the intelligence of electronic devices and thus enhance the user experience.
[0022] In some possible implementations, the method further includes: detecting a user's operation to save the second frame image; in response to the user's operation, saving the second frame image in the electronic device; or, detecting a user's operation to share the second frame image with other users, and in response to the operation, sending the second frame image to the other user's electronic device.
[0023] In conjunction with the second aspect, in some implementations of the second aspect, the electronic device displays a second interface, including: the electronic device extracting an image of a foreground object from the first frame image; the electronic device extracting a multi-scale feature map of the foreground object image; the electronic device decoding the multi-scale feature map of the foreground object image and removing the shadow of the foreground object image to obtain the second frame image; and the electronic device displaying the second interface.
[0024] In this embodiment, the electronic device extracts multi-scale features from the image of the foreground object, decodes them, and performs shadow removal processing to obtain a second frame image after shadow removal. This process uses an AI-based shadow removal scheme that can accurately remove shadows from the image of the foreground object. Simultaneously, the electronic device can provide the user with a first control, allowing the user to select the degree of shadow removal, thus providing a better user experience.
[0025] In conjunction with the second aspect, in some implementations of the second aspect, the electronic device decodes the multi-scale feature map of the image of the foreground object and removes the shadow of the image of the foreground object to obtain the second frame image, including: the electronic device decodes the multi-scale feature map of the image of the foreground object to obtain a shadow effect map of the image of the foreground object; the electronic device removes the shadow of the image of the foreground object based on the shadow effect map of the image of the foreground object to obtain the second frame image.
[0026] In conjunction with the second aspect, in some implementations of the second aspect, the electronic device extracts a multi-scale feature map of the image of the foreground object, including: the electronic device downsamples the image of the foreground object through a convolutional neural network to obtain a multi-scale feature map of the image of the foreground object; the electronic device decodes the multi-scale feature map of the image of the foreground object, including: the electronic device upsamples the multi-scale feature map of the image of the foreground object through a convolutional neural network.
[0027] In conjunction with the second aspect, in some implementations of the second aspect, the electronic device extracts the image of the foreground object in the first frame image by: the electronic device removing the image of the background object in the first frame image through a region of interest (ROI) algorithm to obtain the image of the foreground object in the first frame image.
[0028] Thirdly, a method for eliminating shadows is provided, which is applied in an electronic device. The method includes: the electronic device acquiring first video information, the first video information including multiple frames of images, each of the multiple frames including an image of a foreground object; the electronic device displaying a first interface, the first interface including a thumbnail corresponding to a first 3D model, the first 3D model being a 3D model obtained by three-dimensional reconstruction of the image of the foreground object in each frame; the electronic device detecting a first operation by a user on the thumbnail; in response to the first operation, the electronic device displaying a second interface, the second interface including the first 3D model, the second interface also including a first control; the electronic device detecting a first operation by a user on the first... The electronic device performs a second operation on the control; in response to the second operation, the electronic device displays a third interface, on which a first pose of a second 3D model is displayed, the shadow value of the second 3D model being less than the shadow value of the first 3D model, wherein the shadow value is used to characterize the shadow intensity of the 3D model; the electronic device detects a third operation by the user on the third interface; in response to the third operation, the electronic device displays a fourth interface, on which a second pose of the second 3D model is displayed; the electronic device detects a fourth operation by the user on the first control; in response to the fourth operation, the electronic device displays a fifth interface, on which a third 3D model is displayed, the shadow value of the third 3D model being different from the shadow value of the second 3D model.
[0029] In this embodiment, thumbnails displayed on the interface allow users to quickly locate the desired 3D model. Users can also adjust the shadow intensity of the 3D model using controls on the interface to select a satisfactory model. The shadow removal process eliminates the need for manual selection of shadow areas, thus enhancing the intelligence of the electronic device and improving the user experience. Users can also view the 3D model in different poses through interface operations. When a user wishes to further adjust the shadow values of the 3D model in a certain pose, they can do so by manipulating the first control. The shadow removal process eliminates the need for manual selection of shadow areas, further enhancing the intelligence of the electronic device and improving the user experience.
[0030] In conjunction with the third aspect, in certain implementations of the third aspect, the electronic device displays a second interface, including: the electronic device extracting an image of a foreground object from each frame of the multi-frame images; the electronic device extracting a multi-scale feature map of the image of each foreground object; the electronic device decoding the multi-scale feature map of the image of each foreground object and removing the shadows of the image of each foreground object to obtain a multi-frame image after shadow removal; the electronic device performing three-dimensional reconstruction of the foreground object based on the multi-frame image after shadow removal to obtain the first 3D model; and the electronic device displaying the second interface; wherein, the electronic device decoding the multi-scale feature map of the image of each foreground object and removing the shadows of the image of each foreground object to obtain a multi-frame image after shadow removal includes: the electronic device decoding the multi-scale feature map of the image of each foreground object to obtain an image of the foreground object... The electronic device removes the shadows from the images of each foreground object based on the shadow effect map of each foreground object, obtaining the multi-frame images after shadow removal; or, wherein the electronic device extracts multi-scale feature maps of the images of each foreground object, including: the electronic device downsamples the first image information through a convolutional neural network to obtain multi-scale feature maps of the foreground objects in each frame of the image; the electronic device decodes the multi-scale feature maps of the images of each foreground object, including: the electronic device upsamples the multi-scale feature maps of the foreground objects in each frame of the image through a convolutional neural network; or, wherein the electronic device extracts the image of the foreground object in each frame of the multi-frame images, including: the electronic device removes the image of the background object in each frame of the image through a region of interest (ROI) algorithm to obtain the image of the foreground object in each frame of the image.
[0031] In this embodiment, the electronic device extracts multi-scale features from the image of the foreground object, decodes them, and performs shadow removal processing to obtain multiple frames of images after shadow removal, thereby obtaining a 3D model through 3D reconstruction. This process uses an AI-based shadow removal scheme that can accurately remove shadows from the image of the foreground object. Simultaneously, the electronic device can provide the user with a first control, allowing the user to select the degree of shadow removal for the 3D model, thus providing a better user experience.
[0032] Fourthly, an electronic device is provided, comprising: one or more processors; a memory; and one or more computer programs. The one or more computer programs are stored in the memory and include instructions. When the instructions are executed by the electronic device, the electronic device performs the following steps: acquiring first video information, the first video information including multiple frames of images, each of the multiple frames including an image of a foreground object; displaying a first interface, the first interface including a first 3D model, the first 3D model being a 3D model obtained by three-dimensional reconstruction of the image of the foreground object in each frame, the first interface also including a first control; detecting a first operation by a user on the first control; and in response to the first operation, displaying a second interface, the second interface including a second 3D model, the second 3D model having a shadow value less than the shadow value of the first 3D model, wherein the shadow value is used to characterize the shadow intensity of the 3D model.
[0033] In conjunction with the fourth aspect, in some implementations of the fourth aspect, when the instruction is executed by the electronic device, the electronic device performs the following steps: before displaying the first interface, displays a third interface, the third interface including a thumbnail corresponding to the first 3D model; detects a second operation by the user on the thumbnail; and in response to the second operation, displays the first interface.
[0034] In conjunction with the fourth aspect, in some implementations of the fourth aspect, when the instruction is executed by the electronic device, the electronic device performs the following steps: displaying a first pose of the second 3D model on the second interface; detecting a third operation by the user on the second interface; in response to the third operation, displaying a fourth interface on which a second pose of the second 3D model is displayed; detecting a fourth operation by the user on the first control; and in response to the fourth operation, displaying a fifth interface on which a third 3D model is included, the shadow value of which is different from the shadow value of the second 3D model.
[0035] In conjunction with the fourth aspect, in some implementations of the fourth aspect, when the instruction is executed by the electronic device, the electronic device performs the following steps: extracting the image of the foreground object in each frame of the multi-frame image; extracting the multi-scale feature map of the image of each foreground object; decoding the multi-scale feature map of the image of each foreground object and removing the shadow of the image of each foreground object to obtain the multi-frame image after shadow removal; performing three-dimensional reconstruction of the foreground object based on the multi-frame image after shadow removal to obtain the first 3D model; and displaying the first interface.
[0036] In conjunction with the fourth aspect, in some implementations of the fourth aspect, when the instruction is executed by the electronic device, the electronic device performs the following steps: decoding the multi-scale feature map of the image of each foreground object to obtain the shadow effect map of the image of each foreground object; and removing the shadow of the image of each foreground object based on the shadow effect map of the image of each foreground object to obtain the multi-frame image after shadow removal.
[0037] In conjunction with the fourth aspect, in some implementations of the fourth aspect, when the instruction is executed by the electronic device, the electronic device performs the following steps: downsampling the first image information through a convolutional neural network to obtain a multi-scale feature map of the foreground object image of each frame; and upsampling the multi-scale feature map of the foreground object image of each frame through a convolutional neural network.
[0038] In conjunction with the fourth aspect, in some implementations of the fourth aspect, when the instruction is executed by the electronic device, the electronic device performs the following steps: removing the image of the background object of each frame image by means of the Region of Interest (ROI) algorithm, thereby obtaining the image of the foreground object of each frame image.
[0039] Fifthly, an electronic device is provided, comprising: one or more processors; a memory; and one or more computer programs. The one or more computer programs are stored in the memory, and the one or more computer programs include instructions. When the instructions are executed by the electronic device, the electronic device performs the following steps: displaying a first interface including a first frame image, the first interface also including a first control; detecting a first operation by a user on the first control; and in response to the first operation, displaying a second interface including a second frame image, wherein the shadow value of the second frame image is less than the shadow value of the first frame image, the shadow value being used to characterize the degree of shadow in the image.
[0040] In conjunction with the fifth aspect, in some implementations of the fifth aspect, when the instruction is executed by the electronic device, the electronic device performs the following steps: extracting an image of a foreground object in the first frame image; extracting a multi-scale feature map of the foreground object image; decoding the multi-scale feature map of the foreground object image and removing the shadow of the foreground object image to obtain the second frame image; and displaying the second interface.
[0041] In conjunction with the fifth aspect, in some implementations of the fifth aspect, when the instruction is executed by the electronic device, the electronic device performs the following steps: decoding the multi-scale feature map of the image of the foreground object to obtain the shadow effect map of the image of the foreground object; and eliminating the shadow of the image of the foreground object based on the shadow effect map of the image of the foreground object to obtain the second frame image.
[0042] In conjunction with the fifth aspect, in some implementations of the fifth aspect, when the instruction is executed by the electronic device, the electronic device performs the following steps: downsampling the image of the foreground object through a convolutional neural network to obtain a multi-scale feature map of the image of the foreground object; and upsampling the multi-scale feature map of the image of the foreground object through a convolutional neural network.
[0043] In conjunction with the fifth aspect, in some implementations of the fifth aspect, when the instruction is executed by the electronic device, the electronic device performs the following steps: removing the image of the background object in the first frame image by using the Region of Interest (ROI) algorithm to obtain the image of the foreground object in the first frame image.
[0044] A sixth aspect provides an electronic device, comprising: one or more processors; a memory; and one or more computer programs. The one or more computer programs are stored in the memory, and the one or more computer programs include instructions. When the instructions are executed by the electronic device, the electronic device performs the following steps: acquiring first video information, the first video information including multiple frames of images, each of the multiple frames including an image of a foreground object; displaying a first interface, the first interface including a thumbnail corresponding to a first 3D model, the first 3D model being a 3D model obtained by 3D reconstruction of the image of the foreground object in each frame; detecting a first operation by a user on the thumbnail; responding to the first operation, displaying a second interface, the second interface including the first 3D model, the second interface also including a first control; detecting a second operation by a user on the first control. The user performs the following actions: A third interface is displayed in response to the second action, showing the first pose of the second 3D model. The shadow value of the second 3D model is less than the shadow value of the first 3D model, where the shadow value characterizes the degree of shadow in the 3D model. A third user action on the third interface is detected. In response to the third action, a fourth interface is displayed, showing the second pose of the second 3D model. A fourth user action on the first control is detected. In response to the fourth action, a fifth interface is displayed, showing the third 3D model, whose shadow value differs from that of the second 3D model.
[0045] In conjunction with the sixth aspect, in some implementations of the sixth aspect, when the instruction is executed by the electronic device, the electronic device performs the following steps: extracting the image of the foreground object in each frame of the multi-frame image; extracting the multi-scale feature map of the image of each foreground object; decoding the multi-scale feature map of the image of each foreground object and removing the shadows of the image of each foreground object to obtain the multi-frame image after shadow removal; performing three-dimensional reconstruction of the foreground object based on the multi-frame image after shadow removal to obtain the first 3D model; and displaying the second interface; wherein, when the instruction is executed by the electronic device, the electronic device performs the following steps: decoding the multi-scale feature map of the image of each foreground object to obtain the multi-frame image of each foreground object. The image contains shadow effects; based on the shadow effects of each foreground object's image, the shadows of each foreground object's image are eliminated to obtain the shadow-removed multi-frame image; or, when the instruction is executed by the electronic device, the electronic device performs the following steps: downsampling the first image information using a convolutional neural network to obtain a multi-scale feature map of the foreground object's image in each frame; upsampling the multi-scale feature map of the foreground object's image in each frame using a convolutional neural network; or, when the instruction is executed by the electronic device, the electronic device performs the following steps: removing the background object's image in each frame using a Region of Interest (ROI) algorithm to obtain the foreground object's image in each frame.
[0046] A seventh aspect provides an electronic device comprising: one or more processors; a memory; and one or more computer programs. The one or more computer programs are stored in the memory, and the one or more computer programs include instructions. When the instructions are executed by the electronic device, the electronic device performs the method in any possible implementation of the first aspect described above.
[0047] Eighthly, an electronic device is provided, comprising: one or more processors; a memory; and one or more computer programs. The one or more computer programs are stored in the memory, and the one or more computer programs include instructions. When the instructions are executed by the electronic device, the electronic device performs the method in any possible implementation of the second aspect described above.
[0048] A ninth aspect provides an electronic device comprising: one or more processors; a memory; and one or more computer programs. The one or more computer programs are stored in the memory, and the one or more computer programs include instructions. When the instructions are executed by the electronic device, the electronic device performs the method in any possible implementation of the third aspect described above.
[0049] In a tenth aspect, a computer program product comprising instructions is provided, which, when run on an electronic device, causes the electronic device to perform the method described in the first aspect; or, when run on an electronic device, causes the electronic device to perform the method described in the second aspect; or, when run on an electronic device, causes the electronic device to perform the method described in the third aspect.
[0050] Eleventhly, a computer-readable storage medium is provided, comprising instructions that, when executed on an electronic device, cause the electronic device to perform the method described in the first aspect; or, when executed on an electronic device, cause the electronic device to perform the method described in the second aspect; and when executed on an electronic device, cause the electronic device to perform the method described in the third aspect.
[0051] In a twelfth aspect, a chip is provided for executing instructions, wherein when the chip is running, the chip executes the method described in the first aspect above; or, when the chip is running, the chip executes the method described in the second aspect above; or, when the chip is running, the chip executes the method described in the third aspect above. Attached Figure Description
[0052] Figure 1 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application.
[0053] Figure 2 This is a set of graphical user interfaces provided in the embodiments of this application.
[0054] Figure 3 This is another set of graphical user interfaces provided in the embodiments of this application.
[0055] Figure 4 This is a schematic structural diagram of the system framework of the electronic device provided in the embodiments of this application.
[0056] Figure 5 This describes the process of shadow elimination implemented internally in the electronic device provided in this application embodiment.
[0057] Figure 6 This is a schematic flowchart of the method for eliminating shadows provided in the embodiments of this application.
[0058] Figure 7 This is the process of acquiring training data by setting up a collection environment, as provided in the embodiments of this application. Detailed Implementation
[0059] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in this text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "plural" or "multiple" refers to two or more than two.
[0060] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this embodiment, unless otherwise stated, "a plurality of" means two or more.
[0061] The methods provided in this application can be applied to electronic devices such as mobile phones, tablets, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). This application does not impose any restrictions on the specific type of electronic device.
[0062] For example, Figure 1A schematic diagram of the structure of electronic device 100 is shown. Electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, antenna 1, antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0063] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0064] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0065] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.
[0066] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0067] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0068] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple I2C buses. The processor 110 can couple to the touch sensor 180K, charger, flash, camera 193, etc., through different I2C bus interfaces. For example, the processor 110 can couple to the touch sensor 180K through the I2C interface, enabling the processor 110 and the touch sensor 180K to communicate through the I2C bus interface, thereby realizing the touch function of the electronic device 100.
[0069] The I2S interface can be used for audio communication. In some embodiments, the processor 110 may include multiple I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to enable communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the I2S interface to enable the function of answering phone calls through a Bluetooth headset.
[0070] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled via the PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 via the PCM interface, enabling the function of answering phone calls through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0071] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the UART interface to enable music playback through Bluetooth headphones.
[0072] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI) and a display serial interface (DSI). In some embodiments, the processor 110 and the camera 193 communicate via the CSI interface to enable the electronic device 100 to capture images. The processor 110 and the display screen 194 communicate via the DSI interface to enable the electronic device 100 to display images.
[0073] The GPIO interface can be configured via software. It can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to a camera 193, a display screen 194, a wireless communication module 160, an audio module 170, a sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0074] USB port 130 is a USB standard compliant interface, specifically a Mini USB port, Micro USB port, USB Type-C port, etc. USB port 130 can be used to connect a charger to charge electronic device 100, and can also be used for data transfer between electronic device 100 and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other electronic devices, such as AR devices.
[0075] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0076] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device via the power management module 141.
[0077] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, external memory, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.
[0078] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.
[0079] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.
[0080] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.
[0081] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through an audio device (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.
[0082] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.
[0083] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling electronic device 100 to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).
[0084] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0085] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1.
[0086] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.
[0087] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.
[0088] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0089] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.
[0090] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0091] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.
[0092] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.
[0093] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.
[0094] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.
[0095] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.
[0096] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or make hands-free calls through the speaker 170A.
[0097] The receiver 170B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the electronic device 100 answers a telephone call or voice message, the receiver 170B can be brought close to the ear to listen to the voice.
[0098] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Electronic device 100 may have at least one microphone 170C. In some embodiments, electronic device 100 may have two microphones 170C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, electronic device 100 may also have three, four, or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.
[0099] The 170D headphone jack is used to connect wired headphones. The 170D headphone jack can be a USB 130 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.
[0100] The pressure sensor 180A is used to sense pressure signals and can convert the pressure signals into electrical signals. In some embodiments, the pressure sensor 180A can be disposed on the display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc.
[0101] The gyroscope sensor 180B can be used to determine the motion attitude of the electronic device 100. In some embodiments, the angular velocity of the electronic device 100 about three axes (i.e., the x, y, and z axes) can be determined by the gyroscope sensor 180B.
[0102] The barometric pressure sensor 180C is used to measure air pressure. In some embodiments, the electronic device 100 calculates altitude using the air pressure value measured by the barometric pressure sensor 180C to assist in positioning and navigation.
[0103] The magnetic sensor 180D includes a Hall sensor. The electronic device 100 can use the magnetic sensor 180D to detect the opening and closing of the flip cover.
[0104] The 180E accelerometer can detect the magnitude of acceleration of electronic device 100 in various directions (typically three axes). When electronic device 100 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the posture of electronic devices and applied to applications such as screen orientation switching and pedometers.
[0105] A distance sensor 180F is used to measure distance. Electronic device 100 can measure distance via infrared or laser. In some embodiments, during a shooting scene, electronic device 100 can utilize the distance sensor 180F to measure distance for rapid focusing.
[0106] The proximity light sensor 180G may include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The LED may be an infrared LED.
[0107] The 180L ambient light sensor is used to detect ambient light intensity.
[0108] The fingerprint sensor 180H is used to collect fingerprints. The electronic device 100 can utilize the characteristics of the collected fingerprints to achieve fingerprint unlocking, accessing application locks, taking photos with fingerprints, answering calls with fingerprints, etc.
[0109] Temperature sensor 180J is used to detect temperature. In some embodiments, electronic device 100 uses the temperature detected by temperature sensor 180J to execute a temperature processing strategy.
[0110] Touch sensor 180K, also known as a "touch panel". Touch sensor 180K can be set on display screen 194. Touch sensor 180K and display screen 194 together form a touch screen, also known as a "touch screen". Touch sensor 180K is used to detect touch operations on or near it.
[0111] The bone conduction sensor 180M can acquire vibration signals.
[0112] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Electronic device 100 can receive button input and generate key signal inputs related to user settings and function control of electronic device 100.
[0113] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. For example, different vibration feedback effects can correspond to touch operations performed on different applications (such as taking photos, playing audio, etc.). Motor 191 can also correspond to different vibration feedback effects for touch operations performed on different areas of the display screen 194. Different application scenarios (such as time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also be customized.
[0114] Indicator 192 can be an indicator light, used to indicate charging status, power changes, or to indicate messages, missed calls, notifications, etc.
[0115] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to make contact with or separate from the electronic device 100.
[0116] Figure 2 This is a set of graphical user interfaces (GUIs) provided in the embodiments of this application.
[0117] See Figure 2 The GUI shown in (a) is the phone's home screen. This GUI includes icons for multiple applications, including icon 201 for a 3D reconstruction application. When the phone detects a user tapping icon 201 on the home screen, it displays... Figure 2 The GUI shown in (b) is shown in the image.
[0118] See Figure 2 The GUI shown in (b) is the display interface of the 3D reconstruction application. This interface includes three functional areas: "Acquisition," "Modeling Results," and "My Account." Currently, the phone displays the acquisition interface, which includes two modeling methods: video modeling and photo modeling. When the phone detects that the user clicks on video modeling 202, it displays the following... Figure 2 The GUI shown in (c) is shown in the image.
[0119] See Figure 2The GUI shown in (c) is a video capture interface. This interface includes two video capture methods: shooting and selecting from recorded videos. If the phone detects that the user selects the shooting control 203, the phone can activate the camera to capture video data.
[0120] See Figure 2 In (d), users can use their mobile phones to record real-time video of the object, thereby obtaining a 3D model after removing shadows from the recorded video.
[0121] It should be understood that the method by which the user photographs the object is not limited in the embodiments of this application. For example... Figure 2 As shown in (d), users can take a panoramic photo of an object perpendicular to the table; or, users can take a panoramic photo of an object parallel to the tabletop; or, users can take a panoramic photo of an object at a certain angle to the tabletop.
[0122] In this embodiment, the user can also choose an offline photo or video reconstruction method. If the phone detects that the user selects from a recorded video, the phone can display the recorded videos or photos saved on the phone, and the phone can perform 3D reconstruction based on the user's selection of the saved videos or photos.
[0123] See Figure 2 The GUI shown in (e) is the display interface after video recording of the object is completed. When the phone detects that the user has finished recording the video, it can display a reminder box. This reminder box can prompt the user, "Video recording is complete. Do you want to start 3D reconstruction?". When the phone detects that the user clicks on control 204, it can perform 3D reconstruction based on the real-time video data.
[0124] See Figure 2 The GUI shown in (f) is the interface for displaying the modeling results. These results include the modeling of objects (e.g., a bear) from a real-time recorded video. When the user clicks on the thumbnail 205 of the 3D model obtained through modeling, the phone can display... Figure 2 The GUI shown in (g) is shown in the image.
[0125] It should be understood that the thumbnail 205 can be a thumbnail of the 3D model from any angle. For example, the thumbnail 205 can be a thumbnail of the front view of the 3D model.
[0126] See Figure 2The GUI shown in (g) is the display interface for the 3D model of the reconstructed object. When the phone detects that the user clicks on the thumbnail 205, the phone can display the 3D model of the object after 3D reconstruction. For example, as shown in... Figure 2 As shown in (g), the phone currently displays the front view of the 3D model. Users can slide their fingers across the touchscreen to view the 3D model. The display also includes a sliding control 206 for adjusting the shadow removal effect. When the phone detects a leftward swipe on the touchscreen, it can display... Figure 2 The GUI shown in (h) is shown in the image.
[0127] See Figure 2 The GUI shown in (h) responds to the phone detecting a user swiping left on the touchscreen; the phone can then rotate the 3D model in a preset direction. For example... Figure 2 In the (h) section, the phone can display the back of the 3D model.
[0128] It should be understood that in the embodiments of this application, the mobile phone may rotate the 3D model in a preset direction when it detects that the user is sliding on the touch screen in any direction; or, the mobile phone may also display other controls on the touch screen and realize the rotation display of the 3D model by detecting the user's control of other controls.
[0129] It should also be understood that when the slider control 206 is at the far left, the shadow of the 3D model is not eliminated; as the slider control moves to the right, the shadow of the 3D model is gradually eliminated. When the slider control 206 is at the far right, the shadow of the 3D model is completely eliminated. Figure 2 As shown in (g) and (h), the sliding control 206 is located on the far left, and the shadow of the 3D model has not been eliminated. Figure 2 In (g) and (h), the slashes represent shadows on the object.
[0130] Figure 3 This is another set of GUIs provided in the embodiments of this application.
[0131] When the phone detects that the user has dragged the slider 206 to the right to the middle position, the phone can display as follows: Figure 3 The GUI shown in (a) is shown in the image.
[0132] See Figure 3 The GUI shown in (a) is another display interface for the 3D model of the reconstructed object. Compared to Figure 2 The 3D model shown in (g) is as follows. Figure 3 The 3D model shown in (a) has better shadow removal.
[0133] It should be understood that, compared to Figure 2 The 3D model shown in (g) is as follows. Figure 3 The better shadow removal effect of the 3D model shown in (a) can also be understood as... Figure 3 The shadow value of the 3D model shown in (a) is less than Figure 2 The shadow value of the 3D model shown in (g) is the value of the shadow value of the model.
[0134] The shadow value of the 3D model in this embodiment can be understood as including the degree of shadow in the image of the 3D model. A description of the shadow value is provided below and will not be repeated here.
[0135] It should be understood that Figure 3 (a) shows a schematic diagram of the front view of the 3D model, where some shadows have been removed. When the phone detects a user swiping on the touchscreen, it can display other angles of the 3D model (e.g., the back view). The back view of the 3D model viewed by the user at this time has the same de-shading effect as the front view before rotation. If the user is satisfied with the current de-shading effect of the 3D model, they can share or save it. If the user is not satisfied with the current de-shading effect, they can continue to swipe the control 206 to the right.
[0136] It should also be understood that the mobile phone can adjust the overall shadow removal effect of the 3D model through control 206; or, the mobile phone can only remove shadows from the part of the 3D model displayed on the current screen.
[0137] When the phone detects that the user continues to drag the slider 206 to the far right, the phone can display as follows: Figure 3 The GUI shown in (b) is shown in the image.
[0138] See Figure 3 The GUI shown in (b) is another display interface for the 3D model of the reconstructed object. Compared to Figure 3 The modeling results shown in (a) are as follows: Figure 3 The shadows of the 3D model shown in (b) have been completely eliminated.
[0139] The shadow elimination method provided in this application embodiment helps to solve the problem of uneven texture brightness after 3D reconstruction, thereby improving the user's senses and experience.
[0140] The method for eliminating shadows in the embodiments of this application is described below.
[0141] Figure 4A schematic structural diagram of the system framework of the electronic device provided in an embodiment of this application is shown. Figure 4 As shown, the electronic device includes a main algorithm model and a region of interest (ROI) algorithm. The electronic device acquires keyframe images required for 3D reconstruction, removes the background from each keyframe image using an existing depth sensor-based ROI algorithm, and then inputs it into the main algorithm model. The main algorithm model outputs a shadow effect image, which is then multiplied by a user-adjustable weight λ on the user interface and added to the input keyframe image. This achieves the removal of shadows after object reconstruction. Simultaneously, the user interface allows users to interactively adjust the shadow removal effect.
[0142] Figure 5 The process of shadow elimination implemented internally in an electronic device according to an embodiment of this application is illustrated. For example... Figure 5 As shown, the multi-scale feature encoder and multi-scale feature decoder constitute a... Figure 4 The main algorithm model is shown in the figure.
[0143] The following is combined Figure 6 The method 600 for eliminating shadows shown in the embodiments of this application illustrates the process of shadow elimination within an electronic device. The method 600 includes:
[0144] S601, the electronic device obtains multiple frames of images required for object modeling from the storage unit and inputs them into the processing unit.
[0145] It should be understood that a storage unit can be Figure 1 The internal memory 121 or the memory card connected to the external memory interface 120 shown. This processing unit can output... Figure 1 The processor 110 shown is an example of an NPU.
[0146] For example, such as Figure 2 As shown in (d), the mobile phone can select multiple frames from a video recorded by the user in real time. For example, when the pose of an object is perpendicular to the table, one frame can be selected for every 20° change, so 18 frames can be selected when the pose changes 360°.
[0147] It should be understood that in the embodiments of this application, the electronic device may select some frames from the recorded video for three-dimensional reconstruction; or, the electronic device may select all frames from the recorded video for three-dimensional reconstruction, and the embodiments of this application do not limit this.
[0148] It should also be understood that when an electronic device selects multiple frames from a recorded video, it may do so by selecting multiple frames in a direction perpendicular to the desktop; or, it may do so in a direction parallel to the desktop; or, it may do so in a direction at any angle to the desktop. This application does not limit the scope of the embodiments.
[0149] In one embodiment, the electronic device can also select the multi-frame images according to user settings. For example, the user can set the selection of the multi-frame images in a certain circumferential direction, at 20° or 30°; or the user can set the selection of the multi-frame images from a certain direction of an object (e.g., front, back, top, bottom, or side); or when the electronic device detects a preset operation by the user during video recording (e.g., a click operation on the touchscreen), it can use the currently displayed frame as one of the multi-frame images; or when the electronic device detects that the user has ended recording, it can display all the image information in the recorded video to the user, allowing the user to select the multi-frame images from all the image information.
[0150] In one embodiment, the pose change of an object in this application embodiment can be determined by a simultaneous localization and mapping (SLAM) algorithm.
[0151] S602, the electronic device selects a first keyframe image from the multi-frame images, removes the background object image from the first keyframe image, and obtains the foreground object image.
[0152] It should be understood that in the embodiments of this application, the process of the electronic device removing the background object image of the first keyframe image is to completely (or 100%) remove the background object image; the process of the electronic device eliminating the shadow of the foreground object image can be a process of gradually eliminating the shadow according to a certain proportion as the position of the control 206 (for example, the control 206 from the leftmost to the rightmost), until the shadow of the foreground object image is completely removed when the control 206 is located at the rightmost position.
[0153] It should be understood that, optionally, control 206 represents the degree of shadow removal. When it is on the far right, it may not be completely removed. The degree of shadow removal on the far right can be set by the default setting of the electronic device, or it can be manually set and changed by the user in the settings options of the electronic device, or it can be automatically generated by the electronic device according to the habits of most users. There is no limitation here.
[0154] It should be understood that the first keyframe in the embodiments of this application may be the first image selected by the electronic device in a multi-frame image, or it may be any one of the frames, or each frame in the multi-frame image may be the first keyframe image.
[0155] For example, an electronic device may select a frame image between 0° and 20° from 18 frames as the first keyframe image.
[0156] The first keyframe image includes images of foreground objects (wherein, the images of foreground objects are the images of the objects in the first keyframe image that need to be reconstructed in 3D, such as...). Figure 2 (d) shows the image of the bear on the table and the image of the background object (where the image of the background object is the image of the object other than the foreground object, for example...). Figure 2 (As shown in image (d) of the table), electronic devices can remove background objects from the image using ROI algorithms based on depth sensors (e.g., time-of-flight (TOF) sensors, structured light sensors, or LiDAR sensors); alternatively, they can use ROI algorithms based on red-green-blue (RGB) sensors; or they can use other artificial intelligence (AI) segmentation algorithms (e.g., saliency segmentation algorithms, portrait segmentation algorithms, etc.). The input to AI segmentation algorithms can be data detected by RGB sensors, depth sensors, or a combination of both. AI segmentation algorithms can employ deep learning algorithms instead of traditional image segmentation algorithms.
[0157] It should be understood that the above only introduces a few ways to remove background objects from images. Electronic devices can also use other methods to remove background images, and the embodiments of this application are not limited to these.
[0158] S603, the electronic device inputs an image of the foreground object into the encoder to extract multi-scale feature maps.
[0159] Optionally, a convolutional neural network can be used as the encoder to leverage its feature extraction capabilities to obtain multi-scale feature maps. For example, the multi-scale feature maps can be divided into N layers. Later layers contain more global feature maps, while earlier layers contain more local feature maps. For instance, when N is 3, the first layer includes relatively local feature maps (e.g., ...). Figure 2 (g) The bear's eyes in the middle; the second layer includes more global feature maps compared to the first layer (e.g., Figure 2(g) The bear's head in the middle); the third layer includes global feature maps (e.g., Figure 2 (g) the whole of the bear cub.
[0160] A layer can also be divided into multiple feature maps by gradients in a preset direction (e.g., horizontal gradients, vertical gradients, or gradients in the 45° direction). For example, for the first layer, it can be divided into multiple feature maps according to gradients in a preset direction, including feature maps corresponding to the bear's eyes, the bear's nose, and the bear's mouth.
[0161] It should be understood that N, and the number of feature maps in each layer, are determined by the network architecture of the convolutional neural network.
[0162] It should also be understood that the multi-scale feature maps extracted by the electronic device are the sum of the feature maps of each layer. For example, if the multi-scale feature maps are divided into 3 layers, and each layer includes 4 feature maps, then the electronic device will extract 12 multi-scale feature maps.
[0163] For example, for the j-th feature map of the L-th layer, 1≤L≤N, the L-th layer includes P feature maps, 1≤j≤P, and its calculation formula is as follows:
[0164]
[0165] Where f is the activation function, and in this embodiment, the Leak ReLUI activation function can be used. M j It is in the L-1 layer and The corresponding feature map; For the corresponding convolution kernel; * represents the convolution operation; For the corresponding bias terms; and The value can be obtained during the training phase using the backpropagation algorithm.
[0166] In this embodiment of the application, the electronic device can use fully convolutional repeated and continuous operations, so that the encoder can obtain a more accurate feature map from the combination of contextual information and detailed information.
[0167] Electronic devices can employ multiple downsampling operations to increase robustness to small perturbations in the image, reduce the risk of overfitting, decrease computational load, and increase the receptive field of view, while simultaneously acquiring multi-scale features of the image.
[0168] Electronic devices can achieve downsampling by increasing the convolution stride instead of pooling, allowing the downsampling process to also learn.
[0169] In this embodiment, the encoder can downsample the first keyframe image after shadow removal to obtain multiple multi-scale feature maps, thereby realizing the extraction of global features to local features of the object.
[0170] S604, the electronic device inputs the multi-scale feature map into the decoder to obtain the shadow effect map corresponding to the image of the foreground object.
[0171] In one embodiment, the shadow effect image can be an image with negative shadow values generated based on the image of the foreground object. By superimposing the shadow effect image and the image of the foreground object, an image with the shadow effect removed can be obtained.
[0172] It should be understood that the shadow value of the image in the embodiments of this application can represent the degree of shadow in the image. A larger shadow value indicates a greater degree of shadow in the image; a smaller shadow value indicates a less degree of shadow in the image. For example, Figure 2 The shadow value of the foreground object in (g) is greater than Figure 3 Image shadow values of the foreground object shown in (a) are... Figure 3 The shadow value of the foreground object in (a) is greater than Figure 3 The shadow value of the foreground object shown in (b) is a graph. This shadow effect is related to the object's material, texture color, lighting when the object was photographed, and the degree of occlusion of the object in the photographed image (e.g., occlusion by the object itself or occlusion by other objects).
[0173] It should also be understood that the shadow value of the 3D model in the embodiments of this application can be understood as including the degree of shadow of the image of the 3D model. For example, Figure 2 The shading value of the 3D model shown in (g) is greater than Figure 3 The shadow values of the 3D model shown in (a) are... Figure 3 The shading value of the 3D model shown in (a) is greater than Figure 3 The shadow values of the 3D model shown in (b) are shown in the image.
[0174] For example, when users view 3D models in the same environment, such as Figure 3 In (a) of the above, when the phone detects that the user has adjusted the sliding control 206 to the middle position, Figure 3 The shading level of the 3D model shown in (a) is greater than... Figure 2 The shading level of the 3D model shown in (g) is reduced. Figure 3 (a) compared to Figure 2 The reduction of the (g) diagonal lines indicates a decrease in the degree of shading. When the phone detects that the user has moved the slider 206 to the far right... Figure 3 The shading level of the 3D model shown in (b) is greater than... Figure 3 The shading level of the 3D model shown in (a) is reduced. Figure 3 (b) in the diagram has no slash, indicating that the shading has been completely eliminated.
[0175] like Figure 5 As shown, the decoder in this embodiment can adopt a network structure symmetrical to the encoder. Each decoding layer is step-connected to the encoding layer. The decoder output is guided by the input image, thereby realizing the layer-by-layer decoding of multi-scale feature maps.
[0176] The step connection method is used to connect to the encoder so that only one encoder needs to be trained, while the feature maps of different scales can be learned and selected by the decoder.
[0177] Similar to the encoder, the decoder can also employ full convolution operations, using upsampling to decode multi-scale feature maps, thereby obtaining the shadow effect map corresponding to the first keyframe image.
[0178] It should be understood that in the embodiments of this application, the encoder and decoder constitute the main algorithm model. The main algorithm model can be obtained from a cloud server when a 3D reconstruction application is installed on the mobile phone; or it can be stored on a cloud server, and after the electronic device uploads the video or photo to the server, the server uses the main algorithm model to perform 3D reconstruction of the objects in the video or photo. After completing the 3D reconstruction of the object, the server sends the modeling result to the mobile phone.
[0179] The following describes the process of obtaining training data for the main algorithm model.
[0180] In this embodiment of the application, the training data of the main algorithm model can be obtained in two ways: one is by setting up a collection environment, and the other is by rendering simulation.
[0181] Figure 7 This demonstrates the process of acquiring training data by setting up a data acquisition environment. The acquisition environment can be constructed using a supplementary lighting device 701 and a lighting device 702 to create a 360° shadowless lighting environment. An image acquisition device is fixed using a tripod and controlled by a remote control to acquire images of the object 703 on the table that needs to be reconstructed in 3D. When all the lights are on, one image is captured; this image can be used as the ground truth image. Keeping the acquisition device stationary, the lights are randomly turned off, and images are captured; these images can be used as the shadowed images. For rendering simulation, rendering tools are used to generate the corresponding shadowed and shadowless images by controlling virtual lighting and shadows.
[0182] For both of these training data acquisition methods, rendering simulation has a lower cost and can quickly acquire a large amount of data, but it still differs from the real image. Therefore, it is possible to combine these two methods to acquire training data.
[0183] The shaded and unshaded images obtained from the training data are used to construct the encoder and decoder in the main algorithm model. Updates to the main algorithm model can then be achieved using a truth discriminator.
[0184] Since the output of the main algorithm model may not be realistic—for example, if the input is a toy cat image with shadows, the output might be a blurry image, or the appearance of the toy cat in the output image might differ significantly from that in the original image—a realism discriminator can mitigate this issue. After inputting the output image into the realism discriminator, if the discriminator determines that the image is not realistic, the main algorithm model can be updated using backpropagation.
[0185] In this embodiment, to make the final generated image more realistic and alleviate the problem of high acquisition cost of real images, a realism discriminator is added during training. The input to the realism discriminator is the real image and the output image after encoder and decoder processing. The real image is labeled 1, and the output image is labeled 0.
[0186] The encoder and decoder constitute the main algorithm model, which is updated separately and alternately with the truth discriminator during training. First, the main algorithm model is trained and updated, then the parameters of the main algorithm model are fixed and the truth discriminator model is trained and updated, and so on, alternating between training.
[0187] S605, the electronic device obtains an image after removing the shadow effect based on the shadow effect diagram and the image of the foreground object, and uses the obtained keyframe image after removing the shadow effect as a new keyframe image.
[0188] It should be understood that the shadow of an object is related to the object's material, texture, color, the lighting when the object was photographed, and the degree of occlusion of the object in the photographed image (e.g., occlusion by the object itself or occlusion by other objects).
[0189] In this embodiment, the shadow value of the keyframe image after shadow removal is less than the shadow value of the first keyframe image. This shadow value can be used to represent the degree of shadow in the image. For example, when a user views the keyframe image after shadow removal and the first keyframe image under the same external conditions (such as using the same mobile phone), they can find that the degree of shadow is significantly reduced.
[0190] For example, the 3D model when the sliding control 206 is at the far right (e.g.) Figure 3The shadow value of the 3D model in (b) is less than that of the 3D model when the sliding control 206 is at the far left (e.g., Figure 2 The shadow value of the 3D model in (g) can be used to represent the degree of shadow of the 3D model. For example, when a user views the 3D model with the motion control 206 on the far right and the 3D model with the sliding control 206 on the far left in the same environment, the degree of shadow of the 3D model can be found to be significantly reduced.
[0191] In one embodiment, the electronic device can obtain an image after removing the shadow effect based on the shadow effect map, the weight λ, and the image of the foreground object.
[0192] For example, the image after removing the shadow effect can be obtained by multiplying the effect image by a weight λ and then overlaying it with the image of the foreground object.
[0193] In this embodiment of the application, the weight λ can be obtained through the user interface of the electronic device. For example, as shown... Figure 2 As shown in (g), when the sliding control is located at the far left, the weight λ can be 0, and the shadow of the object after 3D reconstruction is not eliminated.
[0194] For example, such as Figure 3 As shown in (b), when the sliding control is located on the far right, the weight λ can be 1, and the shadow of the object after 3D reconstruction is completely eliminated.
[0195] For example, such as Figure 3 As shown in (a), when the sliding control is in the middle position, the weight λ can be 0.5, and the shadow effect of the 3D reconstructed object is reduced by half.
[0196] It should also be understood that the sliding control can be a process of weight λ changing from 0 to 1 from left to right. The weight λ can also be obtained by the sliding position and a preset algorithm. The embodiments of this application are not limited to this.
[0197] In this embodiment of the application, a sliding UI component (e.g., ...) is added to the user interface. Figure 2 The component (g) in the control 206) is not limited to a specific location, but depends on the layout of the user interface; the component is linked with the weight λ to control the size value of λ, thereby achieving different degrees of shadow removal effect.
[0198] Similarly, the electronic device acquires the second keyframe image from multiple frames of images, and repeats the steps (2)-(5) above to obtain multiple new keyframe images.
[0199] S606 performs 3D reconstruction of foreground objects using multiple keyframe images after shadow removal.
[0200] It is understood that the 3D reconstruction technology can be any existing 3D reconstruction technology, such as the KinectFusion algorithm, ElasticReconstruction algorithm, BundleFusion algorithm, or InfiniTAM algorithm. The embodiments of this application do not specifically limit the 3D reconstruction algorithm.
[0201] The shadow elimination method in this application embodiment can effectively eliminate the influence of ambient light changes on the texture after 3D reconstruction. For general users, the ambient lighting is often not ideal, which inevitably leads to uneven brightness in the reconstructed object texture.
[0202] In this embodiment, a 3D reconstruction algorithm can be embedded in the electronic device to process keyframe images under user interaction, ultimately solving the problem of uneven texture brightness and greatly improving the user's sensory experience. Furthermore, this embodiment can also be used in other scenarios requiring texture image processing, exhibiting a degree of portability and universality.
[0203] This application provides an electronic device comprising: one or more processors, one or more memories, wherein the one or more memories store one or more computer programs, and the one or more computer programs include instructions. When the instructions are executed by the one or more processors, the first electronic device performs the technical solution described in the above embodiments. Its implementation principle and technical effects are similar to those of the related embodiments of the above methods, and will not be repeated here.
[0204] It should be understood that one or more processors in an electronic device can provide... Figure 1 The processor 110 in the electronic device 100 shown; one or more memories in the electronic device may be Figure 1 The internal memory 121 in the electronic device 100 shown.
[0205] This application provides a computer program product that, when run by a first electronic device, causes the first electronic device to execute the technical solutions described in the above embodiments. Its implementation principle and technical effects are similar to those of the related embodiments described above, and will not be repeated here.
[0206] This application provides a readable storage medium containing instructions that, when executed by a first electronic device, cause the first electronic device to perform the technical solution described in the above embodiments. The implementation principle and technical effects are similar and will not be repeated here.
[0207] This application provides a chip for executing instructions. When the chip is running, it executes the technical solutions described in the above embodiments. Its implementation principle and technical effects are similar and will not be repeated here.
[0208] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0209] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0210] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0211] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0212] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0213] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0214] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for eliminating shadows, said method being applied in an electronic device, characterized in that, The method includes: Obtain first video information, which includes multiple frames of images, each of which includes an image of a foreground object; The first interface is displayed, which includes a first 3D model. The first 3D model is a 3D model obtained by reconstructing the image of the foreground object in each frame of the image in three dimensions. The first interface also includes a first control, which is used to adjust the shadow removal effect. The user's first action on the first control was detected; In response to the first operation, a second interface is displayed, which includes a second 3D model. The shadow value of the second 3D model is smaller than the shadow value of the first 3D model, wherein the shadow value is used to characterize the shadow degree of the 3D model. The user's fifth operation on the first control was detected; In response to the fifth operation, a sixth interface is displayed, which includes a fourth 3D model whose shadow value is less than that of the second 3D model.
2. The method according to claim 1, characterized in that, Before displaying the first interface, the method further includes: A third interface is displayed, which includes a thumbnail of the first 3D model. A second user action on the thumbnail was detected; In response to the second operation, the first interface is displayed.
3. The method according to claim 1 or 2, characterized in that, The second interface display includes: The first pose of the second 3D model is displayed on the second interface; The method further includes: A third user action on the second interface was detected; In response to the third operation, a fourth interface is displayed, on which the second pose of the second 3D model is displayed; A fourth user action on the first control was detected; In response to the fourth operation, a fifth interface is displayed, which includes a third 3D model whose shadow values are different from those of the second 3D model.
4. The method according to claim 1 or 2, characterized in that, The first interface for display includes: Extract the image of the foreground object from each of the multiple frames of images; Extract multi-scale feature maps from the image of each foreground object; The multi-scale feature map of the image of each foreground object is decoded, and the shadows of the image of each foreground object are removed to obtain a multi-frame image after shadow removal; Based on the multiple frames of images after shadow removal, the foreground object is reconstructed in three dimensions to obtain the first 3D model; The first interface is displayed.
5. The method according to claim 4, characterized in that, The process of decoding the multi-scale feature map of the image of each foreground object and removing the shadows of the image of each foreground object to obtain a multi-frame image after shadow removal includes: Decode the multi-scale feature map of the image of each foreground object to obtain the shadow effect map of the image of each foreground object; Based on the shadow effect diagram of each foreground object image, the shadow of each foreground object image is removed to obtain the multi-frame image after shadow removal.
6. The method according to claim 4, characterized in that, The extraction of multi-scale feature maps from the image of each foreground object includes: The first video information is downsampled using a convolutional neural network to obtain a multi-scale feature map of the foreground object in each frame of the image; Decoding the multi-scale feature map of the image for each foreground object includes: The multi-scale feature map of the foreground object in each frame of the image is upsampled using a convolutional neural network.
7. The method according to claim 4, characterized in that, The step of extracting the foreground object image from each of the multiple frames includes: The background object image of each frame is removed by the Region of Interest (ROI) algorithm to obtain the foreground object image of each frame.
8. A method for eliminating shadows, said method being applied in an electronic device, characterized in that, The method includes: Obtain first video information, which includes multiple frames of images, each of which includes an image of a foreground object; Display a first interface, which includes a thumbnail corresponding to a first 3D model. The first 3D model is a 3D model obtained by reconstructing the image of the foreground object in each frame of the image. The user's first action on the thumbnail was detected; In response to the first operation, a second interface is displayed, which includes a first 3D model and a first control for adjusting the deshading effect. A second user action on the first control was detected; In response to the second operation, a third interface is displayed, on which the first pose of the second 3D model is displayed, and the shadow value of the second 3D model is less than the shadow value of the first 3D model, wherein the shadow value is used to characterize the shadow degree of the 3D model; A third user action on the third interface was detected. In response to the third operation, a fourth interface is displayed, on which the second pose of the second 3D model is displayed; A fourth user action on the first control was detected; In response to the fourth operation, a fifth interface is displayed, on which a third 3D model is displayed, the shadow values of which are different from those of the second 3D model; The user's fifth operation on the first control was detected; In response to the fifth operation, a sixth interface is displayed, on which a fourth 3D model is displayed, the shadow value of which is less than the shadow value of the third 3D model.
9. The method according to claim 8, characterized in that, The second interface display includes: Extract the image of the foreground object from each of the multiple frames of images; Extract multi-scale feature maps from the image of each foreground object; The multi-scale feature map of the image of each foreground object is decoded, and the shadows of the image of each foreground object are removed to obtain a multi-frame image after shadow removal; Based on the multiple frames of images after shadow removal, the foreground object is reconstructed in three dimensions to obtain the first 3D model; Display the second interface; The process of decoding the multi-scale feature map of the image of each foreground object and removing the shadows of the image of each foreground object to obtain a multi-frame image after shadow removal includes: Decode the multi-scale feature map of the image of each foreground object to obtain the shadow effect map of the image of each foreground object; Based on the shadow effect map of each foreground object image, the shadows of each foreground object image are removed to obtain the multi-frame image after shadow removal; or, The step of extracting the multi-scale feature map of the image of each foreground object includes: The first video information is downsampled using a convolutional neural network to obtain a multi-scale feature map of the foreground object in each frame of the image; Decoding the multi-scale feature map of the image for each foreground object includes: The multi-scale feature maps of the foreground objects in each frame are upsampled using a convolutional neural network; or... The step of extracting the image of the foreground object in each of the multiple frames includes: The background object image of each frame is removed by the Region of Interest (ROI) algorithm to obtain the foreground object image of each frame.
10. An electronic device, characterized in that, include: One or more processors; One or more memories; the one or more memories store one or more computer programs, the one or more computer programs including instructions that, when executed by the one or more processors, cause the electronic device to perform the method as described in any one of claims 1 to 9.
11. A computer-readable storage medium, characterized in that, Includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1 to 9.
12. A computer program product, characterized in that, When the computer program product is run on an electronic device, it causes the electronic device to perform the method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Image processing method and mobile terminal
CN106952235A
Control method and device of sliding control, equipment and storage medium
CN110908571A
Image shadow elimination method based on content perception information
CN111626951A
Face modeling method and device, storage medium and electronic equipment
CN117830503A
Automated Guide For Image Capturing For 3D Model Creation
US20180198976A1