Image processing method, electronic device, chip system, storage medium and computer program product
By merging the target object portion in the target image layer and generating a motion video using camera movement parameters, the problem of time-consuming and labor-intensive manual camera movement is solved, achieving efficient generation of motion video and improving user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HONOR DEVICE CO LTD
- Filing Date
- 2024-10-24
- Publication Date
- 2026-04-24
AI Technical Summary
Manual camera movement is cumbersome, resulting in long video generation times and high manpower costs, which negatively impacts the user experience of video sharing.
By acquiring the 3D image representation of the target image and the region of the target object, the parts belonging to the target object in multiple image layers are merged into a single target image layer to generate a camera movement video. Using camera movement parameters and the updated image layer, the size of the target object is kept constant, and the camera movement effect is simulated through camera intrinsic and extrinsic parameters.
It enables the generation of motion video without manual camera movement, shortening the generation time and reducing manpower costs, reducing the probability of misalignment in motion video, and improving the user experience.
Smart Images

Figure CN121924243A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal technology, and in particular to an image processing method, electronic device, chip system, storage medium and computer program product. Background Technology
[0002] The rise of video-related platforms has led to an increasing number of users sharing videos on these platforms.
[0003] In some implementations, users manually move the camera to capture 3D videos with camera movement effects, i.e., camera movement videos.
[0004] However, manual camera movement is cumbersome, which leads to long production times and high manpower costs for generating video, potentially affecting the user experience of video sharing. Summary of the Invention
[0005] This application provides an image processing method, electronic device, chip system, storage medium, and computer program product, which are applied in the field of terminal technology and can reduce the time and manpower required for generating motion video.
[0006] In a first aspect, embodiments of this application propose an image processing method. The method includes: acquiring a three-dimensional image representation of a target image and a region of a target object in the target image; the three-dimensional image representation includes multiple image layers, each corresponding to a depth; based on the region of the target object, merging the portions belonging to the target object from the multiple image layers into one of the target image layers to obtain updated multiple image layers; and generating a camera movement video corresponding to the target image based on camera movement parameters and the updated multiple image layers. In the camera movement video, the size of the target object remains unchanged.
[0007] This allows for the generation of motion video from a single target image, eliminating the need for manual camera movement and reducing generation time and manpower costs. Furthermore, by merging the portions of the target object from multiple image layers into a single target image layer, updated image layers are obtained. This ensures that in the updated image layers, the target image layer contains all the content of the target object (e.g., the RGB values of all pixels), and the depth of all pixels in the target object is unified to the depth of the target image layer. This reduces the probability of misaligned target objects in motion video generated from updated image layers.
[0008] In one possible implementation, the target image layer satisfies any of the following conditions: When multiple image layers are sorted in descending order of the area of the target object within each image layer to obtain a first image layer sequence, the target image layer is any one of the first n image layers in the first image layer sequence. n is an integer and n≥2. When multiple image layers are sorted in descending order of the number of pixels of the target object within each image layer to obtain a second image layer sequence, the target image layer is any one of the first n image layers in the second image layer sequence. The target image layer contains the target object.
[0009] This allows for the merging of the target object portions from multiple image layers into a single target image layer, resulting in updated image layers. This ensures that in the updated image layers, the corresponding target image layer contains the entire content of the target object, and the depth of all pixels in the target object is unified to the depth of the target image layer. This reduces the probability of misaligned target objects in motion-operated videos generated based on the updated image layers.
[0010] In one possible implementation, the image layer contains the RGB values and opacity of each pixel in the image layer. For a non-target image layer among multiple image layers, the updated image layer corresponding to the non-target image layer includes the RGB values of each pixel of the target object contained in the non-target image layer, and the opacity of each pixel of the target object in the updated image layer corresponding to the non-target image layer is 0. Alternatively, for a non-target image layer among multiple image layers, the RGB value of each pixel of the target object in the updated image layer corresponding to the non-target image layer is 0.
[0011] In this way, the opacity of each pixel of the target object in the updated image layer corresponding to the non-target image layer is 0. This allows the pixels of the target object in the updated image layer corresponding to the non-target image layer to not be rendered during the rendering of multiple updated image layers to generate image frames. Alternatively, the pixels of the target object in the updated image layer corresponding to the non-target image layer can be masked by the pixels of the target object in the updated image layer corresponding to the target image layer. This reduces the probability of the target object appearing out of place in the moving video. The RGB value of each pixel of the target object in the updated image layer corresponding to the non-target image layer is 0. During the rendering of multiple updated image layers to generate image frames, the rendering of pixels with RGB values of 0 in the updated image layer corresponding to the non-target image layer will not adversely affect the target object. This also reduces the probability of the target object appearing out of place in the moving video.
[0012] In one possible implementation, the camera movement parameters include the camera intrinsic and extrinsic parameters corresponding to each camera position among multiple camera positions. The camera intrinsic parameter of the i-th camera position is related to the camera extrinsic parameter of the i-th camera position, where i ≥ 0 and i is an integer.
[0013] In this way, the camera movement parameters include the intrinsic and extrinsic parameters of each camera position across multiple camera locations. When generating camera movement videos using these parameters, the videos can achieve the effect of the camera moving across multiple camera positions. The intrinsic parameters of the i-th camera position are related to the extrinsic parameters of the i-th camera position, allowing the intrinsic parameters to change as the extrinsic parameters change, thus simulating the camera movement effect and giving the video a true camera movement effect.
[0014] In one possible implementation, the camera movement trajectory includes multiple camera positions, with intrinsic camera parameters including focal length and extrinsic camera parameters including object distance; the focal length f at the i-th camera position... i The object distance S between the i-th camera position and the i-th camera position i Satisfying the formula:
[0015]
[0016] Where f0 is the focal length corresponding to the starting position of the camera movement trajectory, or f0 is the focal length corresponding to the target image layer, and S0 is the object distance corresponding to the starting position, or S0 is the object distance corresponding to the target image layer, S i =S0±d i d i Let be the distance between the i-th camera position and the starting position.
[0017] Thus, as Figure 7 The illustrated embodiment describes the focal length f of the i-th camera position along the camera movement trajectory. i and object distance S i Satisfying the formula This allows the size of the target object in the i-th image frame, obtained by rendering multiple updated image layers using the camera intrinsic and extrinsic parameters at the i-th camera position, to be the same as the size of the target object in the target image. This ensures that the size of the target object remains constant in the moving video, thus giving the video a Hitchcock zoom effect.
[0018] In one possible implementation, a camera movement video corresponding to the target image is generated based on camera movement parameters and updated multiple image layers. This includes: rendering the updated multiple image layers using camera intrinsic and extrinsic parameters corresponding to each camera position to obtain image frames corresponding to each camera position; and then synthesizing the image frames corresponding to each camera position into a camera movement video corresponding to the target image.
[0019] This allows for the generation of motion video based on a single target image, reducing the time and manpower required for motion video generation.
[0020] In one possible implementation, the method further includes: displaying an interface of a first application, the interface of which includes a target image; and, in response to an operation on the target image, displaying a first control for instructing the generation of a camera movement video.
[0021] In this way, users can select a target image by operating on the target image in the first application on the electronic device, and generate a camera movement video based on the target image by operating on the first control, which can reduce the time spent generating the camera movement video and the manpower cost of manual camera movement.
[0022] In one possible implementation, the method further includes: displaying an interface of a first application, the interface of which includes a target image and a first control, the first control being used to instruct the generation of a camera movement video.
[0023] In this way, users can generate camera movement videos based on target images by operating the first control, which can reduce the time spent on camera movement video generation and the manpower cost of manual camera movement.
[0024] In one possible implementation, the method further includes: displaying a first pop-up window in response to an operation on the first control, the first pop-up window including information for prompting input of camera movement video-related parameters and an input box for at least one parameter.
[0025] In response to an operation on the first control, a first pop-up window is displayed. This pop-up window includes information prompting the user to input parameters related to the video movement and an input box for at least one parameter. This allows the electronic device to obtain the user-inputted parameter values after the user completes the input and generate the video movement based on those values. This achieves video movement generation from a single target image, reducing generation time and manpower costs while also improving the user experience.
[0026] In one possible implementation, the method further includes: displaying a second pop-up window in response to an operation on the first control, the second pop-up window including information for prompting selection of camera movement video-related parameters and at least one parameter.
[0027] In response to the operation of the first control, a second pop-up window is displayed. This second pop-up window includes information prompting the user to select parameters related to the video movement, and at least one parameter. After the user completes the selection of these parameters, the electronic device obtains the selected parameter values and uses them to generate the video movement video. This achieves video movement video generation based on a single target image, reducing generation time and manpower costs while also improving the user experience.
[0028] In one possible implementation, obtaining a three-dimensional image representation of the target image includes: acquiring depth data of the target image; performing image layering on the target image using the depth data to obtain a three-dimensional image representation of the target image; and obtaining the region of the target object in the target image includes: classifying the objects in the target image layers using a segmentation algorithm to obtain the region of the target object in the target image.
[0029] This process involves acquiring depth data from the target image. The depth data is then used to layer the target image, resulting in a 3D image representation. This allows for the conversion of an RGB image (or the target image) into a 3D image representation, facilitating the subsequent generation of motion-operation video based on this representation. By obtaining the region of the target object within the target image, multiple image layers corresponding to the 3D image representation can be updated, resulting in updated image layers. This reduces the probability of misalignment of the target object in the motion-operation video generated based on the updated image layers.
[0030] Secondly, embodiments of this application provide an image processing apparatus, which may be an electronic device, a chip, or a chip system within an electronic device. The image processing apparatus may include a display unit and a processing unit. When the image processing apparatus is an electronic device, the display unit may be a display screen. The display unit is used to perform display steps to cause the electronic device to implement an image processing method described in the first aspect or any possible implementation of the first aspect. When the image processing apparatus is an electronic device, the processing unit may be a processor. The image processing apparatus may further include a storage unit, which may be a memory. The storage unit is used to store instructions, and the processing unit executes the instructions stored in the storage unit to cause the electronic device to implement an image processing method described in the first aspect or any possible implementation of the first aspect. When the image processing apparatus is a chip or a chip system within an electronic device, the processing unit may be a processor. The processing unit executes the instructions stored in the storage unit to cause the electronic device to implement an image processing method described in the first aspect or any possible implementation of the first aspect. The storage unit can be a storage unit within the chip (e.g., a register, cache, etc.) or a storage unit located outside the chip within the electronic device (e.g., a read-only memory, random access memory, etc.).
[0031] Thirdly, embodiments of this application provide an electronic device including one or more processors and a memory, the memory being coupled to one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, and one or more processors being used to invoke the computer instructions to cause the electronic device to perform the methods described in the first aspect or any possible implementation of the first aspect.
[0032] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program or instructions that, when executed on a computer, cause the computer to perform the methods described in the first aspect or any possible implementation thereof.
[0033] Fifthly, embodiments of this application provide a computer program product including a computer program, which, when run on a computer, causes the computer to perform the methods described in the first aspect or any possible implementation thereof.
[0034] Sixthly, this application provides a chip or chip system including at least one processor and a communication interface. The communication interface and the at least one processor are interconnected via a circuit. The at least one processor is used to run computer programs or instructions to perform the methods described in the first aspect or any possible implementation of the first aspect. The communication interface in the chip can be an input / output interface, pins, or circuits, etc.
[0035] In one possible implementation, the chip or chip system described above in this application further includes at least one memory storing instructions. The memory can be an internal storage unit of the chip, such as a register or cache, or it can be a storage unit of the chip itself (e.g., read-only memory, random access memory, etc.).
[0036] It should be understood that the second to sixth aspects of this application correspond to the technical solutions of the first aspect of this application, and the beneficial effects achieved by each aspect and the corresponding feasible implementation are similar, and will not be repeated here. Attached Figure Description
[0037] Figure 1 This is a schematic diagram of the structure of the electronic device 100 provided in the embodiments of this application;
[0038] Figure 2 A schematic diagram of the software structure of the electronic device 100 provided in the embodiments of this application;
[0039] Figure 3 A scene diagram related to the image processing method provided in the embodiments of this application;
[0040] Figure 4 This is another schematic diagram related to the image processing method provided in the embodiments of this application;
[0041] Figure 5 This is a schematic flowchart of an image processing method provided in an embodiment of this application;
[0042] Figure 6AA schematic diagram illustrating the MPI expression and the updated MPI expression provided in the embodiments of this application;
[0043] Figure 6B A schematic diagram of a first human body mask image and a second human body mask image provided for embodiments of this application;
[0044] Figure 7 A schematic diagram of a camera movement trajectory and camera movement parameters provided for an embodiment of this application;
[0045] Figure 8 This is another scenario illustration related to the image processing method provided in the embodiments of this application;
[0046] Figure 9 This is another scenario diagram related to the image processing method provided in the embodiments of this application. Detailed Implementation
[0047] To facilitate a clear description of the technical solutions in the embodiments of this application, some terms and technologies involved in the embodiments of this application will be briefly introduced below:
[0048] 1. Hitchcock zoom effect
[0049] The Hitchcock zoom effect can also be called the Hitchcock effect.
[0050] The Hitchcock effect can be understood as a dolly zoom effect. For example, the Hitchcock effect occurs when the size of the subject in the frame remains constant, but the perspective of the background changes. Dolly zoom can also be called a push-pull shot or a push-pull camera. It's a special photographic technique used, for example, in film to create visual compression or stretching effects. Dolly zoom involves moving the camera's position (extrinsic parameters) forward or backward while adjusting the camera's intrinsic parameters (such as focal length) to achieve the effect of keeping the subject's size constant in the frame while altering the perspective of the background.
[0051] 2. Other terms
[0052] In the embodiments of this application, terms such as "first" and "second" are used to distinguish identical or similar items with substantially the same function and purpose. For example, "first chip" and "second chip" are used only to distinguish different chips and do not limit their order of execution. Those skilled in the art will understand that terms such as "first" and "second" do not limit the quantity or execution order, and that "first" and "second" do not necessarily imply that they are different.
[0053] It should be noted that, in the embodiments of this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0054] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, a--c, bc, or abc, where a, b, and c can be single or multiple.
[0055] 3. Electronic equipment
[0056] The electronic devices in this application embodiment may include handheld devices with display functions, vehicle-mounted devices, etc. For example, some electronic devices include: mobile phones, tablets, PDAs, laptops, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving vehicles, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, cellular phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), handheld devices with wireless communication capabilities, computing devices or other processing devices connected to wireless modems, in-vehicle devices, wearable devices, terminal devices in 5G networks, or future evolution of public land mobile communication networks. Terminal devices in a network (PLMN), etc., are not limited to this in the embodiments of this application.
[0057] By way of example and not limitation, in this embodiment, the electronic device can also be a wearable device. Wearable devices, also known as wearable smart devices, are a general term for devices that utilize wearable technology to intelligently design and develop everyday wearables, such as glasses, gloves, watches, clothing, and shoes. Wearable devices are portable devices that are worn directly on the body or integrated into the user's clothing or accessories. Wearable devices are not merely hardware devices, but also achieve powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable smart devices include those that are feature-rich, large in size, and can achieve complete or partial functions without relying on a smartphone, such as smartwatches or smart glasses, as well as those that focus on a specific type of application function and require the use of other devices such as smartphones, such as various smart bracelets and smart jewelry for vital sign monitoring.
[0058] Furthermore, in this embodiment of the application, the electronic device can also be a terminal device in the Internet of Things (IoT) system. IoT is an important part of the future development of information technology. Its main technical feature is to connect objects to the network through communication technology, thereby realizing an intelligent network of human-machine interconnection and object-to-object interconnection.
[0059] The electronic devices in the embodiments of this application may also be referred to as: terminal equipment, user equipment (UE), mobile station (MS), mobile terminal (MT), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent, or user device, etc.
[0060] In this embodiment, the electronic device or various network devices include a hardware layer, an operating system layer running on top of the hardware layer, and an application layer running on top of the operating system layer. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and memory (also called main memory). The operating system can be any one or more computer operating systems that implement business processing through processes, such as Linux, Unix, Android, iOS, or Windows. The application layer includes applications such as browsers, address books, word processing software, and instant messaging software.
[0061] In some implementations, users can capture 3D videos with camera movement effects (such as cinematic camera movement effects) by manually moving the camera.
[0062] Taking dolly zoom as an example, it can include pushing the camera forward while zooming out or pulling the camera back while zooming in. The resulting camera movement effect is similar to the Hitchcock zoom effect. Pushing the camera forward while zooming out allows the foreground subject to appear closer to the background while the subject remains the same size. Lenses used in this technique include, for example, camera lenses or camcorder lenses.
[0063] Manual zooming may require manually pushing or pulling the camera or lens while zooming. Therefore, manual camera movement is cumbersome, potentially leading to lengthy video generation times and high manpower costs, which could negatively impact the user experience when sharing videos.
[0064] In view of this, embodiments of this application provide an image processing method. Based on a region of a target object in an image, the portion belonging to the target object in multiple image layers corresponding to the three-dimensional image representation of the image is merged into one target image layer, resulting in updated multiple image layers. Taking the target object as the main subject, based on camera movement parameters and the updated multiple image layers, a camera movement video with the subject's size remaining unchanged can be generated. This enables the generation of a camera movement video with camera movement effects from a single image, eliminating the need for manual camera movement, reducing the time required to generate the video, decreasing manpower costs, and ultimately improving the user experience of video sharing.
[0065] The following is combined Figure 1 and Figure 2 The hardware and software architecture of the electronic devices in the embodiments of this application will be described.
[0066] Figure 1 A schematic diagram of the structure of the electronic device 100 provided in an embodiment of this application is shown.
[0067] Electronic device 100 may include a processor 110, internal memory 121, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, display screen 194, and card interface 195 for user identification module card, etc. The sensor module 180 may include a pressure sensor 180A, gyroscope sensor 180B, accelerometer sensor 180E, touch sensor 180K, ambient light sensor 180L, bone conduction sensor 180M, depth sensor 180N, etc.
[0068] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0069] Processor 110 may include one or more processing units, such as application processors (APs), graphics processing units (GPUs), image signal processors (ISPs), controllers, digital signal processors (DSPs), baseband processors, modem processors, and satellite protocol stacks. These different processing units may be independent devices or integrated into one or more processors.
[0070] The processor 110 may also include a memory for storing instructions and data.
[0071] GPUs can be used to perform mathematical and geometric calculations, as well as for graphics rendering. In some embodiments, electronic devices can use GPUs to generate 3D images, 3D videos, and so on.
[0072] Electronic device 100 implements display functions through GPU, display screen 194, and application processor.
[0073] Internal memory 121 can be used to store executable program code, including instructions. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image and video playback, etc.). The data storage area may store data created during the use of electronic device 100 (such as audio data, image data, phone book, etc.). Furthermore, internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of electronic device 100 by running instructions stored in internal memory 121 and / or instructions stored in memory located within the processor.
[0074] Camera 193 can be used to capture still images or videos.
[0075] The depth sensor 190N can be used to acquire depth information of a scene. In some embodiments, the depth sensor can be disposed in the camera 193, so that the camera can not only acquire images, but also acquire depth data of the images.
[0076] The software system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This embodiment of the invention uses the layered architecture Android system as an example to illustrate the software structure of electronic device 100.
[0077] Figure 2 This is a schematic diagram of the software structure of the electronic device 100 according to an embodiment of this application.
[0078] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.
[0079] The application layer can include a series of application packages.
[0080] like Figure 2 As shown, the application package may include applications such as camera, gallery, call, map, navigation, WLAN, Bluetooth, music, video, SMS and display.
[0081] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.
[0082] like Figure 2 As shown, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.
[0083] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.
[0084] Content providers store and retrieve data, making that data accessible to applications. This data can include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, and more.
[0085] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.
[0086] The phone manager is used to provide communication functions for electronic device 100. For example, it manages call status (including connection and disconnection).
[0087] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.
[0088] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.
[0089] The Android Runtime consists of core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.
[0090] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.
[0091] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0092] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.
[0093] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.
[0094] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.
[0095] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0096] A 2D graphics engine is a graphics engine for 2D drawing.
[0097] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.
[0098] The following example, using the generation of video movement footage, illustrates the workflow of the software and hardware of electronic device 100.
[0099] When the touch sensor 180K receives a touch operation, a corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including touch coordinates, timestamp of the touch operation, etc.). The raw input event is stored in the kernel layer. The application framework layer retrieves the raw input event from the kernel layer and identifies the control corresponding to the input event. Taking a touch click as an example, and the control corresponding to the click as a control in the gallery application representing the generation of a motion video, the application framework layer calls the interface of the application framework layer to transmit the input event to the gallery application. The gallery application obtains the 3D image representation of the target image and the region of the target object in the target image. Based on the region of the target object, the gallery application merges the parts belonging to the target object from multiple image layers corresponding to the 3D image representation into one of the target image layers, resulting in updated multiple image layers. Based on the motion parameters and the updated multiple image layers, the gallery application generates a motion video corresponding to the target image. In this motion video, the size of the target object remains unchanged.
[0100] It should be understood that the target image can be an image pre-captured or stored by an electronic device. In this way, in response to a user's touch operation, the gallery application generates a motion video based on the target image, maintaining the same size as the target object. This allows electronic devices to generate motion video from a single image without requiring manual camera movement, reducing the time required for video generation and minimizing human labor costs.
[0101] The following is combined Figures 3-9 The image processing method provided in the embodiments of this application will be described.
[0102] Figure 3 This illustration shows a scenario related to the image processing method provided in an embodiment of this application.
[0103] like Figure 3As shown in section a, the electronic device can display a desktop, which may include a status bar, the text "xx area", information indicating the time, information indicating the date, icons and text indicating the weather, icons for settings applications, an icon for the gallery application 301, an icon for the calendar application, an icon for the cloud storage application, an icon for the video application, an icon for the music application, an icon for the browser application, an icon for the theme application, a screen switching bar 302, and a quick action bar 303. The quick action bar 303 may include: an icon for the camera application 304, an icon for the SMS application 305, and an icon for the phone application 306. For example, the information indicating the time is 08:00. For example, the information indicating the date is December 3rd, Thursday, the 19th day of the 10th month of the Gengzi year. For example, the text indicating the weather is -8℃, cloudy, good air quality.
[0104] exist Figure 3 On the interface shown in Figure 'a', the user can click icon 301, and the electronic device can display... Figure 3 The interface shown in b. Figure 3 The interface shown in b can include a status bar, the text "Album", a search box, gallery category information 307, controls for photos, controls for albums, controls for time, and controls for discovery.
[0105] The search box can include a search term and text such as "photos, people, locations...". When the data in the gallery is categorized by camera, all photos, and videos, the gallery's categorization information 307 can include information for the camera category, information for the all photos category, and information for the video category. As shown in the gallery's categorization information 307, the camera category information can include: one image from the camera category, the text "camera," and the number of images (or photos) taken by the camera (e.g., 3103). The all photos category information can include: one image from the all photos category, the text "all photos," and the number of images in the all photos category (e.g., 3598). The video category information can include: one image from the video category, the text "video," and the number of videos in the video category (e.g., 98).
[0106] exist Figure 3 On the interface shown in b, users can click on the information for all photo categories in category 307 of the gallery, and the electronic device can display... Figure 3 The interface shown in 'c'. Figure 3 The interface shown in 'c' may include a status bar, a return control, the text "All Photos", a more control, and a selection of images from all photo categories. The selection of images from all photo categories may include image 308.
[0107] exist Figure 3 On the interface shown in 'c', the user can click on image 308, and the electronic device can display... Figure 3 The interface shown by d in the figure. Figure 3 The interface shown as d in the diagram may include a status bar, a return control, time and location information related to image 308, a control for image parameters, image 308 itself, a share control, a favorite control, an edit control, a camera movement control 309, a delete control, and a more control. Image 308 may include a human body 3081, a human body 3082, and an object 3083 in the background.
[0108] Taking human body 3081 as the main body as an example, in Figure 3 On the interface shown as d, the user can click control 309. In response to the user's operation of control 309, the electronic device can use the image processing method provided in this embodiment to generate a motion video based on image 308. The electronic device can display... Figure 3 The interface shown by 'e' in the diagram. Figure 3 The interface shown in 'e' may include a status bar, a return control, and a camera movement video 310. The camera movement video 310 includes a playback control 311. The camera movement video 310 may also include a human body 3081. When the user clicks the control 311, the electronic device can play the camera movement video 310. During the playback of the camera movement video 310, the size of the human body 3081 remains unchanged.
[0109] like Figure 3 As shown in a and b, users can open the Gallery app by clicking the Gallery app icon on the desktop. Figure 3 As shown in b, c, and d, the user can select a pre-stored target image (such as image 308) from the gallery application. Figure 3 As shown in d and e, the user can click on the control representing the motion video. Responding to the user's touch operation on the control representing the motion video, the electronic device can generate a motion video based on the target image while maintaining the same main body size. This allows for the generation of motion video from a pre-stored (or captured) image without the need for manual camera movement, reducing the time required for motion video generation and minimizing manpower costs.
[0110] Figure 4 This illustration shows another scenario related to the image processing method provided in the embodiments of this application.
[0111] Figure 4 The content of the interface shown in 'a' can be found in [reference 1]. Figure 3 The content of the interface shown in 'a' will not be described again here.
[0112] exist Figure 4 On the interface shown in Figure 'a', the user can click the camera application icon 304, and the electronic device can display... Figure 4 The interface shown in b. Figure 4 The interface shown in b includes a status bar, controls for intelligent object recognition, controls for recognizing shooting scenes, controls for turning the flash on or off, controls for filters, controls for setting camera applications, a preview window 401, controls for selecting aperture mode, controls for selecting night scene mode, controls for selecting portrait mode, controls for selecting shooting mode, controls for selecting video mode, controls for selecting more modes, controls for displaying thumbnails 402, controls for taking photos 403, and controls for switching between the front and rear cameras 404. Below the controls for selecting shooting modes is an indicator 407 indicating that the mode is selected. It should be understood that indicator 407, located below the controls for selecting shooting modes, indicates that the shooting mode has been selected. Above the preview window 401 are controls 405 for displaying image parameters and controls 406 for adjusting the shooting magnification. Control 402 can display a thumbnail of the user's last taken photo.
[0113] exist Figure 4 On the interface shown in b, the user can click the control 403 for taking a picture, and the electronic device can display... Figure 4 The interface shown in 'c'. Figure 4 The interface shown in c is the same as Figure 4 The difference between the interfaces shown in b is: Figure 4 The thumbnail displayed in control 402 of the interface shown in c is a thumbnail of the photo taken from the scene displayed in the preview window 401.
[0114] exist Figure 4 On the interface shown in Figure c, the user can click control 402, and the electronic device can display... Figure 4 The interface shown by d in the figure. Figure 4 The interface shown as 'd' includes a status bar and image 308. It should be understood that image 308 is a photograph taken of the scene displayed in preview window 401.
[0115] exist Figure 4 On the interface shown in d, the user can click (e.g., single-click or double-click) image 308, and the electronic device can display... Figure 4 The interface shown in 'e'. Figure 4 The content of the interface shown in 'e' can be found in [reference]. Figure 3 The contents of the interface shown as d in the diagram will not be described again here.
[0116] exist Figure 4On the interface shown in Figure 'e', the user can click on control 309. In response to the user's operation on control 309, the electronic device can use the image processing method provided in this embodiment to generate a motion video based on image 308. The electronic device can display... Figure 4 The interface shown in f is shown in the image. Figure 4 The content of the interface shown in f can be found in [reference]. Figure 3 The content of the interface shown by 'e' in the figure will not be described again here.
[0117] Optionally, in Figure 4 On the interface shown as d, the user can long-press image 308. In response to the user's long-press operation on image 308, the electronic device can use the image processing method provided in this application embodiment to generate a motion video based on image 308. The electronic device can display... Figure 4 The interface shown in f is shown in the image.
[0118] like Figure 4 As shown in a and b, users can open the camera app and take images using it by clicking the camera app icon on the desktop (as in image 308). Figure 4 As shown in c, the user can select the most recently captured image by the camera app as the target image (e.g., image 308). Figure 4 As shown by d, e, and f in the figure, or as shown by Figure 4 As shown in d and f, in response to a user's touch operation on a control representing a motion video or a long press operation on a target image, the electronic device can generate a motion video based on the target image while maintaining the same main body size. This allows for the generation of a motion video from a pre-stored (or captured) image without the need for manual camera movement, thus reducing the time required for motion video generation and minimizing manpower costs.
[0119] Figure 3 This illustrates a scenario where a moving video is generated from a single image using a gallery application and the image processing method described in this application. Figure 4 This illustrates a scene where a camera application and the image processing method of this application generate a moving video based on a single image. It is understood that... Figure 3 and Figure 4 The scenario shown is merely an example and is not intended to limit the scenarios related to the image processing method of this application embodiment. In one possible implementation, other applications (such as video-related applications) and the image processing method of this application embodiment can also be used to generate a camera movement video based on an image, which will not be listed in detail in this application embodiment.
[0120] The following is combined Figure 5 The image processing method provided in the embodiments of this application will be described.
[0121] Figure 5 A schematic flowchart of an image processing method provided in an embodiment of this application is shown.
[0122] like Figure 5 As shown, the image processing method may include:
[0123] S501, Acquire target image
[0124] For example, an electronic device can acquire a target image.
[0125] Using the target image as Figure 3 or Figure 4 For example, in image 308, Figure 3 d in or such Figure 4 As shown in 'e', in response to an operation on control 309, the electronic device can acquire the target image. Or, as... Figure 3 As shown in d, in response to an operation on image 308, the electronic device can acquire the target image.
[0126] It is understood that the target image in the embodiments of this application can be a red-green-blue (RGB) image.
[0127] When an electronic device includes a device for acquiring depth data, the electronic device can acquire depth data of the target image in addition to acquiring the target image. The device for acquiring depth data is, for example, a depth sensor.
[0128] Having obtained the target image and its depth data, the electronic device can perform step S503.
[0129] Optionally, if the electronic device does not include a device for acquiring depth data, the electronic device may perform S502 to perform depth estimation on the target image to obtain depth data of the target image.
[0130] S502, Depth Estimation
[0131] For example, when an electronic device acquires a target image but is not equipped with a device for acquiring depth data, the electronic device can input the target image into a first model. The first model can then output the depth data of the target image.
[0132] The first model can be a model based on a depth prediction network. The first model can be used to predict the depth of the image input to the first model in order to obtain the depth data corresponding to the image.
[0133] Optionally, when the electronic device includes a binocular camera, the target image can be a binocular image. The electronic device can calculate the relative depth data of the image captured by the main camera based on a binocular matching algorithm. The relative depth data of the image captured by the main camera is the depth data of the target image.
[0134] S503. Obtain a three-dimensional image representation of the target image. The three-dimensional image representation can be simply referred to as three-dimensional representation.
[0135] For example, an electronic device can divide a target image into layers (or image layers) based on the depth data of the target image to obtain multiple image layers. Each image layer in the multiple image layers corresponds to a depth, that is, the depth corresponding to each image layer in the multiple image layers is different.
[0136] For example, an electronic device can divide a target image into m image layers based on the depth data of the target image and the number of image layers m. Here, m is an integer greater than 1.
[0137] The number of image layers, m, can be preset. Alternatively, the number of image layers, m, can be determined based on the types of depth values corresponding to multiple pixels in the target image.
[0138] Optionally, before layering the target image, the electronic device can preprocess the depth data of the target image to obtain preprocessed depth data, thereby reducing image noise, etc. Preprocessing may include nonlinear mapping, noise reduction, edge-preserving filtering, and / or hole repair.
[0139] Electronic devices can divide a target image into m layers based on the preprocessed depth data and the number of image layers m.
[0140] In this context, the depth difference between any two adjacent image layers in the m image layers can be the same or different.
[0141] For example, with m=3, multiple image layers including image layer a, image layer b, and image layer c, where the depth of image layer a > the depth of image layer b > the depth of image layer c.
[0142] In m image layers, the depth difference between any two adjacent image layers is the same. For example, the depth difference between image layer a and image layer b is the same as the depth difference between image layer b and image layer c.
[0143] The depth difference between any two adjacent image layers in m image layers is different. For example, the depth difference between image layer a and image layer b is different from the depth difference between image layer b and image layer c.
[0144] For example, taking a preset number of image layers m as an example, the depth data of the target image may include the depth value of each pixel in the target image.
[0145] In the embodiments of this application, the depth value can be simply referred to as depth.
[0146] Electronic devices can obtain the minimum depth S in the depth data of the target image. min and maximum depth S max An electronic device can divide a target image into m image layers based on the depth of each pixel and a preset number of image layers, m in total. The depth S of each of the m image layers is... j The formula can be satisfied:
[0147]
[0148] Where 1 ≤ j ≤ m and j is an integer. Or, j = 1, 2, ..., m.
[0149] The depth S of the image layer j The depth S of pixels from the target image in the image layer k The absolute value of the difference can be less than or equal to the preset depth threshold.
[0150] For example, with m=4, the electronic device can divide the target image into 4 image layers. When sorted by depth from smallest to largest, the depths of these 4 image layers are S0, S1, S2, and S3, respectively. min , and S max .
[0151] In this way, the depth difference between any two adjacent image layers in the m image layers can be the same.
[0152] Optionally, taking the case where the number of image layers m is determined based on the types of depth values corresponding to multiple pixels in the target image as an example, if the electronic device obtains the depth of each pixel in the multiple pixels of the target image, and the depth data of the target image includes m types of depth values, the electronic device can divide the target image into m image layers according to the depth data of the target image, and the depth of each of the m image layers is the same as one of the m depth values.
[0153] Thus, the depth difference between any two adjacent image layers in the m image layers may be different.
[0154] Optionally, when the electronic device obtains the depth data of the target image, it can input the target image and its depth data into a second model. The second model can be a multiplane image (MPI) model. The second model can be used to upscale the target image (RGB image) into multiple planes containing RGB values and alpha values, based on the depth data. Each plane can represent the content of the scene at a certain depth. A plane can be called a planar image or an image layer. Each image layer can include the RGB and alpha values of each pixel in multiple pixels. The alpha value can be used to represent the pixel's transparency or masking information. The range of the alpha value can be 0 to 1 (or 0 to 255).
[0155] Taking an α value ranging from 0 to 1 as an example, an α value of 0 indicates that the pixel is completely transparent. An α value of 1 indicates that the pixel is completely opaque. An α value between 0 and 1 indicates that the pixel is semi-transparent to varying degrees.
[0156] Given the depth data of the target image, the second model can output an MPI representation. MPI representation stands for Multiplanar Image Representation. It can be understood as a specific form of three-dimensional representation.
[0157] This allows for MPI estimation of the target image, resulting in multiple image layers corresponding to the MPI representation. For details on the multiple image layers corresponding to the MPI representation, please refer to... Figure 6A For ease of understanding, Figure 6A This will be described later.
[0158] S504. Merge image layers based on the subject segmentation results to update multiple image layers.
[0159] For example, when an electronic device acquires a target image, the electronic device can input the target image into a third model to obtain a mask image, and can also obtain the bounding box of the object in the foreground of the target image, as well as the position information of the bounding box.
[0160] In this context, the mask image can be understood as the image after segmenting the foreground objects from the target image. The mask image can be a single-channel image where the pixel value of the foreground object region is 1 and the pixel value of the background region is 0. The third model can be a pre-trained model for image segmentation. The third model can include segmentation algorithms. Segmentation algorithms can be semantic segmentation algorithms and / or panoptic segmentation algorithms, etc., that can distinguish different objects in a scene.
[0161] A bounding box can be a rectangle. The bounding box's position information can include the coordinates of its two opposite corners, such as the top-left and bottom-right corners. The bounding box's position information can also include the coordinates of its vertices, width, and height.
[0162] If the foreground of the target image contains k objects that are not connected to each other, the mask image can include k object regions, and the number of bounding boxes for the objects in the foreground of the target image is k. Each bounding box corresponds to one object. Here, k is an integer greater than 1.
[0163] If the foreground of the target image contains k objects that are connected to each other, the mask image can include one object region. The number of bounding boxes of the objects in the foreground of the target image is 1, and the bounding box corresponds to k objects.
[0164] The objects in the foreground of the target image can be human bodies or not. Objects that are not human bodies include, for example, buildings, trees, or animals.
[0165] It should be understood that when the object in the foreground of the target image is a human body, the electronic device inputs the target image into the third model, and the resulting mask image can be called a human body mask image, and the resulting bounding box can be called a human body bounding box.
[0166] For example, if k objects in the foreground of the target image are all human bodies and are not connected to each other, S504 may include S5041-S5043.
[0167] The third model can be a pre-trained model for human segmentation. The third model can include a fourth model and a convolutional neural network (CNN) model based on a binary classification semantic segmentation algorithm. The CNN model can be an encoder-decoder model. The fourth model can be a pre-trained lightweight deep neural network model for human detection. The fourth model can be used to regress all human bounding boxes in the target image.
[0168] S5041. Human body segmentation. Human body segmentation can also be called portrait segmentation.
[0169] For example, taking the case that all k objects in the foreground of the target image are human bodies and that the objects are not connected to each other, the electronic device can input the target image into the third model to obtain the first human body mask image and the bounding boxes of the k human bodies.
[0170] The first human body mask image can include k human body regions. The human body mask image is also called a human body mask image. The human body mask image is a single-channel image where the pixel value of the human body regions is 1, and the pixel value of other regions is 0. The first human body mask image can be found in [reference needed]. Figure 6B As shown, for ease of understanding, Figure 6B This will be described later.
[0171] For example, electronic devices can input target images (e.g., RGB images) into a CNN model and a fourth model, respectively.
[0172] A CNN model can output a first human mask image.
[0173] The fourth model can output k human bounding boxes from the target image, along with the positional information of each bounding box within those k bounding boxes. The positional information of the human bounding boxes can include the coordinates of the two opposite corners of each bounding box. It can also include the vertex coordinates, width, and height of each bounding box.
[0174] It is understandable that the processing of the target image by the CNN model and the processing of the target image by the fourth model can be performed concurrently or sequentially.
[0175] S5042. Select the human body with the largest area as the main subject (i.e., the target object).
[0176] For example, the electronic device can calculate the position information of each human bounding box in the k human bounding boxes to obtain the area of each human bounding box in the k human bounding boxes.
[0177] Electronic devices can identify the human body or target object as the one with the largest area within the bounding box of k human bodies.
[0178] The electronic device can update the first human body mask image based on the target object to obtain a second human body mask image. In the second human body mask image, the pixel value of the target object region is 1, and the pixel value of other regions is 0. That is, the second human body mask image only contains the mask of the target object. The second human body mask image can be found in subsequent sections. Figure 6B The description.
[0179] For example, an electronic device can set the pixel values of non-target human body regions in a first human body mask image to 0 to obtain a second human body mask image.
[0180] In this way, the electronic device can obtain the region of the target object in the target image.
[0181] S5043. Obtain the target image layer from among multiple image layers corresponding to the 3D image representation or MPI representation.
[0182] For example, the electronic device can obtain the image layer corresponding to the target object based on the intersection between the second human body mask image and each of the multiple image layers, and select the target image layer from the image layers of the target object. The image layer corresponding to the target object can be understood as the image layer containing the target object. The image layer corresponding to the target object can also be called the subject-corresponding layer.
[0183] For example, an electronic device can select at least one image layer containing pixels of the target object from multiple image layers corresponding to an MPI representation. Alternatively, the electronic device can select the image layer with the highest distribution of the target object from at least one image layer as the target image layer.
[0184] Optionally, the electronic device can select the image layer with the most distribution of the target object from multiple image layers corresponding to the MPI expression based on the region of the target object in the second human body mask image.
[0185] For example, an electronic device can sort multiple image layers corresponding to an MPI representation according to the area of the target object in each image layer from largest to smallest, obtaining a first image layer sequence. The electronic device can select the first image layer in the first image layer sequence as the target image layer. Optionally, the electronic device can also determine any one of the first n image layers in the first image layer sequence as the target image layer. Here, n is an integer and n≥2.
[0186] Optionally, the electronic device can sort the multiple image layers corresponding to the MPI representation in descending order of the number of pixels of the target object in the image layers to obtain a second image layer sequence. The electronic device can select the first image layer in the second image layer sequence as the target image layer. Optionally, the electronic device can also determine any one of the first n image layers in the second image layer sequence as the target image layer.
[0187] Alternatively, the electronic device may choose any one of the at least one image layer containing the target object as the target image layer.
[0188] It is understandable that the target image layer can be called the main layer corresponding to the subject. The target image layer contains the target object or the pixels of the target object.
[0189] S5044. Merge the image layers corresponding to the target object to update the MPI representation.
[0190] It should be understood that the image layer corresponding to the target object can also be called the layer where the subject is located or the layer corresponding to the subject.
[0191] For example, an electronic device can, based on a region of a target object, merge the portions belonging to the target object from the corresponding image layer into a target image layer, resulting in updated multiple image layers. These updated multiple image layers constitute the updated MPI representation. Further details on the updated multiple image layers can be found later. Figure 6A The description in the text.
[0192] Alternatively, the electronic device can merge the portions of the target object from multiple image layers corresponding to the MPI representation into the target image layer based on the region of the target object, thus obtaining updated multiple image layers.
[0193] For example, an image layer may include the RGB values of each pixel and the transparency (or α value) of each pixel. Among the multiple image layers corresponding to the MPI expression, all image layers other than the target image layer can be called non-target image layers.
[0194] The electronic device can copy the RGB values of the pixels of the target object in the non-target image layer to the corresponding pixels in the target image layer, update the transparency of each pixel of the target object in the target image layer, and set the transparency of the pixels of the target object in the non-target image layer to 0, thus obtaining multiple updated image layers.
[0195] For example, consider an image layer corresponding to a target object comprising q image layers, where q is an integer greater than 1. Updating the transparency of each pixel of the target object in the target image layer can include: obtaining the transparency of the pixel at the target location in each of the q image layers corresponding to the target object, resulting in q transparency values. The maximum transparency value is selected from these q transparency values, and the transparency of the pixel at the target location in the target image layer is updated to the maximum transparency value, thus updating the transparency of the pixel at the target location in the target image layer. The target location can be the position of any pixel of the target object in the target image layer.
[0196] It should be understood that, among the q image layers, the coordinates of the target position in each image layer are the same except for the coordinates in the depth direction.
[0197] Copying the RGB values of pixels representing a target object in a non-target image layer to the corresponding pixels in the target image layer can include: copying the RGB values of the target pixels of the target object in the non-target image layer to the pixels projected onto the target image layer. It should be understood that the coordinates of the pixels projected onto the target image layer are identical to the coordinates of the target pixel, except for the depth direction.
[0198] In this way, in the updated multiple graphics layers, the target image layer contains the RGB values of all pixels of the target object in the updated image layer, and the depth of all pixels of the target object is unified to the depth of the target image layer. The non-target image layer retains (or includes) the RGB values of each pixel of the target object contained in the non-target image layer in the updated image layer, and the transparency of each pixel of the target object in the non-target image layer is 0.
[0199] Optionally, if the electronic device copies the RGB values of the pixels of the target object in the non-target image layer to the corresponding pixels in the target image layer, the electronic device may also set the RGB values of the pixels of the target object in the non-target image layer to 0 to obtain multiple updated image layers.
[0200] Thus, in the updated multiple graphics layers, the target image layer corresponds to the updated image layer containing the RGB values of all (or all) pixels of the target object. In the non-target image layers, the RGB values of each pixel of the target object in the updated image layer are 0, meaning that the RGB values of the target object pixels contained in the non-target image layers are not retained. Optionally, the electronic device can also set the transparency of the pixels of the target object in the non-target image layers to 0.
[0201] In one possible implementation, the multiple image layers corresponding to the MPI expression might not be updated in the manner shown in S504; instead, the motion video might be generated by rendering the multiple image layers corresponding to the MPI expression. In these multiple image layers, due to the inconsistent depth of pixels in the target object, the RGB values of some pixels are contained in the target image layer, while the RGB values of others are contained in a non-target image layer. After rendering the motion video from the multiple image layers corresponding to the MPI expression, there might be a situation where the size of one part of the target object remains unchanged, while the size of another part changes with the background. That is, the target object appears misaligned (or layered) in the motion video.
[0202] Compared to the multiple image layers corresponding to MPI expression, in the updated multiple image layers obtained using the method shown in S504, the target image layer corresponding to the updated image layer contains the RGB values of all pixels of the target object, and the depth of all pixels of the target object is unified to the depth of the target image layer. In the updated image layers corresponding to non-target image layers, the transparency of each pixel of the target object is 0, or the RGB value of each pixel of the target object in the updated image layers corresponding to non-target image layers is 0. Thus, when rendering the updated multiple image layers to generate a motion video, the interference of pixels of the target object in the updated image layers corresponding to non-target image layers on the updated image layers corresponding to the target image layer can be reduced, thereby reducing the probability of misalignment of the target object in the subsequently generated motion video.
[0203] In another possible implementation, the electronic device can assign each pixel of the target object to an image layer based on the depth data of the target image. This reduces the probability of misalignment of the target object in the motion video generated based on multiple image layers corresponding to the MPI expression. However, the inconsistency in the depth of each pixel of the target object is unavoidable. If each pixel of the target object is assigned to an image layer based on the depth data of the target image, the number of image layers corresponding to the MPI expression will be small, resulting in poor motion effects in the motion video generated based on the image layers corresponding to the MPI expression.
[0204] Therefore, in the image processing method provided in this application embodiment, the number of image layers corresponding to the MPI expression can be large, and the image layers corresponding to the MPI expression are updated in the manner shown in S504. This not only reduces the probability of misalignment of the target object in the subsequently generated motion video, but also helps to improve the motion effect of the motion video.
[0205] In the image processing method provided in this application embodiment, the process of rendering multiple updated image layers to generate a camera movement video can be found in S505-S507.
[0206] S505, Obtaining Foreground Depth
[0207] For example, the electronic device can obtain the depth S0 of the target image layer and the focal length f0 of the target image to facilitate the subsequent generation of camera movement parameters.
[0208] Here, the focal length f0 of the target image is the focal length corresponding to the target image layer. The focal length f0 of the target image can be preset.
[0209] Camera movement parameters can be used to generate camera movement videos. These parameters can include the intrinsic and extrinsic parameters of each camera position along a camera movement trajectory.
[0210] In the embodiments of this application, camera intrinsic parameters can be simply referred to as intrinsic parameters, and camera extrinsic parameters can be simply referred to as extrinsic parameters.
[0211] When the camera moves along the z-axis in the world coordinate system, and the z-axis direction can represent the depth direction, the camera movement trajectory can represent the path the camera takes in the depth direction. (See also: [link to camera movement trajectory diagram]). Figure 7 For ease of understanding, Figure 7 This will be described later.
[0212] The camera movement trajectory can include multiple camera positions. Among the multiple camera positions on the camera movement trajectory, the intrinsic parameters of the i-th camera position are related to the extrinsic parameters of the i-th camera position, where i ≥ 0 and i is an integer.
[0213] For details on determining camera movement parameters, please refer to S506.
[0214] S506, Generate camera movement trajectory
[0215] For example, the generation of the camera movement trajectory may include S5061-S5063.
[0216] S5061, Generate camera extrinsic parameters or extrinsic parameter sequences
[0217] Camera extrinsic parameters can include camera position and rotation matrix.
[0218] Camera position can be understood as the specific location of the camera in three-dimensional space, and the camera position can be represented by three-dimensional coordinates in the world coordinate system.
[0219] A rotation matrix can be used to describe the rotation of a camera relative to the world coordinate system, and it can also be used to represent the orientation of the coordinate axes of the camera coordinate system relative to the coordinate axes of the world coordinate system.
[0220] Understandably, taking a camera movement trajectory comprising p camera positions, where p is an integer greater than 1, as an example, the i-th camera position among the p camera positions can be represented as: (x i y i , z i ).
[0221] Given p camera positions, the 0th camera position is (x0, y0, z0), and the camera moves along the z-axis of the world coordinate system, the i-th camera position can be: (x0, y0, z0). i For example, the position of the first camera can be (x0, y0, z1). The position of the second camera can be (x0, y0, z2).
[0222] With the camera positioned on the z-axis, the position of the i-th camera can be: (0, 0, z) iFor ease of understanding, the image processing method provided in this application embodiment will be described below using the example of a camera located on the z-axis. It should be understood that the camera is located on the z-axis, i.e., x... i =0 and y i =0 is merely an example and is not intended to limit the camera position in the embodiments of this application.
[0223] For example, the camera moves along the z-axis of the world coordinate system, the rotation matrix remains unchanged, and the minimum z-axis coordinate of the camera position is z. min The maximum z-axis coordinate of the camera position is z max Let p be the number of camera positions, for example. min To z max The distance between them can be equal to the camera translation distance. Where z min z max Both p and p can be preset.
[0224] Electronic devices can obtain a preset z min z max and p.
[0225] Electronic devices can access z min z max The number of camera positions, p, is calculated to obtain p different camera positions, and then p extrinsic parameters or a sequence of extrinsic parameters composed of p extrinsic parameters are obtained.
[0226] For example, consider a camera moving along the z-axis of the world coordinate system, with the rotation matrix remaining unchanged, and the distance between any two adjacent camera positions in the p camera positions being the same.
[0227] When the camera movement involves pushing the lens (or pushing the camera) along with zooming, z0 can be z min , z i The formula can be satisfied:
[0228]
[0229] When the camera movement involves zooming in (either pulling the lens or the camera) combined with zooming, z0 can be z. max , z i The formula can be satisfied:
[0230]
[0231] When the camera is not rotated relative to the world coordinate system, the rotation matrix can be a 3×3 identity matrix. An example of a 3×3 identity matrix is:
[0232] The information that the camera has not rotated relative to the world coordinate system can be carried in the parameters of the target image when the electronic device acquires the target image. In this way, the electronic device can obtain the rotation matrix based on the information that the camera has not rotated relative to the world coordinate system.
[0233] For example, consider a camera moving along the z-axis of the world coordinate system, with the rotation matrix remaining unchanged, the camera always pointing to the origin, and the camera not rotating relative to the world coordinate system, with z_min = 1.0, z_max = 10.0, and p = 5.
[0234] When the camera movement involves pushing the lens (or pushing the camera) in conjunction with zooming, the sequence of extrinsic parameters obtained by the electronic device can be shown in Table 1 below.
[0235] Table 1 Examples of extrinsic parameter sequences
[0236] Camera position number i External reference 0 Rotation matrix, (0, 0, 1.0) 1 Rotation matrix, (0, 0, 3.25) 2 Rotation matrix, (0, 0, 5.5) 3 Rotation matrix, (0, 0, 7.75) 4 Rotation matrix, (0, 0, 10.0)
[0237] Table 1 shows that when the camera movement involves pushing the lens (or pushing the camera) with zoom, the extrinsic parameter sequence can include p (e.g., 5) tuples. Each tuple can include a rotation matrix and a 3D vector corresponding to a camera position. The 3D vector, also called the translation vector, represents the camera position. As shown in Table 1, camera position index i is 0, indicating the 0th camera position. Camera position index i is 1, indicating the 1st camera position. The extrinsic parameters for the 0th camera position include the rotation matrix and translation vector (0, 0, 1.0). The extrinsic parameters for the 1st camera position include the rotation matrix and translation vector (0, 0, 3.25). The rotation matrix and (0, 0, 1.0) form one tuple of the extrinsic parameter sequence shown in Table 1. The rotation matrix and (0, 0, 3.25) form another tuple of the extrinsic parameter sequence shown in Table 1.
[0238] When the camera movement involves zooming in (or pulling the camera) and using a zoom mechanism, the sequence of extrinsic parameters obtained by the electronic device can be shown in Table 2 below.
[0239] Table 2 Examples of external parameter sequences
[0240] Camera position number i External reference 0 Rotation matrix, (0, 0, 10.0) 1 Rotation matrix, (0, 0, 7.75) 2 Rotation matrix, (0, 0, 5.5) 3 Rotation matrix, (0, 0, 3.25) 4 Rotation matrix, (0, 0, 1.0)
[0241] As shown in Table 2, the extrinsic parameters for the 0th camera position include the rotation matrix and translation vector (0, 0, 10.0), and the extrinsic parameters for the 1st camera position include the rotation matrix and translation vector (0, 0, 7.75). The rotation matrix and (0, 0, 10.0) form one tuple of the extrinsic parameter sequence shown in Table 2. The rotation matrix and (0, 0, 7.75) form another tuple of the extrinsic parameter sequence shown in Table 2.
[0242] Here, the 0th camera position can be called the initial position of the camera or the initial camera position. i can be the index of the extrinsic parameter sequence.
[0243] It should be understood that a sequence of extrinsic parameters in the form of a list, as shown in Table 1 or Table 2, can also be called an extrinsic parameter list.
[0244] S5062, Calculate camera intrinsic parameters (or adjust focal length)
[0245] Camera intrinsics can include focal length. In dolly zoom, by moving the camera and adjusting the focal length, the size of the subject remains constant.
[0246] For example, when an electronic device generates or obtains an extrinsic parameter sequence, the electronic device can calculate the initial object distance and the initial focal length to obtain the focal length of each camera position during the camera's movement, that is, to obtain the focal length of each camera position in the extrinsic parameter sequence. The initial object distance can be understood as the distance from the subject (or object) to the initial position of the camera.
[0247] For example, the initial object distance can be the depth S0 of the target image layer. The initial focal length can be the focal length f0 of the target image.
[0248] Given p camera positions, the focal length f of the i-th camera position out of p camera positions. i Satisfying the formula:
[0249]
[0250] Among them, S i S is the object distance at the i-th camera position. i =S0±d i . d i Let be the distance between the i-th camera position and the initial camera position.
[0251] When the camera movement involves pushing the lens (or pushing the camera) in conjunction with zooming, f i Satisfying the formula:
[0252]
[0253] When the camera movement involves pulling the lens (or pulling the camera) in conjunction with zooming, f i Satisfying the formula:
[0254]
[0255] The camera moves along the z-axis, with the initial camera position at (x0, y0, z0) and the i-th camera position at (x0, y0, z0). i For example, d i =|z i -z0|. || represents the absolute value.
[0256] In this way, the electronic device calculates the focal length at each camera position during the camera's movement based on the initial object distance and initial focal length. The focal length changes with the object distance, keeping the size of the subject constant, thus achieving a dolly zoom effect. This allows the subject to remain the same size while the background appears scaled during subsequent rendering of multiple updated image layers using extrinsic and intrinsic parameters, creating a unique visual experience and simulating the Hitchcock zoom effect in the resulting video footage.
[0257] Among them, formula The derivation can be found in [reference]. Figure 7 For ease of understanding, Figure 7 This will be described later.
[0258] It is understandable that camera extrinsic parameters may also include object distance.
[0259] S5063, Obtain camera parameters or camera parameter sequence
[0260] Camera parameters can include camera extrinsic parameters and camera intrinsic parameters.
[0261] The camera movement parameters in the embodiments of this application can be camera parameters.
[0262] For example, through steps S5061-S5062, the electronic device can obtain the camera extrinsic and intrinsic parameters of each of the p camera positions during the camera movement, so that the electronic device can use the camera extrinsic and intrinsic parameters of each camera position to render the updated multiple image layers, obtain the image frames of each camera position, and thus obtain a video with dollyzoom effect.
[0263] For example, still assuming the camera is located on the z-axis, moves along the z-axis, the rotation matrix remains unchanged, the camera always points to the origin, and the camera does not rotate relative to the world coordinate system, with p=5, the camera movement parameters or camera parameters obtained by the electronic device can be found in Table 3.
[0264] Table 3 Examples of camera movement parameters or camera parameters
[0265] Camera position number i External reference Internal Reference 0 <![CDATA[Rotation matrix, (0, 0, z0)]]> <![CDATA[f0]]> 1 <![CDATA[Rotation matrix, (0, 0, z1)]]> <![CDATA[f1]]> 2 <![CDATA[Rotation matrix, (0, 0, z2)]]> <![CDATA[f2]]> 3 <![CDATA[Rotation matrix, (0, 0, z3)]]> <![CDATA[f3]]> 4 <![CDATA[Rotation matrix, (0, 0, z4)]]> <![CDATA[f4]]>
[0266] Table 3 shows a list-based sequence of camera parameters. This list-based sequence of camera parameters can be called a camera parameter list. As shown in Table 3, the extrinsic parameters for the 0th camera position include the rotation matrix and translation vector (0, 0, z0), and the intrinsic parameters for the 0th camera position include the focal length f0. The camera position number i can also be called the index in the camera parameter list.
[0267] Optionally, S5063 is an optional step. In some implementations, S506 may include S5061-S5062, but not S5063.
[0268] As shown in S506, the electronic device generates a sequence of camera extrinsic parameters (the camera moves along the z-axis) based on preset z_min and z_max, and calculates the focal length corresponding to each of the p camera positions based on the initial object distance S0 and initial focal length f0 of the subject (or target object). The camera extrinsic and intrinsic parameters (such as focal length) corresponding to each of the p camera positions obtained by the electronic device can be used for subsequent rendering of multiple updated image layers, thereby obtaining a Hitchcock effect video, realizing the simulation of a Hitchcock effect video based on a single target image. The Hitchcock effect can be understood as adjusting the focal length while moving the camera, keeping the subject size constant, but changing the perspective effect of the background.
[0269] It is understandable that the number of image frames in a video movement can be the same as the number of camera positions. The p camera positions during the camera's movement can constitute a camera movement trajectory. The distance between the initial camera position and the last camera position among the p camera positions along the direction of camera movement can be called the camera translation distance. The initial camera position can be the starting position of the camera movement trajectory, and the last camera position can be the ending position of the camera movement trajectory.
[0270] The camera panning distance can be preset, or it can be input by the user or selected by the user. The number of image frames in the camera movement video can also be preset, or it can be input by the user or selected by the user. For example, Figure 3 or Figure 4 In the scenario shown, the camera panning distance and the number of image frames used by the electronic device to generate the motion video can be preset. An example of an electronic device generating motion video using user-inputted or user-selected camera panning distance and the number of image frames can be found in [link to example]. Figure 8 and Figure 9 For ease of understanding, Figure 8 and Figure 9 This will be described later.
[0271] It should be understood that one image frame in a motion video corresponds to one camera position. Therefore, if the camera movement trajectory includes p camera positions, the motion video can be synthesized from p image frames.
[0272] S507, Rendering 3D Camera Movement
[0273] For example, the electronic device can use the extrinsic and intrinsic parameters of each camera position obtained in S506 to render the updated multiple image layers and obtain the image frames corresponding to each camera position.
[0274] Electronic devices can synthesize the image frames corresponding to each camera position into a video of camera movement corresponding to the target image.
[0275] Taking the camera movement parameters obtained by the electronic device as shown in Table 3 as an example, the electronic device can use the extrinsic and intrinsic parameters of the initial camera position (i.e., the 0th camera position) to render multiple updated image layers and obtain the image frame corresponding to the initial camera position. The image frame corresponding to the initial camera position can be called the 0th image frame.
[0276] The electronic device can use the extrinsic and intrinsic parameters of the first camera position to render multiple updated image layers, obtaining the image frame corresponding to the first camera position. The image frame corresponding to the first camera position can be called the first image frame.
[0277] The electronic device can use the extrinsic and intrinsic parameters of the second camera position to render multiple updated image layers, obtaining the image frame corresponding to the second camera position. The image frame corresponding to the second camera position can be referred to as the second image frame.
[0278] The electronic device can use the extrinsic and intrinsic parameters of the third camera position to render multiple updated image layers, obtaining the image frame corresponding to the third camera position. The image frame corresponding to the third camera position can be referred to as the third image frame.
[0279] The electronic device can use the extrinsic and intrinsic parameters of the fourth camera position to render multiple updated image layers, obtaining the image frame corresponding to the fourth camera position. The image frame corresponding to the fourth camera position can be referred to as the fourth image frame.
[0280] In this way, the electronic device can obtain image frames corresponding to the positions of the five cameras along the camera movement trajectory.
[0281] Electronic devices can synthesize motion video using image frames corresponding to five camera positions along a camera movement trajectory. For example, the electronic device can concatenate the image frames corresponding to each camera position along the trajectory in order to obtain the motion video. The order of the camera positions along the trajectory represents the sequential order of the camera movements. Examples of such an order are: camera position 0 (or initial camera position), camera position 1, camera position 2, camera position 3, and camera position 4.
[0282] Understandably, when the updated image layer corresponding to the non-target image layer retains the RGB values of each pixel of the target object contained in the non-target image layer, since the transparency of each pixel of the target object in the updated image layer corresponding to the non-target image layer is 0, during the rendering of multiple updated image layers to generate image frames, the pixels of the target object in the updated image layer corresponding to the non-target image layer may not be rendered, or the pixels of the target object in the updated image layer corresponding to the non-target image layer may be masked by the pixels of the target object in the updated image layer corresponding to the target image layer. The pixels of the target object in the updated image layer corresponding to the target image layer may be rendered and will not be masked. This reduces the probability of misaligned target objects in moving video footage.
[0283] When the RGB values of all pixels of the target object in the updated image layer corresponding to a non-target image layer are 0, the rendering of pixels with RGB values of 0 in the updated image layer corresponding to the non-target image layer will not adversely affect the target object during the rendering of multiple updated image layers to generate image frames. This also reduces the probability of misalignment of the target object in moving video footage.
[0284] Understandably, the electronic device uses the extrinsic and intrinsic parameters corresponding to each camera position to render multiple updated image layers, which can be executed concurrently. This reduces the time required to generate motion video.
[0285] It is understandable that the specific implementation principle of generating a video movement video when the camera movement method is zooming in (or zooming out) is similar to that when the camera movement method is zooming in (or zooming in), the specific implementation principle of generating a video movement video will not be repeated here. In the embodiments of this application, the size of the target object in the video movement video generated when the camera movement method is zooming in (or zooming out) can also remain unchanged.
[0286] It should be understood that when an electronic device generates a camera movement video, it can also set the video duration according to a preset video length. The video duration can be preset, or it can be input by the user or selected by the user. For example, Figure 3 and Figure 4 In the scenario shown, the video length of the motion video generated by the electronic device can be preset by the electronic device. Examples of electronic devices generating motion videos using user input or user-selected video lengths can be found later. Figure 8 and Figure 9 The description.
[0287] like Figure 5As shown, the image processing method provided in this application embodiment obtains multiple image layers corresponding to the three-dimensional representation (or MPI expression) of the target image, as well as the region of the target object in the target image. Based on the region of the target object, the parts belonging to the target object in the multiple image layers are merged into the target image layer, and the target image layer is updated to obtain multiple image layers. With the updated multiple image layers, the depth of the target image layer is calculated to obtain the extrinsic and intrinsic parameters of each camera position in the p camera positions on the camera movement trajectory. The updated multiple image layers are rendered using the extrinsic and intrinsic parameters corresponding to each camera position on the camera movement trajectory to obtain image frames corresponding to each camera position. The camera movement video is generated using the image frames corresponding to each camera position. This enables the generation of a camera movement video based on a single target image, eliminating the need for manual camera movement, thus shortening the generation time of the camera movement video and reducing manpower costs.
[0288] Furthermore, in the updated image layers, the target image layer contains the RGB values of all pixels of the target object in the updated image layer, and the depth of all pixels of the target object is unified to the depth of the target image layer. The opacity of all pixels of the target object in the updated image layer corresponding to the target image layer is updated to the maximum opacity. The non-target image layers retain the RGB values of each pixel of the target object contained in the non-target image layer, and the opacity of each pixel of the target object in the updated image layer corresponding to the non-target image layer is 0. Alternatively, the RGB values of each pixel of the target object included in the updated image layer corresponding to the non-target image layer are 0. This reduces the probability of misaligned target objects in the generated video footage rendered from the updated multiple image layers.
[0289] Understandable Figure 5 The image processing method provided in the embodiments of this application is described using an electronic device as the execution subject. In some possible implementations, the execution subject of the image processing method provided in the embodiments of this application may be a camera application or a gallery application in an electronic device, or other applications that can generate video of camera movement based on an image. The embodiments of this application do not specifically limit this.
[0290] Figure 6A The diagram illustrates the MPI expression provided in the embodiments of this application and the updated MPI expression.
[0291] like Figure 6A As shown, with the number of image layers m being 2, the target image being image 308, and image 308 including human body 3081, human body 3082 and object 3083 in the background, the target object being human body 3081 in image 308 is used as an example.
[0292] When the electronic device inputs image 308 into the second model, the second model can output... Figure 6A The MPI representation shown corresponds to image layer a and image layer b. Image layer a may include the RGB values and transparency of a portion of the pixels of human body 3081. Image layer b may include the RGB values and transparency of another portion of the pixels of human body 3081, the RGB values and transparency of the pixels of human body 3082, and the RGB values and transparency of the pixels of object 3083.
[0293] When the electronic device updates image layer a and image layer b corresponding to the MPI expression in the manner shown in S504, the electronic device can obtain Figure 6A The updated MPI representation is shown. The image layers corresponding to the updated MPI representation include: updated image layer a and updated image layer b. Updated image layer a is the updated image layer corresponding to image layer a. Updated image layer b is the updated image layer corresponding to image layer b.
[0294] like Figure 6A As shown, the updated image layer a includes the RGB values and transparency of all pixels of the human body 3081. In the updated image layer b, the RGB value of the human body 3081 pixels is 0, or the transparency of the human body 3081 pixels is 0.
[0295] Figure 6B A schematic diagram of a first human body mask image and a second human body mask image provided in an embodiment of this application is shown.
[0296] like Figure 6B As shown, the target image is still image 308. Image 308 can include human body 3081, human body 3082 and object 3083 in the background. Taking human body 3081 in image 308 as an example, the target object is human body 3081 in image 308.
[0297] Electronic devices can perform human body segmentation on 308 (RGB) images to obtain, for example... Figure 6B The first human body mask image is shown. For example, when the electronic device inputs image 308 into the third model, the third model can output, as shown... Figure 6B The first human body mask image shown. Figure 6B The first human body mask image shown may include human body region 601 and human body region 602.
[0298] When the electronic device updates the first human body mask image based on the human body 3081 (i.e., the target object) or based on the human body region 601, the electronic device can obtain, as follows: Figure 6B The second human body mask image is shown. Figure 6BThe second human body mask image shown may include a human body region 601. The pixel values of the human body region 602 are set to 0.
[0299] Figure 7 This illustration shows a schematic diagram of a camera movement trajectory and camera movement parameters provided in an embodiment of this application.
[0300] Taking the number of camera positions p on the camera movement trajectory as 3, i as 0, 1 or 2, the camera moves along the z-axis, the camera always points to the origin, the rotation matrix remains unchanged, the camera movement method is a push-camera combined with zoom, and the size of the target object is the height of the target object as an example.
[0301] Figure 7 Figure 'a' shows a schematic diagram of the camera movement trajectory and the positions of each camera along that trajectory.
[0302] Figure 7 b in the figure shows the focal length f0, object distance S0, image height u0, and object height x0 corresponding to the initial camera position (the 0th camera position) on the camera movement trajectory. Figure 7 The 'c' in the figure shows the focal length f1, object distance S1, image height u1, and object height x1 corresponding to the first camera position on the camera movement trajectory. Figure 7 In the figure, d represents the focal length f2, object distance S2, image height u2, and object height x2 corresponding to the second camera position on the camera movement trajectory.
[0303] Image height can be understood as the height of the object being photographed (such as the target object or subject) on the imaging plane. Object height can be understood as the actual height of the object being photographed. The z-axis direction can be understood as the depth direction.
[0304] like Figure 7 As shown in a, therefore,
[0305] Similarly, such as Figure 7 As shown in b, therefore,
[0306] Similarly, such as Figure 7 As shown in c, therefore,
[0307] Since the size of the target object needs to remain constant throughout the video movement, we have: u0 = u1 = u2
[0308] It should be understood that the actual size (such as height) of the subject is constant; therefore, x0 = x1 = x2
[0309] Therefore, Right now therefore,
[0310] Thus, the focal length f of the i-th camera position along the camera movement trajectory i and object distance S i Satisfying the formula This allows for the rendering of multiple updated image layers using the extrinsic and intrinsic parameters of each of the p camera positions to achieve a Hitchcock effect in the resulting video.
[0311] Figure 8 This illustration shows yet another scenario related to the image processing method provided in the embodiments of this application.
[0312] by Figure 3 Image 308 in the example is the target image. Figure 8 The interface shown in 'a' is the same as... Figure 3 The interface shown in d is the same. Displayed on electronic devices. Figure 8 On the interface shown in Figure 'a', the user can click on control 309. In response to the user's action on control 309, the electronic device can display... Figure 8 The interface shown in b is as follows. Figure 8 The interface shown in b can include a pop-up window 801. The pop-up window 801 can include information for prompting input of camera movement video parameters, an input box for at least one parameter, and a control 805 indicating confirmation. Figure 8 The rest of the content of the interface shown in b can be found in [reference]. Figure 8 The content of the interface shown in 'a' will not be described again.
[0313] This includes information prompting the user to input parameters related to the camera movement video, such as the text "Please enter parameters related to the camera movement video." There are at least one input box for each parameter, for example: input box 802, input box 803, and input box 804. The input boxes are surrounded by names indicating the parameter category to which the input value belongs. For example, input box 802 has the parameter name "Number of Image Frames" on its left, indicating that the number of image frames can be entered in input box 802. Input box 803 has the parameter name "Camera Pan Distance" on its left, indicating that the camera pan distance can be entered in input box 803. Input box 804 has the parameter name "Video Duration" on its left, indicating that the video duration can be entered in input box 804.
[0314] exist Figure 8 On the interface shown in b, users can enter parameter values for camera movement video-related parameters in the input box of pop-up window 801. The electronic device can display... Figure 8 The interface shown in c is shown in the image. Figure 8 The interface shown in c is the same as Figure 8 The difference between the interfaces shown in b is that: Figure 8In the interface shown in c, the parameter value 10 was entered in input box 802, the parameter value 20 was entered in input box 803, and the parameter value 5 was entered in input box 804.
[0315] exist Figure 8 On the interface shown in Figure c, the user can click control 805. In response to the user's action on control 805, the electronic device can obtain the parameter value displayed in the input box of pop-up window 801. The parameter value displayed in the input box of pop-up window 801 represents the parameter value confirmed by the user. The electronic device can use the parameter value in the input box of pop-up window 801 to generate a motion video. The electronic device can display... Figure 8 The interface shown in d is shown in the image. Figure 8 The content of the interface shown in d can be found in [reference]. Figure 3 The contents of the interface shown in 'e' will not be described again here.
[0316] like Figure 8 As shown in a, b, c, and d, in response to user interaction with the controls representing the motion video, the electronic device can display a pop-up window prompting the user to input parameters related to the motion video. In response to the user's interaction with the controls representing confirmation after inputting the motion video parameters, the electronic device can obtain the user-input parameter values and generate the motion video using these values. This achieves the generation of motion video based on a single target image, reducing the time required for motion video generation, decreasing manpower costs, and improving the user experience.
[0317] Figure 9 This illustration shows yet another scenario related to the image processing method provided in the embodiments of this application.
[0318] Still with Figure 3 Image 308 in the example is the target image. Figure 9 The interface shown in 'a' is the same as... Figure 3 The interface shown in d is the same.
[0319] In electronic device display Figure 9 On the interface shown in Figure 'a', the user can click on control 309. In response to the user's action on control 309, the electronic device can display... Figure 9 The interface shown in b is shown in the image. Figure 9 The interface shown in b is the same as Figure 8 The difference between the interfaces shown in b is: Figure 9 The interface shown in b includes pop-up 901 but does not include pop-up 801.
[0320] The pop-up window 901 may include information prompting the user to select camera movement video-related parameters, a selection box for at least one parameter, and a confirmation control 905. The information prompting the user to select camera movement video-related parameters may include the text "Please select camera movement video-related parameters." The selection box for at least one parameter may include selection boxes 902, 903, and 904. The selection box is surrounded by the name of the camera movement video-related parameter to indicate the parameter category to which the selected parameter value belongs. The selection box may display the parameter value and a drop-down control. The drop-down control may include control 906. The user can click the drop-down control in the selection box to select a parameter value. The parameter value in the selection box may be the default parameter value of the electronic device or the parameter value previously selected by the user. Figure 9 As shown in b, the parameter name "Number of Image Frames" is on the left side of selection box 902, indicating that the parameter value in selection box 902 (e.g., 5) belongs to the number of image frames. The parameter name "Camera Pan Distance" is on the left side of selection box 903, indicating that the parameter value in selection box 903 (e.g., 20) belongs to the camera pan distance. The parameter name "Video Duration" is on the left side of selection box 904, indicating that the parameter value in selection box 904 (e.g., 5) belongs to the video duration.
[0321] exist Figure 9 On the interface shown in b, the user can click the drop-down control 906 in the selection box 902, and the electronic device can display... Figure 9 The interface shown in 'c'. Figure 9 The interface shown in c is the same as Figure 9 The difference between the interfaces shown in b is: Figure 9 The interface shown in Figure 'c' also includes a drop-down menu 907, which contains parameter values 10, 15, 20, and 30. It should be understood that the parameter values in drop-down menu 907 all pertain to the number of image frames.
[0322] exist Figure 9 On the interface shown in Figure c, the user can click on parameter value 30 in drop-down menu 907 to select the parameter value for the number of image frames. The electronic device can display... Figure 9 The interface shown in d is shown in the figure. Figure 9 The interface shown in d is the same as Figure 9 The difference between the interface shown in 'c' and the interface shown in 'c' is: Figure 9 In the interface shown by d, the parameter value displayed in selection box 902 is 30.
[0323] exist Figure 9On the interface shown as d, the user can click control 905. In response to the user's action on control 905, the electronic device receives the parameter value displayed in the selection box of pop-up window 901. The parameter value displayed in the selection box of pop-up window 901 represents the parameter value selected by the user. The electronic device can use the parameter value displayed in the selection box of pop-up window 901 to generate a motion video. The electronic device can display... Figure 9 The interface shown in 'e'. Figure 9 The content of the interface shown in 'e' can be found in [reference]. Figure 3 The contents of the interface shown in 'e' will not be described again here.
[0324] like Figure 9 As shown in Figures a, b, c, d, and e, in response to user interaction with the controls representing the motion video, the electronic device can display a pop-up window prompting the user to select parameters related to the motion video. In response to the user's interaction with the control representing confirmation after selecting the parameters, the electronic device can obtain the selected parameter values and generate the motion video using these values. This achieves the generation of motion video based on a single target image, reducing the time required for motion video generation, decreasing manpower costs, and improving the user experience.
[0325] It should be understood that the starting position of the camera movement trajectory (e.g., the 0th camera position) can be preset, and the camera movement translation distance d is obtained after user input or selection. max In this case, the electronic device can determine the position of the ending point of the camera movement trajectory. For example, with the 0th camera position as (0,0,z)... min For example, the electronic device can obtain the position of the ending position of the camera movement trajectory as (0,0,z). max ), and z max =z min +d max .
[0326] This application also provides an image processing method, which may include:
[0327] Obtain a 3D image representation of the target image, as well as the region of the target object in the target image. The 3D image representation includes multiple image layers, each corresponding to a depth.
[0328] Based on the region of the target object, the parts belonging to the target object in multiple image layers are merged into one of the target image layers to obtain updated multiple image layers.
[0329] Based on the camera movement parameters and updated image layers, a camera movement video corresponding to the target image is generated. The size of the target object remains unchanged throughout the camera movement video.
[0330] For example, the target image can be image 308 as described above, and the target object can be human body 3081 in image 308. The specific implementation principle for obtaining the three-dimensional image representation of the target image can be found in the specific implementation principle of S503. The specific implementation principle for obtaining the region of the target object in the target image can be found in the specific implementation principles of S5041 and / or S5042. The region of the target object can be... Figure 6B The human body region 601 corresponding to human body 3081 in the first or second human body mask image. Based on the target object region, the portions belonging to the target object from multiple image layers are merged into one of the target image layers to obtain the updated multiple image layers. For the specific implementation principle of this process, please refer to the implementation principles of S5043-S5044. For the specific implementation principle of generating the camera movement video corresponding to the target image based on the camera movement parameters and the updated multiple image layers, please refer to the implementation principles of S505-S507.
[0331] In this way, multiple image layers corresponding to the 3D image representation of the target image are obtained, as well as the region of the target object in the target image. Based on the region of the target object, the parts belonging to the target object in the multiple image layers are merged into the target image layer, and the target image layer is updated to obtain multiple image layers. Based on the camera movement parameters and the updated multiple image layers, a camera movement video corresponding to the target image is generated. This enables the generation of a camera movement video based on a single target image, eliminating the need for manual camera movement, thus shortening the generation time of the camera movement video and reducing manpower costs.
[0332] Furthermore, by merging the portions belonging to the target object from multiple image layers into a single target image layer, updated image layers are obtained. This ensures that in the updated image layers, the target image layer contains all the content of the target object (e.g., the RGB values of all pixels of the target object), and the depth of all pixels of the target object is unified to the depth of the target image layer. This reduces the probability of misaligned target objects in motion-operated videos generated based on the updated image layers.
[0333] Optionally, the target image layer satisfies any of the following conditions:
[0334] When multiple image layers are sorted in descending order of the area of the target object within each image layer to obtain the first image layer sequence, the target image layer is any one of the first n image layers in the first image layer sequence. n is an integer and n≥2.
[0335] When multiple image layers are sorted from largest to smallest according to the number of pixels of the target object in the image layer to obtain the second image layer sequence, the target image layer is any one of the first n image layers in the second image layer sequence.
[0336] The target image layer contains the target object.
[0337] This allows for the merging of the target object portions from multiple image layers into a single target image layer, resulting in updated image layers. This ensures that in the updated image layers, the corresponding target image layer contains the entire content of the target object, and the depth of all pixels in the target object is unified to the depth of the target image layer. This reduces the probability of misaligned target objects in motion-operated videos generated based on the updated image layers.
[0338] Optionally, the image layer contains the RGB values and transparency of each pixel in the image layer.
[0339] For a non-target image layer in multiple image layers, the updated image layer corresponding to the non-target image layer includes the RGB values of each pixel of the target object contained in the non-target image layer, and the opacity of each pixel of the target object in the updated image layer corresponding to the non-target image layer is 0. Alternatively,
[0340] For non-target image layers in multiple image layers, the RGB values of each pixel of the target object in the updated image layer corresponding to the non-target image layer are 0.
[0341] In this way, the opacity of each pixel of the target object in the updated image layer corresponding to the non-target image layer is 0. This allows the pixels of the target object in the updated image layer corresponding to the non-target image layer to not be rendered during the rendering of multiple updated image layers to generate image frames. Alternatively, the pixels of the target object in the updated image layer corresponding to the non-target image layer can be masked by the pixels of the target object in the updated image layer corresponding to the target image layer. This reduces the probability of the target object appearing out of place in the moving video. The RGB value of each pixel of the target object in the updated image layer corresponding to the non-target image layer is 0. During the rendering of multiple updated image layers to generate image frames, the rendering of pixels with RGB values of 0 in the updated image layer corresponding to the non-target image layer will not adversely affect the target object. This also reduces the probability of the target object appearing out of place in the moving video.
[0342] Optionally, the camera movement parameters include the camera intrinsic parameters and camera extrinsic parameters corresponding to each camera position among multiple camera positions. The camera intrinsic parameters of the i-th camera position are related to the camera extrinsic parameters of the i-th camera position, where i ≥ 0 and i is an integer.
[0343] In this way, the camera movement parameters include the intrinsic and extrinsic parameters of each camera position across multiple camera locations. When generating video using these parameters, the video can simulate the effect of a camera moving across multiple camera positions. The intrinsic parameters of the i-th camera position are related to its extrinsic parameters, allowing the intrinsic parameters to change with the extrinsic parameters, thus simulating the camera movement effect and giving the video a true camera movement effect.
[0344] Optionally, the camera movement trajectory includes multiple camera positions, with intrinsic camera parameters including focal length and extrinsic camera parameters including object distance; the focal length f at the i-th camera position. i The object distance S between the i-th camera position and the i-th camera position i Satisfying the formula:
[0345]
[0346] Where f0 is the focal length corresponding to the starting position of the camera movement trajectory, or f0 is the focal length corresponding to the target image layer, and S0 is the object distance corresponding to the starting position, or S0 is the object distance corresponding to the target image layer, S i =S0±d i d i Let be the distance between the i-th camera position and the starting position.
[0347] Thus, as Figure 7 The illustrated embodiment describes the focal length f of the i-th camera position along the camera movement trajectory. i and object distance S i Satisfying the formula This allows the size of the target object in the i-th image frame, obtained by rendering multiple updated image layers using the camera intrinsic and extrinsic parameters at the i-th camera position, to be the same as the size of the target object in the target image. This ensures that the size of the target object remains constant in the moving video, thus giving the video a Hitchcock zoom effect.
[0348] Optionally, based on camera movement parameters and updated multiple image layers, a camera movement video corresponding to the target image is generated, including:
[0349] The updated image layers are rendered using the camera intrinsic and extrinsic parameters corresponding to each camera position to obtain the image frames corresponding to each camera position.
[0350] The image frames corresponding to each camera position are combined to form the camera movement video corresponding to the target image.
[0351] For example, the specific implementation principle of the embodiments of this application can be found in the specific implementation principle of S507. In this way, the generation of camera movement video can be realized.
[0352] Optionally, the method further includes:
[0353] The interface of the first application is displayed, and the interface of the first application includes the target image.
[0354] In response to an operation on the target image, a first control is displayed, which is used to instruct the generation of a motion video.
[0355] For example, the first application could be Figure 3 The gallery application in the illustrated embodiment or Figure 4 Camera application in the illustrated embodiment.
[0356] If the first application is a gallery application, the interface of the first application may include Figure 3 The interface shown in 'c'. The target image can be... Figure 3 Image 308 in the interface shown in Figure c. The first control can be... Figure 3 The control shown as d in the figure is 309.
[0357] If the first application is a camera application, the interface of the first application may include Figure 4 The interface shown as 'd' in the diagram. The target image can be... Figure 4 Image 308 in the interface shown by d. The first control can be... Figure 4 The control shown as 'e' in the figure is 309.
[0358] In this way, users can select a target image by operating on the target image in the first application on the electronic device, and generate a camera movement video based on the target image by operating on the first control, which can reduce the time spent generating the camera movement video and the manpower cost of manual camera movement.
[0359] Optionally, the method further includes:
[0360] The interface of the first application is displayed. The interface of the first application includes a target image and a first control, which is used to instruct the generation of camera movement video.
[0361] For example, in the case where the first application is a gallery application, the interface of the first application may include Figure 3 In the interface shown by d, the first control can be Figure 3 Control 309 is shown as 'd' in the diagram. When the first application is a camera application, the interface of the first application can be... Figure 4 In the interface shown by 'e', the first control can be Figure 4 The control shown as 'e' in the figure is 309.
[0362] In this way, users can generate camera movement videos based on target images by operating the first control, which can reduce the time spent on camera movement video generation and the manpower cost of manual camera movement.
[0363] Optionally, the method further includes:
[0364] In response to an operation on the first control, a first pop-up window is displayed. The first pop-up window includes information for prompting input of parameters related to the camera movement video and an input box for at least one parameter.
[0365] For example, the first control can be Figure 8 Control 309, shown as 'a' in the diagram. The first pop-up can be... Figure 8 The pop-up window 801 shown in b is shown in the image.
[0366] In response to an operation on the first control, a first pop-up window is displayed. This pop-up window includes information prompting the user to input parameters related to the video movement and an input box for at least one parameter. This allows the electronic device to obtain the user-inputted parameter values after the user completes the input and generate the video movement based on those values. This achieves video movement generation from a single target image, reducing generation time and manpower costs while also improving the user experience.
[0367] Optionally, the method further includes:
[0368] In response to an operation on the first control, a second pop-up window is displayed. The second pop-up window includes information for prompting the selection of camera movement video-related parameters and at least one parameter.
[0369] For example, the first control can be Figure 9 Control 309, shown as 'a' in the diagram. The second pop-up can be... Figure 9 The pop-up window 901 shown in b is shown in the image.
[0370] In response to the operation of the first control, a second pop-up window is displayed. This second pop-up window includes information prompting the user to select parameters related to the video movement, and at least one parameter. After the user completes the selection of these parameters, the electronic device obtains the selected parameter values and uses them to generate the video movement video. This achieves video movement video generation based on a single target image, reducing generation time and manpower costs while also improving the user experience.
[0371] Optionally, obtaining a three-dimensional image representation of the target image includes:
[0372] Obtain the depth data of the target image.
[0373] The target image is layered using depth data to obtain a three-dimensional image representation of the target image.
[0374] Obtain the region of the target object in the target image, including:
[0375] A segmentation algorithm is used to classify objects in the target image layer to obtain the regions of the target objects in the target image.
[0376] For example, the specific implementation principle of acquiring the depth data of the target image can be found in S501 or the specific implementation principle of S501-S502. The specific implementation principle of using the depth data to perform image layering on the target image to obtain a three-dimensional image representation of the target image can be found in the specific implementation principle of S503. In this way, the RGB image (or target image) is converted into a three-dimensional image representation, facilitating the subsequent generation of motion video based on the three-dimensional image representation.
[0377] For example, the specific implementation principle of classifying objects in the target image layer using a segmentation algorithm to obtain the region of the target object in the target image can be found in S5041 or S5041-S5042. By obtaining the region of the target object in the target image, multiple image layers corresponding to the 3D image representation can be updated to obtain an updated image layer, thereby reducing the probability of misalignment of the target object in the video generated based on the updated image layer.
[0378] It should be noted that the module names involved in the embodiments of this application can all be defined as other names, as long as they can achieve the function of each module, and no specific restrictions are placed on the module names.
[0379] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0380] The image processing method of the present application embodiments has been described above. The apparatus for performing the above method provided in the present application embodiments is described below. Those skilled in the art will understand that the methods and apparatus can be combined with and referenced by each other, and the related apparatus provided in the present application embodiments can perform the steps in the above image processing method.
[0381] The image processing method provided in this application can be applied to electronic devices with communication functions. The electronic devices include terminal devices, and the specific device form of the terminal devices can be referred to the above-described related descriptions, which will not be repeated here.
[0382] This application provides an electronic device, which includes one or more processors and a memory; the memory is coupled to one or more processors and is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the electronic device to perform the above-described method.
[0383] This application provides a chip. The chip includes a processor, which is used to call a computer program in memory to execute the technical solutions in the above embodiments. Its implementation principle and technical effects are similar to those in the related embodiments described above, and will not be repeated here.
[0384] This application provides a chip system. The chip system includes at least one processor and a communication interface, the communication interface and the at least one processor being interconnected via a circuit, and the at least one processor being used to run computer programs or instructions to perform the above-described method.
[0385] This application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, it implements the methods described above. The methods described in the above embodiments can be implemented wholly or partially by software, hardware, firmware, or any combination thereof. If implemented in software, the functionality can be stored as one or more instructions or code on or transmitted over the computer-readable medium. The computer-readable medium can include computer storage media and communication media, and can also include any medium that can transfer a computer program from one place to another. The storage medium can be any target medium accessible by a computer.
[0386] In one possible implementation, a computer-readable medium may include RAM, ROM, compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage or other magnetic storage devices, or any other medium targeted to carry or to store the required program code in the form of instructions or data structures, and accessible by a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. As used herein, disks and optical discs include optical discs, laser discs, optical discs, Digital Versatile Discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs optically reproduce data using lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0387] This application provides a computer program product, which includes a computer program that, when run, causes a computer to perform the above-described method.
[0388] This application describes embodiments of methods, apparatus (systems), and computer program products according to embodiments of this application with reference to flowchart illustrations and / or block diagrams. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, special-purpose computer, embedded processor, or other programmable device to produce a machine, such that the instructions, which execute via the processing unit of the computer or other programmable data processing device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0389] The above specific embodiments further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solution of the present invention should be included within the scope of protection of the present invention.
Claims
1. An image processing method, characterized in that, include: A three-dimensional image representation of a target image and the region of the target object in the target image are obtained. The three-dimensional image representation includes multiple image layers, each corresponding to a depth. Based on the region of the target object, the portions belonging to the target object in the plurality of image layers are merged into one of the target image layers to obtain updated plurality of image layers; Based on the camera movement parameters and the updated multiple image layers, a camera movement video corresponding to the target image is generated; in the camera movement video, the size of the target object remains unchanged.
2. The method according to claim 1, characterized in that, The target image layer satisfies any of the following conditions: When the plurality of image layers are sorted in descending order of the area of the target object in the image layer to obtain a first image layer sequence, the target image layer is any one of the first n image layers in the first image layer sequence; n is an integer and n≥2; When the plurality of image layers are sorted in descending order of the number of pixels of the target object in the image layer to obtain a second image layer sequence, the target image layer is any one of the first n image layers in the second image layer sequence; The target image layer contains the target object.
3. The method according to claim 1 or 2, characterized in that, The image layer contains the RGB values and transparency of each pixel in the image layer; For the non-target image layer in the plurality of image layers, the updated image layer corresponding to the non-target image layer includes the RGB values of each pixel of the target object contained in the non-target image layer, and the transparency of each pixel of the target object in the updated image layer corresponding to the non-target image layer is 0; or, For the non-target image layer in the plurality of image layers, the RGB value of each pixel of the target object in the updated image layer corresponding to the non-target image layer is 0.
4. The method according to any one of claims 1-3, characterized in that, The camera movement parameters include camera intrinsic parameters and camera extrinsic parameters corresponding to each camera position among multiple camera positions. The camera intrinsic parameters of the i-th camera position are related to the camera extrinsic parameters of the i-th camera position, where i ≥ 0 and i is an integer.
5. The method according to claim 4, characterized in that, The camera movement trajectory includes the multiple camera positions, and the camera intrinsic parameters include focal length, while the camera extrinsic parameters include object distance. The focal length f at the i-th camera position i The object distance S at the i-th camera position i Satisfying the formula: Where f0 is the focal length corresponding to the starting position of the camera movement trajectory, or f0 is the focal length corresponding to the target image layer, and S0 is the object distance corresponding to the starting position, or S0 is the object distance corresponding to the target image layer, S i =S0±d i d i The distance between the i-th camera position and the starting position is given.
6. The method according to claim 4 or 5, characterized in that, The step of generating the camera movement video corresponding to the target image based on the camera movement parameters and the updated multiple image layers includes: The updated multiple image layers are rendered using the camera intrinsic and extrinsic parameters corresponding to each camera position to obtain the image frames corresponding to each camera position. The image frames corresponding to each camera position are combined to form the camera movement video corresponding to the target image.
7. The method according to any one of claims 1-6, characterized in that, Also includes: The interface of the first application is displayed, and the interface of the first application includes the target image; In response to an operation on the target image, a first control is displayed, the first control being used to instruct the generation of a camera movement video.
8. The method according to any one of claims 1-6, characterized in that, Also includes: The interface of the first application is displayed. The interface of the first application includes the target image and a first control, which is used to instruct the generation of a camera movement video.
9. The method according to claim 7 or 8, characterized in that, Also includes: In response to an operation on the first control, a first pop-up window is displayed, the first pop-up window including information for prompting input of camera movement video-related parameters and an input box for at least one parameter.
10. The method according to claim 7 or 8, characterized in that, Also includes: In response to an operation on the first control, a second pop-up window is displayed, the second pop-up window including information for prompting the selection of camera movement video-related parameters and at least one parameter.
11. The method according to any one of claims 1-10, characterized in that, The acquisition of the three-dimensional image representation of the target image includes: Obtain the depth data of the target image; The target image is layered using the depth data to obtain a three-dimensional image representation of the target image; The step of obtaining the region of the target object in the target image includes: A segmentation algorithm is used to classify the objects in the target image layer to obtain the regions of the target objects in the target image.
12. An electronic device, characterized in that, The electronic device includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, and the one or more processors invoking the computer instructions to cause the electronic device to perform the method as described in any one of claims 1 to 11.
13. A chip system, characterized in that, The chip system is applied to an electronic device, the chip system including one or more processors, the one or more processors being used to invoke computer instructions to cause the electronic device to perform the method as described in any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1 to 11.
15. A computer program product, characterized in that, The computer program product includes computer program code that, when run on an electronic device, causes the electronic device to perform the method as described in any one of claims 1 to 11.