Video processing method and device, electronic equipment and computer readable medium
By splicing the original video and the target subject image into the target video, and segmenting the target subject in each image frame as the foreground and background superimposed display, the complexity of synchronous playback in the depth of field video solution is solved, and efficient video display effect is achieved.
Patent Information
- Application Number
- CN202510920237.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-09-05
AI Technical Summary
In the existing depth of field video solution, the main video and the original video are two independent video streams, and precise synchronous playback is required, resulting in increased player complexity and prone to synchronization errors, affecting the viewing experience.
The original video and the predetermined target subject image are spliced into the target video. The target subject is divided into each image frame as the foreground and background superimposed display. The synchronous display of the target subject and the original video is achieved through the stitched image frame, reducing the decoding and rendering complexity.
The automatic synchronous display of the target subject and the original video is achieved, reducing the complexity of the player, reducing synchronization errors and resource consumption, and improving user experience.
Smart Images

Figure CN120602723A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of mobile terminal technology, and more specifically, to a video processing method, device, electronic device, and computer-readable medium. Background Art
[0002] Existing depth-of-field video solutions, which separate and process the subject and background in a video and combine them with advanced video synthesis and synchronization technologies, have gradually become an innovative way to enhance the user's visual experience. Specifically, this solution extracts the main subject from the video and stores it as a separate video stream, then uses synthesis technology to overlay it on top of the original video, achieving a dynamic depth-of-field effect.
[0003] However, in current solutions, the main video and the original video are two independent video streams. Therefore, precise synchronization of the two videos is required during playback. This places high demands on the player, especially ensuring that the timelines of the two videos are completely consistent. This not only increases the complexity of the player but also easily leads to synchronization errors, affecting the viewing experience. Summary of the Invention
[0004] This application proposes a video processing method, device, electronic device and computer-readable medium to improve the above-mentioned defects.
[0005] In a first aspect, the present application provides a video processing method, which is applied to an electronic device, and the method includes: obtaining an original video, wherein the original video includes multiple first image frames; obtaining a target video corresponding to the original video, wherein the target video includes multiple second image frames spliced together based on a predetermined target subject image and each of the first image frames; during the playback of the target video, for each of the second image frames, segmenting the target subject and the first image frame, and displaying the target subject as the foreground and the first image frame as the background superimposed.
[0006] In a second aspect, the present application further provides a video processing device for use in an electronic device, the device comprising: a first acquisition unit, a second acquisition unit, and a display unit. The first acquisition unit is configured to acquire an original video, the original video comprising a plurality of first image frames; the second acquisition unit is configured to acquire a target video corresponding to the original video, wherein the target video comprises a plurality of second image frames spliced together based on a predetermined target subject image and each of the first image frames; and the display unit is configured to, during playback of the target video, segment each of the second image frames into a target subject and a first image frame, and display the target subject as a foreground and the first image frame as a background overlay.
[0007] In a third aspect, the present application also provides an electronic device comprising: one or more processors; a memory; and one or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more applications are configured to execute the above method.
[0008] In a fourth aspect, the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a program code executable by a processor, and when the program code is executed by the processor, the processor executes the above method.
[0009] The video processing method, device, electronic device and computer-readable medium provided by the present application are to obtain an original video, wherein the original video includes a plurality of first image frames; obtain a target video corresponding to the original video, wherein the target video includes a plurality of second image frames spliced based on a predetermined target subject image and each of the first image frames; and in the process of playing the target video, for each of the second image frames, separate the target subject and the first image frame, and display the target subject as the foreground and the first image frame as the background in a superimposed manner. Therefore, since the first image frame of the original video and the target subject image have been synthesized into a target video, when the target subject and the original video are superimposed and displayed by playing the target video, the first image frame and the target subject image are synchronized within the frame, that is, the synchronization of the first image frame and the target subject image is automatically achieved without the need to play the two videos synchronously.
[0010] Other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present application. The purposes and other advantages of the present application can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0012] Figure 1 A flow chart of a video processing method according to an embodiment of the present application is shown; Figure 2 A schematic diagram of a depth-of-field video operation interface provided by an embodiment of the present application is shown; Figure 3 A flow chart of a video processing method provided by another embodiment of the present application is shown; Figure 4 A schematic diagram showing the stacking relationship of various images provided in an embodiment of the present application; Figure 5 A schematic diagram of a depth-of-field video provided by an embodiment of the present application is shown; Figure 6 A schematic diagram of a second image frame provided by an embodiment of the present application is shown; Figure 7 A schematic diagram of the interaction between the lock screen and the colorful engine provided in one embodiment of the present application is shown; Figure 8 A flow chart of a video processing method according to another embodiment of the present application is shown; Figure 9 A schematic diagram of video playback for a wallpaper application and a lock screen application provided in an embodiment of the present application is shown; Figure 10 A flow chart of a video processing method provided by another embodiment of the present application is shown; Figure 11 A schematic diagram showing a video processing architecture provided by an embodiment of the present application is shown; Figure 12 An interaction diagram between an electronic device and the cloud provided by an embodiment of the present application is shown; Figure 13 A schematic diagram of a video depth of field algorithm model provided by an embodiment of the present application is shown; Figure 14 A schematic diagram of a depth of field cutout model provided by an embodiment of the present application is shown; Figure 15 A module block diagram of a video processing device provided by an embodiment of the present application is shown; Figure 16 A structural block diagram of an electronic device provided in an embodiment of the present application is shown; Figure 17 A storage unit for storing or carrying program codes for implementing the method according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0013] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application for which protection is claimed, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work fall within the scope of protection of the present application.
[0014] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.
[0015] Existing depth-of-field video solutions, which separate and process the subject and background in a video and combine them with advanced video synthesis and synchronization technologies, have gradually become an innovative way to enhance the user's visual experience. Specifically, this solution extracts the main subject from the video and stores it as a separate video stream, then uses synthesis technology to overlay it on top of the original video, achieving a dynamic depth-of-field effect.
[0016] However, in current solutions, the main video and the original video are two independent video streams. Therefore, precise synchronization of the two videos is required during playback. This places high demands on the player, especially ensuring that the timelines of the two videos are completely consistent. This not only increases the complexity of the player but also easily leads to synchronization errors, affecting the viewing experience.
[0017] In other words, the current depth of field video technology has the following shortcomings: 1. The main video and the original video are two videos, and it is difficult to compare the depth of field effect. The player needs to accurately control the progress of the two videos to be completely consistent; 2. Code is needed for synchronization during playback, and two players are used for playback, which requires higher decoding performance; 3. The file size of the two videos and the memory occupied during playback will be larger than that of a single video.
[0018] Therefore, in order to overcome the above-mentioned defects, the embodiment of the present application provides a video processing method, which is applied to electronic devices such as Figure 1 As shown, the method includes: S101 to S103.
[0019] S101: Acquire an original video, where the original video includes a plurality of first image frames.
[0020] It is understandable that the purpose of the embodiments of the present application is to obtain the visual effect of dynamic depth of field video. That is, for the user, the visual effect obtained is that when watching the video, they can see a depth of field effect, which can include background blur, gradual depth of field, and visual layering. Among them, background blur refers to highlighting the foreground object by blurring the background. The farther the background is from the focus, the stronger the blur effect. This effect is usually used to focus the viewer's attention on the characters, objects, or certain specific scene details in the video. Gradual depth of field means that the depth of field effect does not change suddenly, but gradually transitions from clear to blurred. This transition is usually used to express the change of scenery from close to far distance, or the shift from one focus to another. Visual layering means that the user can feel that the subject and background are on different levels, and other content can be added between the subject and background, such as adding a clock, date, or other text or pattern video.
[0021] Therefore, you first need to select an original video as the background video. Figure 2 As shown, it shows the depth of field video operation interface, through which the user selects the original video and opens the depth of field video effect for the original video. Figure 2 As shown, the operation interface includes a video display area 201, a first control 202 and a second control 203. Among them, the video display area 201 is used to display the original video selected by the user or the depth of field video processed by the video processing method of this application, the first control 202 is used to select the original video, that is, by clicking the control, the user can determine the original video of this time among the multiple videos displayed later, and the second control 203 is used to start the depth of field video switch, that is, after starting the switch, the subsequent implementation of the present application is executed to realize the depth of field effect of superimposing the main body and the original video. If the switch is not started, the original video is displayed. Taking the lock screen scene as an example, if the original video is configured for the lock screen scene and the depth of field video switch is started, then the depth of field effect of superimposing the main body and the original video is displayed in the lock screen scene. If the depth of field video switch is not started, then the original video is displayed in the lock screen scene.
[0022] It is understood that the original video includes multiple first image frames. It is understood that multiple image frames played at a certain speed (frame rate) form a dynamic image sequence, which is a video. In other words, the first image frame can be multiple consecutive video frames that constitute the original video.
[0023] S102: Acquire a target video corresponding to the original video, wherein the target video includes a plurality of second image frames spliced together based on a predetermined target subject image and each of the first image frames.
[0024] It can be understood that each first image frame is spliced with the target subject image to obtain a second image frame, and multiple second image frames constitute the target video.
[0025] As an embodiment, the stitching method may be to stitch each first image frame and the target subject image together to form a single image frame. For example, in the second image frame, the image content of the first image frame is located in the left half of the second image frame, and the image content of the target subject image is located in the right half of the second image frame. Of course, stitching may also be done top to bottom, and this is not limited to this.
[0026] It should be noted that the target subject image can be a predetermined image containing the target subject. The target subject in the target subject image can be the subject in the original video or not, and there is no limitation on this. If the target subject is not the subject in the original video, the method for obtaining the target subject image can be to determine a preset image selected by the user and extract the target subject image from the preset image. The preset image can be the currently selected image or an image determined based on user data, for example, an image of interest to the user determined based on the user's interests and hobbies. There is no limitation on this.
[0027] In addition, the target subject image can be a static image or can be from a preset video. If the target subject image is a static image, it is equivalent to splicing the static image with each first image frame to obtain a second image frame corresponding to each first image frame, thereby obtaining multiple second image frames.
[0028] If the target subject image is from a preset video, then the target subject image is at least one frame of the preset video. For example, the target subject image can be a single frame of the preset video, or all frames of the preset video, or a portion of all frames. The target subject image can be determined by the user clicking on the screen while the preset video is playing.
[0029] Therefore, when splicing the target subject image with the first image frame, the relationship between the number of the target subject image and the number of the first image frames needs to be considered.
[0030] If the number of target subject images is less than the number of first image frames, the target subject images can be repeated cyclically to fill the gap. In other words, by repeating the target subject images, the number of expanded target subject images is equal to the number of first image frames. For example, if the original video has 50 frames and there are only 10 target subject images, the target subject images can be repeated five times in sequence until the number matches the number of video frames. Alternatively, an image can be constructed using existing target subject images so that the number of constructed images and target subject images equals the number of first image frames. Specifically, the target subject images are interpolated or synthesized. If the number of target subject images is less than the number of frames in the original video, an interpolation algorithm (such as image interpolation, style transfer, etc.) can be used to generate additional target subject images to fill in the video. This method can achieve a visually smooth transition between target subject images and avoid simple repetition.
[0031] If the number of target subject images is equal to the number of first image frames, then the number of target subject images matches the number of first image frames, and each first image frame can be associated with a target subject image. Specifically, since the target subject images belong to a preset video, the target subject images can be sorted one by one based on the playback order of the preset video. The first image frames can be sorted one by one according to the playback order of the original video, and then the target subject images can be associated one by one according to the sorting of the first image frames. Then, the first image frame and the target subject images corresponding to the first image frame are spliced together to form a second image frame, thereby obtaining multiple second image frames.
[0032] If the number of target subject images is greater than the number of first image frames, consider randomly matching the video frames and target subject images or matching them sequentially. For example, for each video frame, randomly select a target subject image for splicing. This has the advantage of preventing a single image from appearing repeatedly in multiple frames.
[0033] As an embodiment, the target subject image can be from the original video, that is, the target subject image can be the subject image corresponding to each image frame of the original video. Exemplarily, each first image frame of the original video is obtained, and image processing operations are performed on each first image frame to identify the subject of the first image frame, thereby obtaining the target subject image corresponding to the first image frame. Then, each first image frame is spliced with the target subject image corresponding to the first image frame to obtain a second image frame. The plurality of second image frames constitute the target video.
[0034] It can be understood that, in the target subject image, the target subject serves as the foreground, and the background of the target subject image may not display any content.
[0035] In addition, the target video generation process can be performed by the electronic device or the cloud, that is, the electronic device sends the original video to the cloud, and the cloud combines the predetermined target subject image and the plurality of second image frames obtained by splicing each of the first image frames to obtain the target video, and then returns the target video to the electronic device. This is not limited to this.
[0036] S103: During the playback of the target video, for each second image frame, a target subject and a first image frame are segmented, and the target subject is used as a foreground and the first image frame is used as a background for superimposed display.
[0037] As an implementation, when playing the target video, every second frame of the decoded target video is temporarily stored in memory or video memory. This process is called "frame buffering." Each frame's image information is stored in memory, ready for rendering. During video playback, the player continuously retrieves frame data from the frame buffer and renders it using the graphics processing unit (GPU). The rendered image is then output to the screen.
[0038] In some embodiments, in the second image frame, the first image frame and the target subject image have been superimposed and synthesized. When playing the video, playing the second image frame in sequence can obtain a superimposed display effect with the target subject as the foreground and the first image frame as the background.
[0039] In other embodiments, in the second image frame, the first image frame and the target subject image are not superimposed, but rather spliced side-by-side or top-to-bottom. During playback of the target video, the target subject image and the first image frame can be superimposed into a single image for display, or the target subject image and the first image frame can be rendered separately and then sequentially superimposed for display. In other words, if the rendering method involves superimposing the two images into a single image for display, this method typically involves first synthesizing the target subject image (foreground) and the background image (e.g., the first image frame) to form a new single image (the composite image). During synthesis, the background image serves as the base layer, with the foreground image superimposed on top. Transparency or a mask is typically used to control the display of the foreground. If the two images are rendered separately and then sequentially superimposed for display, this method involves rendering the two images separately: first rendering the background image (the first image frame), then rendering the target subject image (the foreground). The order of rendering determines the superimposition of the target subject image on the background image.
[0040] Specifically, if the original video and the target main image are both static images, the two images can be superimposed into one image during rendering and then displayed; if at least one of the original video and the target main image is a dynamic video, the two images can be rendered separately and then superimposed and displayed in sequence.
[0041] It is understood that the superimposed display of the target subject as the foreground and the first image frame as the background can be achieved by using the target subject as the foreground and the first image frame as the background, and the first image frame can be subjected to illusory processing to achieve a blurred background effect. Of course, additional content can also be added between the target subject and the first image frame to provide certain occlusion or perspective effects to achieve a three-dimensional effect. This is not limited to this.
[0042] Therefore, compared with the method of saving the target subject as a video, namely the main video, and playing the main video and the original video synchronously to present the effect of the target subject as the foreground and the first image frame as the background superimposed, the embodiment of the present application uses the second image frame method, and the spliced image is already fixed, and all time synchronization problems have been solved in the video production stage. Therefore, during playback, the player only needs to display the second image frame in sequence according to the order of the video frames, without the need to additionally control the synchronization problem between the two videos. The use of the second image frame method can reduce the complexity of decoding and rendering. When the video is played, the target subject and the background image have been synthesized into one frame in advance, and the player only needs to process a single video stream. This saves more resources and reduces delays than decoding and rendering two independent video streams (the main video and the original video) at the same time.
[0043] See also Figure 3 , the embodiment of the present application provides a video processing method, which is applied to electronic devices, such as Figure 3 As shown, the method includes: S301 to S305.
[0044] S301: Acquire an original video, where the original video includes a plurality of first image frames.
[0045] S302: Obtain a target video corresponding to the original video, wherein the target video includes a plurality of second image frames spliced together based on a predetermined target subject image and each of the first image frames.
[0046] S303: Decode the target video to obtain multiple second image frames.
[0047] Specifically, during rendering, the target video is decoded using a decoder, and then the decoded video frame, i.e., the second image frame, is input as an external texture into a multi-color engine, which is a pre-defined module for performing the rendering operation of the present application.
[0048] S304: For each of the second image frames, obtain the first image frame and the designated image.
[0049] It should be noted that the designated image is an image obtained based on the target subject image and including the target subject, and the transparency parameter of the other areas of the designated image other than the target subject is set to a preset transparency. In other words, the designated image is an image obtained by the electronic device based on the target subject image, the designated image includes the target subject, and the transparency parameter of the other areas other than the target subject, i.e., the background area, is set to a preset transparency.
[0050] It is understood that the purpose of setting the transparency parameter of other areas outside the target subject to the preset transparency is to make the image content in other areas outside the target subject invisible or nearly transparent to the user. Therefore, an implementation method of setting the transparency parameter to the preset transparency can be to set the Alpha value to 0, that is, the preset transparency is 100% transparent.
[0051] S305: On the basis of displaying the first image frame, superimpose a designated image corresponding to the first image frame on the first image frame for display.
[0052] It is understandable that, when the superimposed display is performed, the first image frame is displayed frame by frame, and since the designated image is determined based on the target main image, and the target main image is spliced with the first image frame to form the second image frame, during the playback of the second image frame of the target video, the designated image is also displayed frame by frame, and is displayed synchronously with the first image frame. Therefore, the display of multiple designated images is equivalent to the playback of the designated video for the user. In other words, the visual effect of S305 is to superimpose the playback of the designated video on the basis of the playback of the original video. However, the operation method is not that two independent video playback channels are playing two different videos synchronously, but that the target video, i.e., a single video, is played. This does not reduce the user's visual experience, and at the same time avoids the defects caused by the synchronous playback of two videos.
[0053] like Figure 4 As shown, the positional relationship of each image in the stacked display is shown in the figure. It can be seen that the content displayed by the electronic device includes the original video, content 2, transparent video and content 1, wherein the transparent video refers to a video that only retains the foreground and removes the background, and the foreground is clear and retains the original image while the background is a 100% transparent image, that is, the transparent video refers to multiple specified images. It can be understood that this application is not intended to combine multiple specified images into one video. The transparent video in the figure is for the user's visual effect. In the embodiment of the present application, content 1 and content 2 can be set based on demand, wherein content 1 and content 2 can be images, specifically static images or dynamic images.
[0054] See also Figure 5, which shows the effect of superimposing the target subject on the original video (hereinafter referred to as the depth of field video effect). It can be seen that the original video display content 501 is at the bottom layer, followed by the second content 502, the designated image 503, and the first content 504. The second content is the aforementioned content 2, and the first content is the aforementioned content 1.
[0055] That is to say, on the basis of displaying the first image frame, the implementation method of superimposing the designated image corresponding to the first image frame on the first image frame may be, on the basis of displaying the first image frame, superimposing and displaying a preset image; on top of the preset image, superimposing and displaying the designated image corresponding to the first image frame. That is to say, when the stacked display is performed, a preset image can be displayed between the original video and the target subject. The preset image can be a clock, date, notification or other content, and is not specifically limited. The relationship between the first image frame, the preset image and the designated image can refer to the aforementioned Figure 4 and Figure 5 .
[0056] As an embodiment, for each second image frame, obtaining the first image frame and the designated image may be performed by extracting the first texture data of the first image frame within each second image frame, determining the target subject image for each second image frame, and setting the transparency parameter of the target subject image's other regions outside the target subject to a preset transparency to obtain the second texture data of the designated image. Then, upon displaying the first image frame, the designated image corresponding to the first image frame is superimposed on the first image frame by sequentially associating the first texture data and the second texture data to a rendering surface (i.e., a Surface); and displaying the first image frame and the designated image sequentially via the rendering surface in the order in which they were associated.
[0057] As previously described, during rendering, the target video is decoded using a decoder. The decoded video frame, i.e., the second image frame, is then input into the multi-pass engine as an external texture. When processing the current second image frame, the multi-pass engine uses a multi-rendering output method. Specifically, the first pass outputs the first texture data of the first image frame within the second image frame to the second pass. The texture of the target main image is then alpha-filtered to output the extracted main texture to the third pass. Alpha filtering sets the alpha value of the target main area to 1 (opaque) and the alpha value of the background area to 0 (fully transparent). Finally, the second and third passes are each associated with an external surface (i.e., a rendering surface). Thus, during display, the first image frame and the designated image are displayed sequentially via the rendering surface in the order in which they were associated.
[0058] In one embodiment, the target subject image is a binary mask image containing the target subject, i.e., a mask image, and the first image frame is a color image, i.e., an RGB image. Therefore, for each of the second image frames, the target subject image is determined, and the transparency parameter of the target subject image's other regions outside the target subject is set to a preset transparency to obtain the second texture data of the specified image. An embodiment may be to determine the target subject image for each of the second image frames, restore the target subject image to a color image, and set the transparency parameter of the target subject image's other regions outside the target subject to a preset transparency to obtain the second texture data of the specified image.
[0059] That is, for each second image frame, the target subject image of the second image frame is extracted. If the target subject image is currently a mask image, it is restored to an RGB image to obtain the target subject image in RGB format, i.e., the intermediate image. Specifically, the target subject image is extracted from a certain image or video, and the target subject area in the target subject image can be mapped to a region of a certain image, thereby determining the color mapping relationship of the target subject area. For example, the target subject image is extracted from the original video, i.e., each first image frame of the original video corresponds to a target subject image.
[0060] like Figure 6 As shown, a second image frame is shown, which includes a first image frame 610 and a target subject image 620. The first image frame 610 is an RGB image, and the target subject image 620 is a mask image. It can be seen that the target subject 621 (i.e., the white area) of the target subject image 620 corresponds to the foreground object 611 of the first image frame 610, and the background 612 of the first image frame 610 is a black area in the target subject image 620.
[0061] Therefore, in the process of restoring the target subject image to a color image, the target subject area can correspond to the display area of the first image frame, so that the RGB parameters of the target subject in the first image frame can be determined, and thus it can be restored to an image in RGB format. At the same time, by setting the transparency parameter, the transparency of the background area is fully transparent and the subject area is opaque, thereby obtaining the specified image, and then obtaining the second texture data of the specified image.
[0062] like Figure 7 As shown, for example, assuming that the user sets a video depth of field wallpaper through the wallpaper application, the wallpaper APK calls the colorful engine across processes through the binder based on the data protocol to pass the URI of the video resource to be loaded. In Android, the Binder mechanism is used to implement cross-process communication. Through Binder, the application can interact with other processes (such as services) to pass data or call methods. The wallpaper APK interacts with the colorful engine (as a service) by calling the interface and passes the video resource URI to be loaded, that is, the identifier of the video resource. In the embodiment of the present application, the target main image is the main body of the original video, so the video resource URI is the URI of the original video. It can be understood that the URI can be a local path or a remote resource URL.
[0063] When the user enters the system lock screen interface, the lock screen service creates two TextureViews, one of which is used to render the transparent video layer, that is, the specified image, and the other is used to render the background video layer, that is, the original video. The former is named AlphaSurface and the latter is named NormalSurface.
[0064] The lock screen service uses bindService to bind the Colorful Engine's video wallpaper service to the native Android system's WallpaperService. It also binds the foreground and background Surfaces (AlphaSurface and NormalSurface) to Colorful Engine, allowing the video wallpaper service to render the foreground and background. Note that WallpaperService is a native Android service that manages wallpaper display. The lock screen service uses bindService to bind the Colorful Engine's video wallpaper service to the system's WallpaperService. This allows the lock screen service to access the video wallpaper service's functionality.
[0065] In an embodiment of the present application, each first image frame of the original video and its corresponding target subject image are spliced into a second image frame. In the second image frame, the left part is the first image frame and the right part is the second image frame.
[0066] The Colorful Engine uses a video resource composed of a left-right spliced original video and a depth-of-field mask (i.e., the target subject image) to split the video into two parts and then blend them frame by frame, using the second and third passes described above. For details, refer to the previous section and will not be repeated here. The main foreground layer with transparent pixels is then rendered to the foreground AlphaSurface for playback, while the original video is rendered to the background NormalSurface for playback. Simultaneously, the background TextureView's onFrameAvailable callback controls the foreground and background frames to ensure synchronization.
[0067] Therefore, in the practical example of this application, the video created by splicing the binarized black-and-white image of the main subject to the right of the original video is half the size of the original video and the depth-of-field video stored separately. The spliced video is a single video, requiring no additional logic to synchronize the original and depth-of-field videos. During decoding, every frame of the original and depth-of-field videos is the same frame. This eliminates the need for two decoders for decoding, saving CPU resources and memory consumed by the phone's decoding process and improving performance.
[0068] See also Figure 8 , the embodiment of the present application provides a video processing method, which is applied to electronic devices, such as Figure 8 As shown, the method includes: S801 to S804.
[0069] S801: Acquire an original video, where the original video includes a plurality of first image frames.
[0070] S802: Obtain a target video corresponding to the original video, wherein the target video includes a plurality of second image frames spliced together based on a predetermined target subject image and each of the first image frames.
[0071] The implementation of S801 and S802 can refer to the above embodiments and will not be described in detail here.
[0072] S803: When a lock screen scene is detected, the target video is played.
[0073] In the embodiment of the present application, the purpose of the target video is to display the target video in the lock screen interface when the user enters the lock screen interface, so that the user can obtain the visual effect of the depth of field video, for example, Figure 5 Therefore, when it is detected that the electronic device enters the lock screen scene, the target video is played.
[0074] S804: During the playback of the target video, for each of the second image frames, a target subject and a first image frame are segmented, and the target subject is used as a foreground and the first image frame is used as a background for superimposed display.
[0075] The implementation of S804 may refer to the aforementioned embodiment and will not be described in detail here.
[0076] As an embodiment, the target subject image is extracted from the original video, that is, each first image frame of the original video corresponds to a target subject image, and each second image frame of the target video is an image spliced together with the first image frame and the target subject image corresponding to the first image frame, wherein the target subject image corresponding to the first image frame refers to the target subject image extracted from the first image frame.
[0077] Furthermore, considering that users may unlock the screen when the screen is locked, if the target subject image is extracted from the original video, the difference between playing the target video and the original video is that a preset image can be added between the target subject and the original video to create a depth of field effect. If the preset image is removed, the content displayed on the screen is almost the same as when playing the original video.
[0078] So, if Figure 9 As shown in the figure, the content played by the wallpaper app and the lock screen app are respectively shown. Among them, the original video and the transparent video are the target video as a whole, and the clock is the preset image. In other words, after the lock screen interface is unlocked, the electronic device displays the desktop. At this time, the lock screen app is invisible, so the wallpaper app also displays the original video, so that the progress of the video played on the lock screen interface can be seamlessly connected after unlocking.
[0079] Therefore, the above-mentioned method embodiment also includes, in the lock screen scenario, running a desktop wallpaper application in the background, and in the desktop wallpaper application, playing the original video synchronously with the target video; when it is detected that the lock screen scenario is switched to the desktop wallpaper scenario, based on the current playback progress of the original video, continuing to play the original video in the desktop wallpaper scenario.
[0080] Specifically, when a lock screen scene is detected, the target video is played, and at the same time, a desktop wallpaper application is run in the background, and in the desktop wallpaper application, the original video is played synchronously with the target video. That is, the original video is played in the background, and when the electronic device is currently in a lock screen scene, the playback process of the original video is imperceptible to the user. That is, if the original video has audio data, the audio data is played silently, and the display content of the original video is not displayed on the screen of the electronic device. For example, the original video is rendered to an invisible Surface. In other words, the original video is in a state of being played, but its video content is not displayed on the screen of the electronic device.
[0081] Therefore, when it is detected that the lock screen scene switches to the desktop wallpaper scene, the current playback progress of the original video is obtained, and the original video continues to be played in the desktop wallpaper scene. It is understandable that when it is detected that the lock screen scene switches to the desktop wallpaper scene, this means that the lock screen scene has been exited, that is, the content of the target video is no longer visible, then the playback of the target video can be paused at this time, and then the current playback progress of the original video is obtained. Since the original video and the target video are played synchronously, the current playback progress of the original video can be regarded as the playback progress when the target video ends. Therefore, continuing to play the original video in the desktop wallpaper scene is equivalent to continuing to play the original video at the moment when the target video ends. And since the target subject in the target video is extracted from each image frame of the original video, the original video is continued to be played based on the current playback progress of the original video. The content seen by the user in the desktop scene can be better connected with the video content last seen when the lock screen switches to the desktop scene, avoiding visual tearing.
[0082] In addition, in addition to the above-mentioned method, it is also possible that, when it is detected that the lock screen scene switches to the desktop wallpaper scene, the current playback progress of the target video is determined; based on the current playback progress of the target video, the original video is played in the desktop wallpaper scene. Different from the above-mentioned method, in the lock screen scene, the original video does not need to be played. Due to the relationship between the original video and the target video introduced above, when the lock screen scene switches to the desktop wallpaper scene, the original video is started to play based on the current playback progress of the target video, which can also avoid visual tearing. In addition, when it is detected that the lock screen scene switches to the desktop wallpaper scene, the target video stops playing.
[0083] See also Figure 10 , the embodiment of the present application provides a video processing method, which is applied to electronic devices, such as Figure 8 As shown, the method includes: S1001 to S1003.
[0084] S1001: Acquire an original video, where the original video includes a plurality of first image frames.
[0085] S1002: Send the original video to the cloud, so that the cloud determines the target main image, splices the target main image and each of the first image frames into a second image frame, to obtain the target video and return it.
[0086] S1003: During the playback of the target video, for each of the second image frames, a target subject and a first image frame are segmented, and the target subject is used as a foreground and the first image frame is used as a background for superimposed display.
[0087] In the embodiment of the present application, the operation of generating the target video can be handed over to the cloud. That is, the electronic device can send the original video to the cloud, then obtain the target video sent by the cloud, save it, and play the target video when it is needed later.
[0088] As an implementation method, Figure 11 As shown, using the lock screen as an example, users can select the source video in the wallpaper app. The cloud then generates the target video based on the source video. The wallpaper APK uses the cloud-side video depth algorithm model through the end-cloud communication process to obtain the resources for the frame-by-frame depth mask of the current video wallpaper.
[0089] It can be understood that the target subject image is the image corresponding to the target subject extracted from the first image frame. The electronic device then sends the original video to the cloud, which obtains each first image frame of the original video, extracts the subject from each first image frame, and obtains the target subject image corresponding to the first image frame. Each first image frame and the target subject image corresponding to the first image frame are then spliced together into a second image frame, thereby obtaining multiple second image frames and ultimately obtaining the target video.
[0090] In other words, the embodiments of the present application take into account the high computing power required to generate the target video, while the computing power of electronic devices (e.g., mobile phones) is limited. Therefore, the processing of the target video is handled by the cloud. That is, due to the high computational complexity of the video depth of field generation algorithm model, it is deployed on the cloud side, and the data to be processed and the generated results are obtained through HTTPS requests. The video depth of field generation algorithm model refers to the model used to obtain the target video based on the original video, the data to be processed refers to the original video, and the generated result is the target video.
[0091] like Figure 12 As shown in the figure, the service module (wallpaper APK) initiates a video data processing request (cloudVideoMatching). This request includes parameters such as the video URI to be processed, for example, the URI of the original video selected by the user in the wallpaper app. The algorithm SDK, as the core processing unit on the device, receives the service request and performs preliminary parameter verification and data encapsulation.
[0092] Specifically, the user selects a video as wallpaper in the wallpaper app, which determines the original video. The video's Uniform Resource Identifier (URI) is extracted as part of the request. Other parameters may also be included, such as the user-selected wallpaper effect (e.g., whether specific special effects are required, resolution requirements), device information (screen resolution, device model), and other video processing-related parameters. This information is sent to the on-device algorithm SDK through SDK interface calls for further processing. The algorithm SDK performs preliminary parameter validation on the video data processing request, such as checking the video URI to see if it is valid and correctly points to an existing video file. It also checks other relevant parameters for validity. For example, it checks the validity of additional parameters in the request, particularly whether processing-related requirements such as resolution and effect options are within a reasonable range. This may also include device information validation. If the request includes device information, the SDK ensures that the device's relevant information (such as screen resolution and hardware capabilities) is compatible with the requested processing capabilities.
[0093] Afterwards, the algorithm SDK forwards the video data processing request to the AIUnit plug-in collection, and AIUnit calls the Plugin module to load the video depth of field plug-in, and flexible functional expansion can be achieved later. The video depth of field plug-in uploads the video data to the cloud storage offload communication service (OCS) through the inter-process communication (IPC) mechanism. The OCS compresses, encrypts, and blocks the data, and returns the resource URL. The video depth of field plug-in then requests the cloud-side interface through the HTTP / HTTPS protocol, and then schedules the GPU cluster computing resources to execute the video depth of field model algorithm. The algorithm model returns the execution results to the video depth of field plug-in through the cloud-side interface. The plug-in downloads the result video resource through the preset URL, generates a JSON structured result, and passes it to the business module through multi-layer returns, completing the end-cloud collaborative computing closed loop.
[0094] See also Figure 13 , which shows the architecture of the video depth generation algorithm model. Specifically, the video decoding uses the nvidia hardware acceleration library and performs video decoding through FFMPEG. The model includes a depth of field cutout model. Figure 14 As shown in Figure 2, the model's network architecture utilizes an encoder-decoder architecture and multi-scale feature fusion. The encoder extracts multi-level features based on a lightweight backbone network, reducing the number of model parameters while preserving spatial semantic information. The decoder upsamples the feature maps through operations such as deconvolution and bilinear interpolation, and integrates skip connections to fuse shallow, detailed encoder features with deep semantic features.
[0095] Figure 13 The model inference structure in
[15] includes a scale-invariant contextual attention network (SICA) and a Laplacian multi-scale pyramid. SICA uses non-local attention to model long-range pixel dependencies, addressing the lack of global context awareness in traditional convolutional neural networks. The Laplacian multi-scale pyramid decomposes input features into different frequency components (low-frequency contours and high-frequency details), extracting and fusing these features separately to improve the model's adaptability to multi-scale targets. The model's training data includes both synthetic and real data. Synthetic data refers to training samples synthesized from high-quality foreground footage (such as human and animal cutouts) and diverse backgrounds (natural scenery, urban buildings, etc.), annotated with precise masks and transparency labels. Real data refers to real video data collected from business scenarios using public datasets and integrated with real-world scenarios, covering dynamic backgrounds and complex lighting.
[0096] The model's post-processing capabilities utilize the OpenCV library to achieve pixel-level stitching of the original video frame and the mask. Utilizing FFMPEG's nvenc hardware encoding module, the processed video stream is efficiently encoded into H.264 / H.265 format, ensuring output video clarity and compression efficiency. The model's special scene processing capabilities include dynamic background interference and translucent object processing. Dynamic background interference refers to the addition of motion to the training data; translucent objects refer to matting loss training, specifically for human hair and animal fur. The model's rejection capabilities include determining whether depth of field is supported. This means that for target scenes (natural scenery, portraits, cityscapes, cartoons, and pets), material access is performed based on cover frames or keyframes, while non-target scenes are notified to the user.
[0097] See also Figure 15 , which shows a structural block diagram of a video processing device 1500 provided in an embodiment of the present application. The device is applied to an electronic device and may include: a first acquisition unit 1501, a second acquisition unit 1502 and a display unit 1503.
[0098] The first acquisition unit 1501 is configured to acquire an original video, where the original video includes a plurality of first image frames.
[0099] The second acquisition unit 1502 is configured to acquire a target video corresponding to the original video, wherein the target video includes a plurality of second image frames spliced together based on a predetermined target subject image and each of the first image frames.
[0100] Furthermore, the second acquisition unit 1502 is also used to send the original video to the cloud, so that the cloud can determine the target main image, splice the target main image and each of the first image frames into a second image frame to obtain the target video and return it.
[0101] The display unit 1503 is used to segment the target subject and the first image frame for each second image frame during the playback of the target video, and to display the target subject as the foreground and the first image frame as the background. Furthermore, the display unit 1503 is also used to decode the target video to obtain multiple second image frames; for each of the second image frames, a first image frame and a designated image are obtained, wherein the designated image is an image including a target subject obtained based on the target subject image, and the transparency parameters of other areas outside the target subject in the designated image are set to a preset transparency; on the basis of displaying the first image frame, the designated image corresponding to the first image frame is superimposed on the first image frame for display.
[0102] Furthermore, the display unit 1503 is also used to extract the first texture data of the first image frame in each of the second image frames; for each of the second image frames, determine the target subject image, and set the transparency parameters of other areas outside the target subject of the target subject image to a preset transparency to obtain the second texture data of the specified image.
[0103] Furthermore, the display unit 1503 is further configured to sequentially associate the first texture data and the second texture data with a rendering surface; and sequentially display the first image frame and the designated image through the rendering surface in the order of association.
[0104] Furthermore, the target subject image is a binary mask image containing the target subject, the first image frame is a color image, and the display unit 1503 is also used to determine the target subject image for each of the second image frames, restore the target subject image to a color image, and set the transparency parameters of other areas outside the target subject to a preset transparency to obtain the second texture data of the specified image.
[0105] Furthermore, the display unit 1503 is further configured to display a preset image in a superimposed manner on the basis of displaying the first image frame; and to display a designated image corresponding to the first image frame in a superimposed manner on the preset image.
[0106] Furthermore, the display unit 1503 is also used to detect the lock screen scene and play the target video; in the process of playing the target video, for each second image frame, the target subject and the first image frame are separated, and the target subject is displayed as the foreground and the first image frame is displayed as the background.
[0107] Furthermore, the display unit 1503 is also used to run a desktop wallpaper application in the background in the lock screen scenario, and in the desktop wallpaper application, play the original video synchronously with the target video; when it is detected that the lock screen scenario is switched to the desktop wallpaper scenario, the original video continues to be played in the desktop wallpaper scenario based on the current playback progress of the original video.
[0108] Furthermore, the display unit 1503 is also used to determine the current playback progress of the target video when detecting a switch from the lock screen scene to the desktop wallpaper scene; based on the current playback progress of the target video, play the original video in the desktop wallpaper scene.
[0109] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0110] In several embodiments provided in this application, the coupling between modules may be electrical, mechanical or other forms of coupling.
[0111] In addition, the functional modules in the various embodiments of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module. The above-mentioned integrated modules may be implemented in the form of hardware or software functional modules.
[0112] Please refer to Figure 16 , which shows a structural block diagram of an electronic device provided in an embodiment of the present application. The electronic device 100 can be an electronic device capable of running applications, such as a smartphone, a tablet computer, an e-book, etc. The electronic device 100 in the present application may include one or more of the following components: a processor 110, a memory 120, and one or more applications, wherein the one or more applications may be stored in the memory 120 and configured to be executed by one or more processors 110, and one or more programs are configured to execute the method described in the aforementioned method embodiment.
[0113] The processor 110 may include one or more processing cores. The processor 110 utilizes various interfaces and circuits to connect various components within the electronic device 100. It executes instructions, programs, code sets, or instruction sets stored in the memory 120, and accesses data stored in the memory 120 to perform various functions and process data within the electronic device 100. Optionally, the processor 110 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 110 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem handles wireless communications. It is understood that the modem may also be implemented independently of the processor 110 via a separate communications chip.
[0114] The memory 120 may include random access memory (RAM) or read-only memory (ROM). The memory 120 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the various method embodiments described below, and the like. The data storage area may also store data created by the electronic device 100 during use (such as a phone book, audio and video data, and chat history data).
[0115] Please refer to Figure 17 , which shows a block diagram of a computer-readable medium provided in an embodiment of the present application. The computer-readable medium 1700 stores program code, which can be called by a processor to execute the method described in the above method embodiment.
[0116] Computer-readable medium 1700 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, a hard disk, or ROM. Alternatively, computer-readable medium 1700 may include non-transitory computer-readable storage media. Computer-readable medium 1700 has storage space for program code 1710 for executing any of the method steps described above. This program code can be read from or written to one or more computer program products. Program code 1710 may be compressed, for example, in a suitable format.
[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A video processing method, characterized in that: Applied to electronic equipment, the method includes: Acquire an original video, where the original video includes a plurality of first image frames; Acquire a target video corresponding to the original video, wherein the target video includes a plurality of second image frames spliced together based on a predetermined target subject image and each of the first image frames; During the process of playing the target video, for each second image frame, the target subject and the first image frame are segmented, and the target subject is used as the foreground and the first image frame is used as the background for superimposition display.
2. The method according to claim 1, characterized in that In the process of playing the target video, for each second image frame, segmenting the target subject and the first image frame, and displaying the target subject as the foreground and the first image frame as the background in an overlaid manner, including: Decoding the target video to obtain a plurality of second image frames; For each of the second image frames, acquiring the first image frame and a designated image, wherein the designated image is an image including a target subject obtained based on the target subject image, and the transparency parameter of other areas outside the target subject in the designated image is set to a preset transparency; On the basis of displaying the first image frame, the designated image corresponding to the first image frame is superimposed on the first image frame and displayed.
3. The method according to claim 2, characterized in that The acquiring of the first image frame and the designated image for each of the second image frames includes: For each of the second image frames, extracting first texture data of the first image frame in the second image frame; For each of the second image frames, a target subject image is determined, and transparency parameters of other areas of the target subject image other than the target subject are set to a preset transparency, so as to obtain second texture data of a designated image.
4. The method according to claim 3, characterized in that The step of displaying the first image frame by superimposing a designated image corresponding to the first image frame on the first image frame includes: Associating the first texture data and the second texture data to a rendering surface in sequence; The first image frame and the designated image are displayed in sequence through the rendering surface according to the order in which they are associated.
5. The method according to claim 3, characterized in that The target subject image is a binary mask image containing a target subject, the first image frame is a color image, and for each of the second image frames, determining the target subject image, and setting the transparency parameter of other areas of the target subject image outside the target subject to a preset transparency to obtain second texture data of the specified image, including: For each of the second image frames, a target subject image is determined, the target subject image is restored to a color image, and transparency parameters of other areas outside the target subject are set to a preset transparency to obtain second texture data of the specified image.
6. The method according to claim 2, characterized in that The step of displaying the first image frame by superimposing a designated image corresponding to the first image frame on the first image frame includes: On the basis of displaying the first image frame, superimposing and displaying a preset image; The designated image corresponding to the first image frame is superimposed and displayed on the preset image.
7. The method according to any one of claims 1 to 6, characterized in that The obtaining of the target video corresponding to the original video includes: The original video is sent to the cloud, so that the cloud determines the target main image, splices the target main image and each of the first image frames into a second image frame, to obtain the target video and return it.
8. The method according to any one of claims 1 to 6, characterized in that The target subject image is an image corresponding to the target subject extracted from the first image frame.
9. The method according to any one of claims 1 to 6, characterized in that In the process of playing the target video, for each second image frame, segmenting the target subject and the first image frame, and displaying the target subject as the foreground and the first image frame as the background in an overlaid manner, including: When a lock screen scene is detected, the target video is played; During the process of playing the target video, for each second image frame, the target subject and the first image frame are segmented, and the target subject is used as the foreground and the first image frame is used as the background for superimposition display.
10. The method according to claim 9, characterized in that The target subject image is an image corresponding to the target subject extracted from the first image frame, and the method further includes: In a lock screen scenario, running a desktop wallpaper application in the background, and playing the original video synchronously with the target video in the desktop wallpaper application; When it is detected that the lock screen scene is switched to the desktop wallpaper scene, the original video continues to be played in the desktop wallpaper scene based on the current playback progress of the original video.
11. The method according to claim 9, characterized in that The target subject image is an image corresponding to the target subject extracted from the first image frame, and the method further includes: When detecting that the lock screen scene switches to the desktop wallpaper scene, determining the current playback progress of the target video; Based on the current playback progress of the target video, the original video is played in the desktop wallpaper scene.
12. A video processing device, characterized in that: Applied to electronic equipment, the device comprises: A first acquisition unit is used to acquire an original video, where the original video includes a plurality of first image frames; A second acquisition unit is configured to acquire a target video corresponding to the original video, wherein the target video includes a plurality of second image frames spliced together based on a predetermined target subject image and each of the first image frames; The display unit is used to segment the target subject and the first image frame for each of the second image frames during the playback of the target video, and to superimpose and display the target subject as the foreground and the first image frame as the background.
13. An electronic device, characterized in that: include: one or more processors; Memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more applications are configured to execute the method according to any one of claims 1 to 11.
14. A computer-readable medium, characterized in that The computer-readable medium stores a program code executable by a processor, and when the program code is executed by the processor, the processor executes the method according to any one of claims 1 to 11.