Wallpaper display method and electronic equipment
By detecting user operations and displaying coherent wallpapers from AOD to lock screen to desktop, the problem of incoherent wallpaper display in the existing technology is solved, dynamic effects and user custom functions are achieved, and user experience is improved.
Patent Information
- Application Number
- CN202311458292.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-02
- Publication Date
- 2025-05-13
AI Technical Summary
The existing technology is difficult to achieve coherent and dynamic wallpaper displays from AOD to lock screen to desktop, etc., and the wallpaper transition effect is stiff and the user experience is poor.
By detecting user operations, the electronic device displays coherent wallpapers from AOD to lock screen and then to desktop, using video frames to arrange them in sequence to achieve dynamic effects, and supports users to customize wallpapers, including selecting video clips and editing wallpaper subjects and backgrounds.
It realizes coherent dynamic wallpaper displays from AOD to lock screen to desktop and other interfaces, which improves the user experience, provides rich wallpaper creation space, and simplifies user operations.
Smart Images

Figure CN119996558A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of electronic technology, and in particular to a wallpaper display method and an electronic device. Background Art
[0002] When using mobile phones and other electronic devices, wallpaper is the interface that users see most often and can also show the user's personality. At present, wallpaper is mainly used in user interfaces such as always on display (AOD), lock screen and desktop. Users often like to set wallpapers in AOD, lock screen, desktop and other places to pursue beauty and personality. Summary of the invention
[0003] The embodiments of the present application provide a wallpaper display method and an electronic device, which can generate a coherent and dynamic wallpaper from AOD to lock screen to desktop based on user customized operations, meet the user's demand for customized wallpaper, and provide users with a rich and colorful wallpaper creation space.
[0004] In a first aspect, an embodiment of the present application provides a wallpaper display method, which may include: an electronic device in a locked screen state or an unlocked state detects a first operation, and in response to the first operation, the electronic device displays a first interface, and the first interface includes an AOD wallpaper. When the electronic device displays the first interface, in response to a second operation, the electronic device displays a second interface, and the second interface includes a lock screen wallpaper. In response to a third operation for unlocking, the electronic device displays a third interface, and the third interface includes a desktop wallpaper.
[0005] Among them, the AOD wallpaper, lock screen wallpaper and desktop wallpaper can correspond to the first video frame, the second video frame and the third video frame in the first video respectively, the second video frame is after the first video frame, and the third video frame is after the second video frame.
[0006] In the first aspect, the first interface, the second interface, and the third interface may be an AOD interface, a lock screen interface, and a desktop interface, respectively.
[0007] In the first aspect, the first operation may be a screen-off operation, such as pressing the power button to turn off the screen of an electronic device in a locked or unlocked state. In the locked or unlocked state, the screen of the electronic device is in a lit state, wherein the lock screen interface is displayed on the screen in the locked state, and the desktop interface or the user interface of the application opened by the user is displayed on the screen in the unlocked state. Not limited to being triggered by the screen-off operation, if the user does not operate the mobile phone for a long time in the locked or unlocked state, the electronic device may also display the AOD interface.
[0008] In the first aspect, the second operation may be a user operation of lighting up the screen, such as an operation of lightly touching the screen, an operation of pressing the power button, etc., but is not limited thereto. Lighting up the screen in the AOD interface may also be a scheduled arrival, and the user may set the time for lighting up the screen, and when the time arrives, the electronic device lights up the screen. The embodiment of the present application does not limit the operation of lighting up the screen in the AOD interface.
[0009] By implementing the method provided in the first aspect, from AOD to lock screen and then to desktop, the AOD wallpaper, lock screen wallpaper and desktop wallpaper are linked together and are no longer separated from each other.
[0010] In combination with the first aspect, in some embodiments, the first video frame, the second video frame, and the third video frame are all one frame, the frame interval between the first video frame and the second video frame is greater than the first frame interval, and the frame interval between the second video frame and the third video frame is greater than the second frame interval. For example, the first video frame, the second video frame, and the third video frame can be extracted from the first video at a certain frame interval (such as more than 1 second). In this way, the images of these multiple video frames can be different, and a dynamic change effect can be presented when they are played continuously.
[0011] In combination with the first aspect, in some embodiments, the first video frame, the second video frame, and the third video frame are all multiple frames, the start frame in the second video frame is the end frame in the first video frame, and the start frame in the third video frame is the end frame in the second video frame. In other words, the video frames used by the AOD wallpaper, the lock screen wallpaper, and the desktop wallpaper need to be connected end to end according to the interface switching timing, so as to achieve an overall coherent dynamic effect from AOD to the lock screen and then to the desktop.
[0012] Combined with the first aspect, in some embodiments, the wallpaper can be well integrated with interfaces such as AOD, lock screen, and desktop, and the main object (such as a person) in the wallpaper will not be blocked by non-interactive elements such as time in interfaces such as AOD, lock screen, and desktop. Interactive elements in interfaces such as AOD, lock screen, or desktop, such as the fingerprint unlocking key in the AOD interface, the password keyboard in the lock screen interface, and the application icons in the desktop, can be set on the main object of the wallpaper to facilitate users to operate these interactive elements and ensure the realization of the interface interaction function.
[0013] In combination with the first aspect, in some embodiments, the AOD wallpaper, the lock screen wallpaper and the desktop wallpaper may directly adopt the original video frames in the first video, or may use image frames after image processing such as cropping and filtering the original video frames.
[0014] In conjunction with the first aspect, in some embodiments, the first video may be selected from a gallery video. Figure 3BAs shown, the user can select the video clip 14 as the first video for generating a dynamic wallpaper. The first video can be selected from a video in the gallery, or a live photo in the gallery. A live photo is also a video. In this way, the user's need for customizing wallpapers can be met.
[0015] In conjunction with the first aspect, in some embodiments, the first video is selected from videos provided by a setting application.
[0016] In some embodiments, not limited to from AOD to lock screen and then to desktop, the method provided in the first aspect can also be applied to other adjacent interface jump scenarios, such as AOD to desktop and then to certain application interfaces, which will especially occur when the user does not set the lock screen, and for example, from desktop to application main interface and then to application secondary interface. Of course, these application interfaces can support wallpaper display, rather than application interfaces such as game display interface and video playback interface that are not suitable for wallpaper display. For example, the chat interface of "WeChat" can set chat wallpaper. That is: the electronic device can display multiple user interfaces of different levels in turn, and the electronic device also displays wallpapers at multiple user interfaces. The wallpaper is generated based on the first video, and the video frames used by the wallpaper at multiple user interfaces at different levels are arranged in sequence; wherein, the main object in the wallpaper is displayed above the non-interactive elements in the user interface and below the interactive elements in the user interface. In this way, when the user performs multi-level interface jumps, a coherent and dynamic wallpaper display effect is formed as a whole.
[0017] In combination with the first aspect, in some embodiments, the wallpaper display method may further include: the electronic device displays a fourth interface, the fourth interface including a video frame sequence of the second video. A first video is selected from the video frame sequence, wherein the first video is a partial video clip in the second video or the entire second video. The fourth interface may be, for example, Fig.10E In the illustrated “video wallpaper preview”, the second video may be video 16 , and the video frame sequence of the second video may be video frame sequence 18 .
[0018] In some embodiments, there may be a selection box on the video frame sequence, such as Fig.10E The selection box 19 in the selection box is a first video. The user operation of selecting a video segment is specifically an operation of sliding the selection box along the video frame sequence.
[0019] In some embodiments, the fourth interface also includes a first preview, which is used to display the display status of the AOD wallpaper, the lock screen wallpaper, and the desktop wallpaper in the first interface, the second interface, and the third interface respectively. The first preview can be Fig.10E The first preview may be a video, and the first preview presents a dynamic display process of switching from the AOD wallpaper to the lock screen wallpaper and then to the desktop wallpaper when playing.
[0020] In combination with the first aspect, in some embodiments, the wallpaper display method may further include: the electronic device displays a fifth interface, in which the main object and the background of the wallpaper are allowed to be selected separately by the user. The fifth interface may be, for example, Fig.11A The "Wallpaper Editing" interface shown. In some embodiments, the wallpaper display method may also include: in response to a user operation of selecting a subject to edit the subject, the electronic device updates the image of the subject, and refreshes the AOD wallpaper, lock screen wallpaper, and desktop wallpaper based on the updated subject image. In some embodiments, the wallpaper display method may also include: in response to a user operation of selecting a background of a wallpaper to edit the background, the electronic device updates the image of the background, and refreshes the AOD wallpaper, lock screen wallpaper, and desktop wallpaper based on the updated background image.
[0021] In some embodiments, before the electronic device displays the fifth interface, the wallpaper display method may further include: the electronic device cuts out the main object in the first video to separate the main object from the background.
[0022] In combination with the first aspect, in some embodiments, an electronic device performs a cutout on a first video for a main object, specifically including: compressing information on a target video frame in the video to obtain a first image feature; the video is a video clip; based on the first image feature, segmenting the target video frame to obtain features of a contour mask image of the object obtained during the segmentation process and a first contour mask image of the object in the target video frame; the object is the main object; performing feature reconstruction based on the first image feature and features of the obtained contour mask image, fusing the reconstructed features with first hidden state information of the video and updating the first hidden state information to obtain a fusion result, wherein the first hidden state information represents: fusion features of a transparency mask image of an object edge in a video frame that is cutout before the target video frame; based on the fusion result, obtaining a target transparency mask image of an object edge in the target video frame; performing a region cutout on the target video frame according to the target transparency mask image and the first contour mask image to obtain a cutout result.
[0023] In combination with the first aspect, in some embodiments, based on the fusion result, a target transparency mask image of the edge of the object in the target video frame is obtained, including: based on the fusion result, obtaining a second contour mask image of the object in the target video frame; fusing the second contour mask image and the fusion result to obtain a target fusion feature; based on the target fusion feature, obtaining a target transparency mask image of the edge of the object in the target video frame.
[0024] In combination with the first aspect, in some embodiments, based on the first image feature, the target video frame is segmented to obtain the features of the contour mask image of the object obtained in the segmentation process and the first contour mask image of the object in the target video frame, including: performing cascade feature reconstruction based on the first image feature to obtain the features of the contour mask images of the object with successively increasing scales, and obtaining the first contour mask image of the object in the target video frame based on the features obtained by the last processing; the first hidden state information includes: a plurality of first sub-hidden state information, each of which represents the fusion features of a transparency mask image of a scale;
[0025] Based on the first image feature and the feature of the obtained contour mask image, feature reconstruction is performed, and the reconstructed feature and the first hidden state information of the video are fused and the first hidden state information is updated to obtain a fusion result, including: performing a preset number of transparency information fusions in the following manner, and determining the feature obtained by the last processing as the fusion result: based on the first target feature and the second target feature in the feature of the obtained contour mask image, feature reconstruction is performed to obtain a second image feature with an increased scale, wherein the first target feature is the first image feature when the information fusion is performed for the first time, and the first target feature is the feature obtained by the last information fusion when the information fusion is performed for other times, and the scale of the first target feature is the same as the scale of the second target feature; the second image feature and the first sub-state information in the first hidden state information are fused and the first sub-state information is updated to obtain a third image feature, wherein the scale of the transparency mask image corresponding to the fusion feature represented by the first sub-state information is the same as the scale of the second image feature.
[0026] In combination with the first aspect, in some embodiments, feature reconstruction is performed based on the first target feature and the second target feature in the features of the obtained contour mask image to obtain a second image feature with an enlarged scale, including: screening representative features of the edge of the object from the second target feature in the features of the obtained contour mask image; and feature reconstruction is performed based on the first target feature and the representative feature to obtain a second image feature with an enlarged scale.
[0027] In combination with the first aspect, in some embodiments, cascade feature reconstruction is performed based on the first image feature to obtain features of a contour mask image of the object with successively increasing scales, including: performing contour information fusion a preset number of times in the following manner to obtain features of a contour mask image of the object with successively increasing scales: performing feature reconstruction based on a third target feature to obtain a fourth image feature with an increasing scale, wherein the third target feature is the first image feature when information fusion is performed for the first time, and the third target feature is the feature obtained by the previous feature reconstruction when information fusion is performed for other times; fusing the fourth image feature with the second sub-state information in the second hidden state information and updating the second sub-state information to obtain features of the contour mask image of the object, wherein the second hidden state information represents: a fused feature of the contour mask image of the object in a video frame segmented before the target video frame, and the second hidden state information includes: a plurality of second sub-hidden state information, each second sub-hidden state information represents a fused feature of a contour mask image of a scale, and the scale of the contour mask image corresponding to the fused feature represented by the second sub-state information is the same as the scale of the fourth image feature.
[0028] In combination with the first aspect, in some embodiments, the first image feature includes multiple first sub-image features. The target video frame in the video is compressed to obtain the first image feature, including: cascading information compression of the target video frame in the video to obtain first sub-image features with successively decreasing scales; when the feature reconstruction is performed for the first time, the third target feature is the first sub-image feature with the smallest scale. Feature reconstruction is performed based on the third target feature to obtain a fourth image feature with an increased scale, including: when the feature reconstruction is performed for other times, feature reconstruction is performed based on the third target feature and the first sub-image feature with the same scale as the third target feature to obtain a fourth image feature with an increased scale.
[0029] In combination with the first aspect, in some embodiments, the second image feature and the first sub-state information in the first hidden state information are merged and the first sub-state information is updated to obtain the third image feature, including: segmenting the second image feature to obtain the second sub-image feature and the third sub-image feature; fusing the second sub-image feature and the first sub-state information in the first hidden state information and updating the first sub-state information to obtain the fourth sub-image feature; splicing the fourth sub-image feature and the third sub-image feature to obtain the third image feature;
[0030] The fourth image feature and the second sub-state information in the second hidden state information are merged and the second sub-state information is updated to obtain the feature of the contour mask image of the object, including: dividing the fourth image feature to obtain the fifth sub-image feature and the sixth sub-image feature; fusing the fifth sub-image feature and the second sub-state information in the second hidden state information and updating the second sub-state information to obtain the seventh sub-image feature; splicing the seventh sub-image feature and the sixth sub-image feature to obtain the feature of the contour mask image of the object.
[0031] In combination with the first aspect, in some embodiments, information compression is performed on a target video frame in a video to obtain a first feature, including: inputting the target video frame in the video into an information compression network in a pre-trained video matting model to obtain a first image feature output by the information compression network, wherein the video matting model further includes: a first image generation network, a second image generation network, a result output network, a plurality of groups of contour feature generation networks and transparency feature generation networks of the same number, each group of contour feature generation networks corresponds to a scale of a contour mask image, including a first reconstruction subnetwork and a first fusion subnetwork, and each group of transparency feature generation networks corresponds to a scale of a transparency mask image, including a second reconstruction subnetwork and a second fusion subnetwork;
[0032] Performing feature reconstruction based on the first target feature and the second target feature in the feature of the obtained contour mask image to obtain a scale-enlarged second image feature, including: inputting the first target feature and the second target feature in the feature of the obtained contour mask image into a target second reconstruction subnetwork in a target transparency feature generation network to obtain a scale-enlarged second image feature output by the target second reconstruction subnetwork, wherein the scale of the transparency mask image corresponding to the target transparency feature generation network is the same as the scale of the second image feature;
[0033] The second image feature and the first sub-state information in the first hidden state information are fused and the first sub-state information is updated to obtain a third image feature, including: inputting the second image feature into a target second fusion sub-network in a target transparency feature generation network, so that the target second fusion sub-network fuses the second image feature and the first sub-state information provided by itself and updates the first sub-state information, to obtain a third image feature output by the target second fusion sub-network;
[0034] Based on the fusion result, a target transparency mask image of the edge of the object in the target video frame is obtained, including: inputting the fusion result into a second image generation network to obtain a target transparency mask image of the object in the target video frame output by the second image generation network;
[0035] Performing feature reconstruction based on the third target feature to obtain a fourth image feature with an increased scale, including: inputting the third target feature into a target first reconstruction subnetwork in a target contour feature generation network to obtain a fourth image feature with an increased scale output by the target first reconstruction subnetwork, wherein the scale of the contour mask image corresponding to the target contour feature generation network is the same as the scale of the fourth image feature;
[0036] The fourth image feature and the second sub-state information in the second hidden state information are fused and the second sub-state information is updated to obtain the feature of the contour mask image of the object, including: inputting the fourth image feature into the target first fusion sub-network in the target contour feature generation network, so that the target first fusion sub-network fuses the fourth image feature and the second sub-state information provided by itself and updates the second sub-state information, and obtains the feature of the contour mask image of the object output by the target first fusion sub-network;
[0037] Based on the features obtained by the last processing, a first contour mask image of the object in the target video frame is obtained, including: inputting the features obtained by the last processing into a first image generation network, and obtaining a first contour mask image of the object in the target video frame output by the first image generation network;
[0038] According to the target transparency mask image and the first contour mask image, the target video frame is subjected to regional cutout to obtain the cutout result, including: inputting the target transparency mask image and the first contour mask image into the result output network, so that the result output network performs regional cutout on the target video frame based on the obtained image, and obtains the cutout result output by the result output network.
[0039] In conjunction with the first aspect, in some embodiments, the transparency feature generation network further includes a feature screening subnetwork;
[0040] Before inputting the first target feature and the second target feature in the features of the obtained contour mask image into the target second reconstruction subnetwork in the target transparency feature generation network, the method further includes: inputting the second target feature in the features of the obtained contour mask image into the target feature screening subnetwork in the target transparency feature generation network, and obtaining a target screening feature that is representative of the edge contour of the object in the second target feature output by the target feature screening subnetwork;
[0041] Inputting the first target feature and the second target feature of the obtained contour mask image into the target second reconstruction subnetwork in the target transparency feature generation network, including: inputting the first target feature and the target screening feature into the target second reconstruction subnetwork in the target transparency feature generation network.
[0042] In combination with the first aspect, in some embodiments, the first fusion subnetwork is: a gated recurrent unit GRU or a long short-term memory LSTM unit; and / or, the second fusion subnetwork is: a gated recurrent unit GRU or a long short-term memory LSTM unit; and / or, the first reconstruction subnetwork is implemented based on the QARepVGG network structure, or, the first reconstruction subnetwork in the specific contour feature generation network is implemented based on the QARepVGG network structure, wherein the specific contour feature generation network is: a contour feature generation network whose scale of the corresponding contour mask image is smaller than the first preset scale; and / or, the second reconstruction subnetwork is implemented based on the QARepVGG network structure, or, the second reconstruction subnetwork in the specific transparency feature generation network is implemented based on the QARepVGG network structure, wherein the specific transparency feature generation network is: a transparency feature generation network whose scale of the corresponding transparency mask image is smaller than the second preset scale.
[0043] In combination with the first aspect, in some embodiments, the video cutout model is trained in the following manner:
[0044] Inputting a first sample video frame in the sample video into an initial model of the video matting model for processing, and obtaining a first sample contour mask image of the object in the first sample video frame output by a first image generation network in the initial model;
[0045] Obtaining a first difference between the annotated mask image corresponding to the first sample video frame and the annotated mask image corresponding to the second sample video frame, wherein the second sample video frame is: a video frame in the sample video that is before the first sample video frame and is separated by a preset number of frames;
[0046] Obtaining a second difference between the first sample contour mask image and the second sample contour mask image, wherein the second sample contour mask image is: a mask image output by the first image generation network when the initial model processes the second sample video frame;
[0047] Obtaining a third difference between the first sample transparency mask image and the second sample transparency mask image, wherein the first sample transparency mask image is: a mask image output by the second image generation network when the initial model processes the first sample video frame, and the second sample transparency mask image is: a mask image output by the second image generation network when the initial model processes the second sample video frame;
[0048] Calculate training loss based on first difference, second difference and third;
[0049] Based on the training loss, the model parameters of the initial model are adjusted to obtain the video cutout model.
[0050] In combination with the first aspect, in some embodiments, the first sample contour mask image includes: a first mask sub-image identifying an area where an object is located in a first sample video frame and a second mask sub-image identifying an area outside the object in the first sample video frame; the second sample contour mask image includes: a third mask sub-image identifying an area where an object is located in the second sample video frame and a fourth mask sub-image identifying an area outside the object in the second sample video frame; obtaining a second difference between the first sample contour mask image and the second sample contour mask image includes: obtaining a difference between the first mask sub-image and the third mask sub-image, and obtaining a difference between the second mask sub-image and the fourth mask sub-image, to obtain a second difference including the obtained differences.
[0051] In combination with the first aspect, in some embodiments, information compression is performed on a target video frame in a video to obtain a first feature, including: performing a convolution transformation on the target video frame in the video to obtain a fifth image feature; performing a linear transformation on the fifth image feature based on a convolution kernel to obtain a sixth image feature; performing batch normalization processing on the sixth image feature to obtain a seventh image feature; performing a nonlinear transformation on the seventh image feature to obtain an eighth image feature; and performing a linear transformation on the eighth image feature based on the convolution kernel to obtain the first feature of the target video frame.
[0052] In combination with the first aspect, in some embodiments, the convolution kernel is: a 1x1 convolution kernel; and / or, performing a nonlinear transformation on the seventh image feature to obtain an eighth image feature, including: performing a nonlinear transformation on the seventh image feature based on a RELU activation function to obtain the eighth image feature.
[0053] In a second aspect, the present application provides an electronic device, comprising a processor and a memory; wherein the memory is coupled to the processor, the memory is used to store computer program code, the computer program code comprises computer instructions, and when the processor executes the computer instructions, the electronic device executes the method described in the first aspect and any possible implementation method of the first aspect.
[0054] In a third aspect, an embodiment of the present application provides a chip system, which is applied to an electronic device, and the chip system includes one or more processors, which are used to call computer instructions so that the electronic device executes the method described in the first aspect and any possible implementation method of the first aspect.
[0055] In a fourth aspect, the present application provides a computer-readable storage medium, including a computer executable program. When the above-mentioned executable program runs on an electronic device, the above-mentioned electronic device executes the method described in the first aspect and any possible implementation method of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1A setting interface for turning off the screen provided in an embodiment of the present application is shown;
[0057] Figure 2 The overall process of the wallpaper display method provided by the embodiment of the present application is shown;
[0058] Figure 3A The wallpaper provided in the embodiment of the present application is exemplarily shown;
[0059] Figure 3B It shows that the wallpaper provided by the embodiment of the present application is generated based on the video clip;
[0060] Figure 4A A method of sequentially arranging video frames used by AOD wallpaper, lock screen wallpaper, and desktop wallpaper is shown;
[0061] Figure 4B Another way of sequentially arranging video frames used by AOD wallpaper, lock screen wallpaper, and desktop wallpaper is shown;
[0062] Figure 5 A method for integrating wallpaper and user interface provided by an embodiment of the present application is shown;
[0063] Figure 6 A method flow for editing a wallpaper background (such as changing the background) provided in an embodiment of the present application is shown;
[0064] Figure 7 The new wallpaper after the background is changed is displayed on the AOD, lock screen or desktop interface;
[0065] Figure 8 A method flow for editing a foreground (such as adding decorations) provided by an embodiment of the present application is shown;
[0066] Fig. 9 The new wallpaper that adds decorations to the main object is shown on the AOD, lock screen or desktop interface;
[0067] Figures 10A-10H The human-computer interaction for a user to select a video clip to generate a wallpaper is exemplarily shown;
[0068] Figures 11A-11D The human-computer interaction of the subject object for the user to edit the wallpaper is exemplified;
[0069] Figure 12A-12E The human-computer interaction for a user to edit a wallpaper background is exemplarily shown;
[0070] Figure 13A-13B The human-computer interaction for a user to change the position of a subject object is exemplarily shown;
[0071] Figure 14A-14BThe human-computer interaction for the user to change the scaling ratio of the wallpaper subject and background is exemplified;
[0072] Fig.15 The first video cutout method provided by the embodiment of the present application is shown;
[0073] Fig.16 An image change provided by an embodiment of the present application is shown;
[0074] Fig.17 A second video cutout method provided in an embodiment of the present application is shown;
[0075] Fig.18 The first feature processing method provided by the embodiment of the present application is shown;
[0076] Fig.19 A second feature processing method provided in an embodiment of the present application is shown;
[0077] Fig. 20 The first cascade feature reconstruction method provided by the embodiment of the present application is shown;
[0078] Fig.21 The second cascade feature reconstruction method provided by the embodiment of the present application is shown;
[0079] Fig. 22 A third video cutout method provided in an embodiment of the present application is shown;
[0080] Fig.23 The structure of the first video cutout model provided in the embodiment of the present application is shown;
[0081] Fig.24 The structure of an information compression network provided by an embodiment of the present application is shown;
[0082] Fig.25 The structure of the second video cutout model provided in the embodiment of the present application is shown;
[0083] Fig.26 A fourth video cutout method provided in an embodiment of the present application is shown;
[0084] Fig. 27 The structure of the third video cutout model provided in the embodiment of the present application is shown;
[0085] Fig.28 The structure of the fourth video cutout model provided in the embodiment of the present application is shown;
[0086] Fig.29 The first model training method provided by the embodiment of the present application is shown;
[0087] Fig.30A second model training method provided in an embodiment of the present application is shown;
[0088] Fig.31 A mask provided by an embodiment of the present application is shown;
[0089] Fig.32 The structure of a second image generation network provided by an embodiment of the present application is shown;
[0090] Fig.33 A comparison of a cutout result provided by an embodiment of the present application is shown;
[0091] Fig.34 An electronic device provided by an embodiment of the present application is shown;
[0092] Fig.35 A software system architecture applied to an electronic device provided by an embodiment of the present application is shown;
[0093] Fig.36 A chip system provided by an embodiment of the present application is shown. DETAILED DESCRIPTION
[0094] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application.
[0095] As described in the background technology section, mobile phones and other electronic devices can provide wallpaper functions to support users to set wallpapers in AOD, lock screen, desktop, etc. One wallpaper function can support users to set wallpapers independently in multiple places such as AOD, lock screen, desktop, etc., but the AOD wallpaper, lock screen wallpaper, desktop wallpaper and other wallpapers are relatively separated, and no coherent display effect of wallpaper is provided. Another wallpaper function, also known as "super wallpaper" and "screen-off one-shot", can provide a coherent and dynamic wallpaper display effect from AOD to lock screen and then to desktop, but the wallpaper transition effect is relatively stiff and the user experience is limited.
[0096] Herein, the AOD interface is an interface displayed by an electronic device when the screen is off, which can be used to display information such as time, date, text messages and call reminders, which can save user operations and is very intuitive and convenient. In the embodiment of the present application, the AOD interface may also include an AOD wallpaper. To provide an AOD interface, the screen of the electronic device needs to have pixel-level luminescence capability and support lighting only some pixels of the screen, such as an organic light-emitting diode (OLED) screen. The lock screen interface is an interface displayed by an electronic device when the screen is on but not unlocked, and may include icons of system status bar, time, date, weather, lock screen logo, unlock prompt, and shortcut functions (such as flashlight, camera, etc.) that do not require user authorization, etc. The desktop is an interface displayed by an electronic device when the screen is on and unlocked, which can also be called the main interface, and may include desktop icons of various applications, system status bar, etc. Not limited to the desktop, the interface displayed by the electronic device when the screen is on and unlocked can also be other interfaces, such as the user interface of the application, the user interface displayed in the last unlocked state (before the lock screen). The interfaces displayed in the bright screen and unlocked state are collectively referred to as unlock interfaces herein.
[0097] like Figure 1 As shown, the electronic device can provide an AOD function switch in the user interface of the setting application, such as providing a "screen off display" switch in the "screen off display" interface of the setting application. The embodiment of the present application does not limit the interface presentation of the AOD function switch, whether it can also appear in other application interfaces, etc. The user can choose whether to use the AOD function by operating the switch. When the AOD function switch is turned on, the electronic device can display the AOD interface when the screen is off, otherwise the AOD interface will not be displayed. Figure 1 As shown, when the AOD function switch is turned on, the electronic device can also display the two screen-off modes of "full screen" and "partial" in the "screen-off display" interface. Among them, in the "full screen" screen-off mode, the entire screen area displays the AOD interface with a darkened effect to save screen power consumption; in the "partial" screen-off mode, only part of the screen is lit to display the AOD interface. In some embodiments, in the "full screen" screen-off mode, the AOD wallpaper in the AOD interface can be replaced by setting the lock screen wallpaper, that is, the picture content of the AOD wallpaper is the same as the picture content of the lock screen wallpaper.
[0098] The embodiment of the present application provides a wallpaper display method, which can achieve a dynamic wallpaper display effect with multiple levels of interfaces such as AOD, lock screen and desktop, and there is no need to design an additional transition method or transition method from AOD to lock screen, lock screen to desktop for the wallpaper. The transition effect between them is built-in to the video, which is more vivid and natural. Moreover, the wallpaper is well integrated with interfaces such as AOD, lock screen, and desktop, and the main object in the wallpaper will not be blocked by non-interactive elements such as time in the interface. In addition, the wallpaper can be generated based on user customization or editing operations, which is convenient for users to set personalized wallpapers, and the user operation is simple, and one editing operation can take effect on the entire dynamic wallpaper.
[0099] Figure 2 The overall process of the wallpaper display method provided by the embodiment of the present application is shown. Figure 2 As shown, the method may include the following steps:
[0100] S11, the electronic device in the locked state or unlocked state detects a first operation, and in response to the first operation, the electronic device displays an AOD interface. The AOD interface includes an AOD wallpaper. In addition to the AOD wallpaper, the AOD interface can also display brief information such as date, time, and weather.
[0101] The first operation may be a screen-off operation, such as pressing the power button to turn off the screen of an electronic device in a locked or unlocked state. In the locked or unlocked state, the screen of the electronic device is lit up, wherein the lock screen interface is displayed on the screen in the locked state, and the desktop interface or the user interface of the application opened by the user is displayed on the screen in the unlocked state. Not limited to being triggered by the screen-off operation, if the user does not operate the mobile phone for a long time in the locked or unlocked state, the electronic device may also display the AOD interface.
[0102] S12, when the electronic device displays the AOD interface, the electronic device detects a second operation, such as a user operation of lighting up the screen.
[0103] Specifically, the user operation of lighting up the screen in the AOD interface may be an operation of lightly touching the screen, pressing the power button, etc., but is not limited thereto. The screen lighting up in the AOD interface may also be a scheduled arrival, and the user may set the time to light up the screen, and when the time arrives, the electronic device lights up the screen. The embodiment of the present application does not limit the operation of lighting up the screen in the AOD interface.
[0104] S13, in response to the second operation, the electronic device displays a lock screen interface, where the lock screen interface includes a lock screen wallpaper.
[0105] S14, the electronic device detects a third operation for unlocking, and the third operation unlocks the screen successfully.
[0106] S15, in response to the third operation, the electronic device displays a desktop interface, where the desktop interface includes a first desktop wallpaper.
[0107] Among them, AOD wallpaper, lock screen wallpaper and desktop wallpaper can respectively use the first video frame, the second video frame and the third video frame in a video, and the second video frame is after the first video frame, and the third video frame is after the second video frame. As the user switches from the AOD interface to the lock screen interface and then to the desktop, a coherent and dynamic wallpaper playback screen is formed.
[0108] In the embodiment of the present application, the AOD interface, the lock screen interface, and the desktop interface may be respectively referred to as the first interface, the second interface, and the third interface; the video (or video clip) used to generate the dynamic wallpaper may be referred to as the first video.
[0109] AOD wallpapers, lock screen wallpapers and desktop wallpapers can directly use the original video frames in the first video, or use image frames that have been processed by cropping, filtering, etc. on the original video frames.
[0110] Figure 3A The wallpaper provided in the embodiment of the present application is exemplified. Figure 3A As shown, (a), (b), and (c) respectively show the display of AOD wallpaper 11, lock screen wallpaper 12, and desktop wallpaper 13. AOD wallpaper 11, lock screen wallpaper 12, and desktop wallpaper 13 are linked in sequence to form a dynamic wallpaper. Figure 3A As shown, from AOD to lock screen and then to desktop, AOD wallpaper 11, lock screen wallpaper 12 and desktop wallpaper 13 are linked together, and they are no longer separated from each other.
[0111] The video clip used to generate the wallpaper can be selected by the user. Figure 3B As shown, the user can select a video clip 14 to generate a wallpaper, and the video clip 14 comes from a video 15. The video clip 14 can also be selected from a live photo.
[0112] later Figures 10A-10H The embodiment will introduce in detail the human-computer interaction method for users to select video clips to generate wallpapers, which will not be expanded here.
[0113] Not limited to from AOD to lock screen and then to desktop, the wallpaper generated by the embodiment of the present application can also be applied to other adjacent interface jump scenarios, such as AOD to desktop and then to certain application interfaces, which will especially occur when the user has not set a lock screen, and for example, from desktop to application main interface and then to application secondary interface. Of course, these application interfaces can support wallpaper display, rather than game display interface, video playback interface and other application interfaces that are not suitable for wallpaper display. For example, the chat interface of "WeChat" can set chat wallpaper. That is, the electronic device can display multiple user interfaces of different levels in turn, and display wallpapers at these multiple user interfaces; the wallpaper is generated based on the video clip selected by the user, and the video frames used by the wallpapers at these multiple user interfaces are arranged in sequence, forming a coherent and dynamic wallpaper display effect as a whole when the user performs multi-level interface jumps.
[0114] The wallpaper display solution provided in the embodiment of the present application will be further explained below from the aspects of wallpaper linkage, main object highlighting, background editing, and main object editing.
[0115] Wallpaper linkage
[0116] Since the video frames used in the AOD interface, lock screen interface, and desktop wallpapers are arranged sequentially, the embodiment of the present application does not need to additionally design a transition method or transition method from AOD to lock screen, and from lock screen to desktop for the wallpaper. The transition effect between them is inherent in the video clips, which is more vivid and natural.
[0117] As an example, the video frames used in the AOD interface, lock screen interface, and desktop wallpaper in this application can be arranged in the following order:
[0118] The first implementation method
[0119] The AOD wallpaper 11, the lock screen wallpaper 12, and the desktop wallpaper 13 may all be static wallpapers, that is, the video frame used by the AOD wallpaper, the video frame used by the lock screen wallpaper, and the video frame used by the desktop wallpaper are all one frame.
[0120] like Figure 4A As shown, during the duration of the AOD interface (t1 to t2), the video frame used by the AOD wallpaper 11 is always frame "1", such as the AOD wallpaper 11 specifically only uses the face image in frame "1"; during the duration of the lock screen interface (t2 to t3), the video frame used by the lock screen wallpaper 12 is always frame "2"; during the duration of the desktop (t3 to t4), the video frame used by the desktop wallpaper 13 is always frame "3". Frame "1", frame "2", and frame "3" are several video frames arranged in sequence in the video clip, wherein frame "1" is before frame "2", and frame "2" is before frame "3".
[0121] The frame interval between the first video frame and the second video frame may be greater than the first frame interval, and the frame interval between the second video frame and the third video frame may be greater than the second frame interval. For example, the first video frame, the second video frame, and the third video frame may be extracted from the first video at a certain frame interval (e.g., more than 1 second). In this way, the images of these multiple video frames may be different, and a dynamic change effect may be presented when they are played continuously.
[0122] With the first implementation method, although the wallpaper seen by the user when he / she stays at a certain place such as AOD, lock screen or desktop is static, as the user switches from AOD to lock screen and then to desktop, AOD wallpaper 11, lock screen wallpaper 12 and desktop wallpaper 13 are displayed in sequence, which will produce a coherent dynamic effect as a whole.
[0123] Moreover, in the embodiment of the present application, when the user switches interfaces in reverse, that is, from the desktop to the lock screen and then to the AOD interface, a coherent dynamic effect will be produced as a whole.
[0124] The second implementation method
[0125] The AOD wallpaper 11, the lock screen wallpaper 12, and the desktop wallpaper 13 can be dynamic wallpapers alone, that is, the video frames used by the AOD wallpaper, the lock screen wallpaper, and the desktop wallpaper are all multiple frames.
[0126] When the user stays at a certain place such as AOD, lock screen or desktop, the electronic device plays the dynamic wallpaper at that place, and the wallpaper seen by the user is dynamic. In order to achieve a coherent dynamic effect from AOD to lock screen and then to desktop as a whole, the video frames used by AOD wallpaper 11, lock screen wallpaper 12, and desktop wallpaper 13 need to be connected end to end according to the timing of interface switching. That is: when switching from AOD to the lock screen interface, the starting frame used by the lock screen wallpaper 12 can be the next frame of the ending frame used by the AOD wallpaper 11; when switching from the lock screen interface to the description interface, the starting frame used by the desktop wallpaper 13 can be the next frame of the ending frame used by the lock screen wallpaper 12. In other words, in two user interfaces of adjacent levels, the starting video frame used by the wallpaper at the user interface of the next level can be the next frame of the ending video frame used by the wallpaper at the user interface of the previous level.
[0127] For example, assume that the AOD wallpaper 11, the lock screen wallpaper 12, and the desktop wallpaper 13 are all generated by frame "1", frame "2", and frame "3". Figure 4B As shown, from AOD to lock screen and then to desktop, the stages and timings of a round of wallpaper display can be as follows:
[0128] t1 to t2: Play AOD wallpaper 11. Figure 4BAs shown, if the time from t1 to t2 is very long, the user stays on the AOD interface for a long time and does not press the power button to light up the screen, then the AOD wallpaper 11 can be played in a loop.
[0129] t2: Detecting that the user switches from the AOD interface to the lock screen interface, the AOD wallpaper 11 ends playing, and the lock screen wallpaper 12 begins playing. The starting frame (such as frame "3") used by the lock screen wallpaper 12 can be the next frame of the ending frame (such as frame "2") used by the AOD wallpaper 11.
[0130] t2 to t3: Play lock screen wallpaper 12. Figure 4B As shown, if the time from t2 to t3 is very short and the user quickly unlocks and successfully enters the desktop, the lock screen wallpaper 12 will not be played in a loop; otherwise, the lock screen wallpaper 12 will be played in a loop.
[0131] t3: Detecting that the user switches from the lock screen interface to the desktop, the lock screen wallpaper 12 ends playing, and the desktop wallpaper 13 begins playing. The starting frame (such as frame "3") used by the desktop wallpaper 13 is the next frame of the ending frame (such as frame "2") used by the lock screen wallpaper 12.
[0132] t3 to t4: Play desktop wallpaper 13. Figure 4B As shown, if the time from t1 to t2 is very long, and the user stays on the desktop for a long time and does not press the power button to turn off the screen, the AOD wallpaper 11 can be played in a loop.
[0133] t4: Detect that the user switches from the desktop to the AOD interface, end the desktop wallpaper 13, start playing the AOD wallpaper 11, and enter the next round of wallpaper display. The starting frame (such as frame "1") used by the next round of AOD wallpaper 11 is the next frame of the ending frame (such as frame "3") used by the previous round of desktop wallpaper 13. Of course, the user may also switch from other application interfaces (such as the video playback interface) to the AOD interface. At this time, the starting frame used by the next round of AOD wallpaper 11 can be frame "1" by default, that is, the first frame in the video clip.
[0134] Moreover, in the embodiment of the present application, when the user switches the interface in reverse, that is, from the desktop to the lock screen and then to the AOD interface, a coherent dynamic effect will be produced as a whole. Specifically, it can be connected end to end according to the timing of the interface switching. That is: when switching from the desktop to the lock screen, the starting frame used by the lock screen wallpaper can be the next frame of the ending frame used by the desktop wallpaper; when switching from the lock screen to the AOD, the starting frame used by the AOD wallpaper can be the next frame of the ending frame used by the lock screen wallpaper.
[0135] like Figure 4BAs shown, the frame rates of the AOD wallpaper 11, the lock screen wallpaper 12 and the desktop wallpaper 13 can be the same, such as changing the video frame used twice every 1 second. But not limited to this, their frame rates can also be different. For example, the frame rate of the AOD dynamic wallpaper can be larger, such as changing the video frame used 5 times every 1 second, to reduce the risk of damage to the fixed area of the screen due to long-term lighting. Originally, the fact that the AOD wallpaper 11 is dynamic can reduce this risk compared to the static AOD wallpaper 11. Figure 4A and Figure 4B In the figure, “T” represents the screen refresh rate. Figure 4A and Figure 4B Just as an example, the wallpaper frame rate can actually be greater than T, such as changing the video frame used by the wallpaper every 10 T; moreover, the duration of AOD can be much longer than several T; based on human reaction delay and operation delay, the duration of lock screen wallpaper and desktop wallpaper can also be much longer than several T.
[0136] Highlight the main object
[0137] In the embodiment of the present application, the wallpaper can be well integrated with interfaces such as AOD, lock screen, and desktop, and the main objects (such as characters) in the wallpaper will not be blocked by non-interactive elements such as time in interfaces such as AOD, lock screen, and desktop.
[0138] Figure 5 A method for integrating wallpaper and user interface provided by an embodiment of the present application is shown. Figure 5 As shown, the electronic device can first cut out the main object of the video clip used to generate the wallpaper, separate the main object and its background, and obtain the foreground layer and the background layer. The foreground layer can be a layer of the main object (such as a person), and the background layer is a layer of image content other than the main object. Then, the electronic device can insert layers of non-interactive elements such as time between the foreground layer and the background layer, synthesize these layers, and render interfaces such as AOD, lock screen or desktop. In this way, the image of the main object will be displayed on non-interactive elements such as time, and will not be blocked by non-interactive elements such as time, which can better show the style of the main object in the wallpaper.
[0139] The layers of interactive elements in interfaces such as AOD, lock screen or desktop, such as the fingerprint unlocking key in the AOD interface, the password keyboard in the lock screen interface, and the application icons in the desktop, can be set on the foreground layer to facilitate users to operate these interactive elements and ensure the realization of the interface interaction function. That is, the layers of interfaces such as AOD, lock screen or desktop can be decomposed into layers of non-interactive elements and layers of interactive elements.
[0140] Regarding how to cut out a main object in a video, please refer to the subsequent embodiments, which will not be expanded here.
[0141] In addition to the main object not being blocked by non-interactive elements such as time, the cutout electronic device can further provide background editing and foreground editing functions.
[0142] Background
[0143] Figure 6 The following is a method flow for editing wallpaper background (such as changing the background) provided by an embodiment of the present application. Figure 6 As shown in FIG. 1 , after obtaining the foreground layer and background layer of the wallpaper by cutting out, the electronic device can replace the background layer, and then insert a layer of non-interactive elements such as time between the foreground layer and the new background layer, synthesize these layers, and render an AOD, lock screen, or desktop interface for changing the wallpaper background. Figure 7 As shown, the wallpaper displayed in the AOD, lock screen or desktop interface completes the background change.
[0144] Further, in order to simplify the editing operation and improve the efficiency of background editing, after detecting that the user edits the background of the wallpaper at a certain user interface (such as the lock screen interface), the electronic device can synchronously apply the editing to the wallpaper at other user interfaces, that is, the same editing operation is also performed on the wallpaper background at other user interfaces, so that one editing operation takes effect on the entire wallpaper. Here, other user interfaces refer to user interfaces other than the user interface (such as the lock screen interface) that receives the editing operation in the multi-level user interface (such as the AOD interface, the lock screen interface, and the desktop) that can display the wallpaper. In other words, the editing operation implemented by the user can be an operation in which the user selects the wallpaper background at a certain user interface in the multi-level user interface for background editing; not limited to updating the wallpaper background at the certain user interface, the electronic device can also update the background image at other user interfaces. That is to say, in response to the user's operation of editing the wallpaper background at a certain user interface, the electronic device can update the wallpaper background image at the multi-level user interface.
[0145] Not limited to changing the background, editing the background may also include changing the background scale, blurring the background, defocusing the background, changing the background color, brightness, and so on. The electronic device may perform corresponding image processing on the background layer separately, such as image scaling, image blurring, image defocusing, adjusting image brightness, changing image color, etc., to update the background layer, thereby obtaining a new wallpaper, and applying the new wallpaper on the AOD interface, lock screen interface, and desktop. Here, the specific meaning of application may refer to synthesizing the background layer, foreground layer of the new wallpaper and the non-interactive element layers and interactive element layers of the AOD interface, lock screen interface, and desktop, and rendering the screen displayed by the new wallpaper on these interfaces.
[0146] later Figure 12A-12E The embodiment will exemplarily introduce the human-computer interaction method for users to edit wallpaper backgrounds, which will not be expanded here.
[0147] Main object
[0148] Figure 8 The figure shows a method flow of editing the foreground (such as adding decorations) provided by an embodiment of the present application. Figure 8 As shown in FIG. 1 , after obtaining the foreground layer and background layer of the wallpaper by cutting out, the electronic device can add an image of a decoration (such as a hat) to the foreground layer to obtain a new foreground layer, and then insert a layer of non-interactive elements such as time between the new foreground layer and the background layer, synthesize these layers, and render an AOD, lock screen, or desktop interface decorated with the main object of the wallpaper. Fig. 9 As shown, the wallpaper displayed in the AOD, lock screen or desktop interface realizes the decoration of the main object. The specific implementation of adding decorations to the main object can be layer synthesis, such as synthesizing the decoration layer and the foreground layer to obtain a new foreground layer.
[0149] Further, in order to simplify the editing operation and improve the efficiency of editing the main object, after detecting that the user edits the main object of the wallpaper at a certain user interface (such as the lock screen interface), the electronic device can synchronously apply the editing to the wallpaper at other user interfaces, that is, the same editing operation is also performed on the wallpaper foreground object at other user interfaces, so as to achieve a main object editing operation that takes effect on the entire wallpaper. Here, other user interfaces refer to user interfaces other than the user interface (such as the lock screen interface) that receives the editing operation in the multi-level user interface (such as the AOD interface, the lock screen interface, and the desktop) that can display the wallpaper. In other words, the editing operation implemented by the user can be an operation in which the user selects the wallpaper main object at a certain user interface in the multi-level user interface to edit the main object; not limited to updating the wallpaper main object at the certain user interface, the electronic device can also update the main object image at other user interfaces. That is to say, in response to the user's operation of editing the wallpaper main object at a certain user interface, the electronic device can update the main object image of the wallpaper at the multi-level user interface.
[0150] Not limited to adding decorations, editing the foreground may also include changing the scale of the main object, changing the position of the main object, changing the color and brightness of the main object, etc. The electronic device may perform corresponding image processing on the foreground layer alone, such as image scaling, changing the position of the main object image in the foreground layer, adjusting the image brightness, changing the image color, etc., to update the foreground layer containing the main object image, thereby obtaining a new wallpaper, and applying the new wallpaper to the AOD interface, lock screen interface, and desktop.
[0151] later Figures 11A-11D The embodiment will exemplarily introduce the human-computer interaction method for the user to edit the subject object, which will not be expanded here.
[0152] Electronic devices can also edit the wallpaper as a whole, such as adding filters and changing the scaling of the main object and its background at the same time. Editing as a whole means that the editing effect takes effect on the wallpaper as a whole, not limited to the foreground or background. For example, when synthesizing layers, add a filter layer on top of the wallpaper background layer, the non-interactive element layer of the lock screen interface, the wallpaper foreground layer, and the interactive element layer of the lock screen interface, so that the filter effect takes effect on the entire lock screen wallpaper. For another example, scale the main object and its background according to the selected scaling ratio to obtain new foreground and background layers, which are then synthesized with the layers of the lock screen interface so that the main object and its background in the lock screen wallpaper are scaled synchronously.
[0153] Based on the wallpaper display solution introduced above, next, a series of human-computer interaction methods provided in the embodiments of the present application are introduced.
[0154] Figures 10A-10H The human-computer interaction method for a user to select a video clip to generate a wallpaper is exemplified.
[0155] First, electronic devices can display Fig. 10A The "Desktop and Wallpaper" setting interface shown in FIG. 100 provides setting options such as "Wallpaper". After detecting that the user selects the setting option "Wallpaper" 101, the electronic device can display Fig. 10B The "wallpaper" setting interface shown in the figure provides wallpaper setting options such as "video wallpaper" 102. In this article, "video wallpaper" is just a name, which refers to the wallpaper setting function that uses video to generate wallpaper, and there is no restriction on its naming. After detecting that the user selects the "video wallpaper" option 102, the electronic device can display Fig. 10C The "Video Wallpaper" setting interface shown in the figure shows some recommended videos. The recommended videos can come from the gallery, the Internet, or some videos pre-stored locally. In order to enrich the user's choices, the electronic device can also Fig. 10C The "video wallpaper" setting interface shown provides a gallery entry "Select from gallery" 103, which can be a control (such as a button), a link, or other types of interface interaction elements in technical implementation. After detecting that the user clicks on the gallery entry, the electronic device can display Fig. 10D The interface shown above will display the videos in the gallery. To further facilitate user selection, Fig. 10D In the interface shown, the videos in the gallery can be classified, such as "people" and "architecture".
[0156] The user can select a video 16 for generating a wallpaper from a plurality of videos. Once the operation of selecting a video 16 by the user is detected, the electronic device can display the video. Fig.10E The “Video Wallpaper Preview” interface shown.
[0157] like Fig.10E As shown, the "Video Wallpaper Preview" interface may display a preview 17, which may be used to display the display status of the AOD wallpaper, lock screen wallpaper, and desktop wallpaper in the AOD, lock screen, desktop, and other interfaces, respectively. The preview may be a video, such as Figure 10F-10H As shown, when the video is played, the dynamic display process of the wallpaper from AOD to lock screen and then to the desktop can be presented. The video can be composed of multiple frames of images, wherein the front image frame can present the display state of the wallpaper on the AOD interface, the middle image frame can present the display state of the wallpaper on the lock screen interface, and the rear image frame can present the display state of the wallpaper on the desktop.
[0158] Furthermore, the preview can also present the reverse dynamic display process of the wallpaper from the desktop to the lock screen and then to the AOD.
[0159] like Fig.10E As shown, the "video wallpaper preview" interface may also include: a video frame sequence 18 of the selected video 16, and a selection box 19. The video frames in the video frame sequence may be presented in the form of video frame thumbnails. Among them, a video in the selection box 19 is the selected first video, and the first video may be part or all of the video 16. In the embodiment of the present application, the "video wallpaper preview" interface may be referred to as the fourth interface, and the video 16 may be referred to as the second video. The user may slide the selection box 19 along the video frame sequence 18 to change the selected video clip. The length of the selection box 19 may be fixed, such as a fixed selection of a video clip of 240 frames or 4 seconds. The length of the selected video clip may also be variable. The selected video clip is at least longer than the sum of the frame interval 1 and the frame interval 2, so that three video frames with a long interval can be selected from the video clip for use in AOD wallpaper, lock screen wallpaper and desktop wallpaper, respectively, so that the pictures of the multiple video frames are different, and a dynamic change effect is presented when they are played continuously. Frame interval 1 is the frame interval between video frames used by the AOD wallpaper and the lock screen wallpaper, and frame interval 2 is the frame interval between video frames used by the lock screen wallpaper and the desktop wallpaper. The interface interaction element for selecting a video segment from the video frame sequence 18 is not limited to the selection box 19, and can also be in other styles. Alternatively, there is no need to design a special control for the user to select a video segment. The user can directly select multiple consecutive frames in the video frame sequence 18 to specify a video segment.
[0160] The above content illustrates that the first video can be selected by the user from the gallery video. Not limited to this, the setting application can also provide a video for generating wallpaper, and the first video can be selected by the user from the videos provided by the setting application. The first video can also be provided by the system without user operation.
[0161] like Fig.10EAs shown, the "Video Wallpaper Preview" interface may also include: a cancel control 20, a completion control 21, a "depth of field effect" control, an "edit" control, and so on. Among them, the cancel control 20 can be used by the user to cancel the video clip in the selection box 19 from generating wallpaper; the completion control 21 can be used by the user to confirm the video clip in the selection box 19 from generating wallpaper; the "depth of field effect" control can be used by the user to adjust the depth of field of the wallpaper. The "edit" control can be the entrance to wallpaper editing. Through this entrance, the user can perform a series of wallpaper editing operations, such as adding decorations to the main object, changing the background, etc. Figures 11A-11D , Figure 12A-12E The implementation example will introduce wallpaper editing, which will not be expanded here.
[0162] After the electronic device detects that the user has changed the video clip used to generate the wallpaper, such as by sliding the selection box 19 to change the selected video clip, the electronic device can generate a new wallpaper based on the new video clip and update the preview 17. The updated preview 17 can be used to display the display status of the new wallpaper in interfaces such as AOD, lock screen, and desktop. Specifically, when the updated preview 17 is played, the dynamic display process of the new wallpaper from AOD to lock screen and then to desktop can be displayed. In this way, it is convenient for users to understand the dynamic display effect of the new wallpaper in a timely manner, so that users can make their favorite wallpapers.
[0163] The video clip used to generate the wallpaper may be a default one, such as the first three video frames of the selected video may constitute the video clip, or the video clip used to generate the wallpaper may be composed of several video frames before and after the wonderful video frame (including the wonderful frame itself). As for how to identify the wonderful video frame from the video, the embodiment of the present application does not limit it.
[0164] After the wallpaper is updated by changing the video clip, when the AOD, lock screen, and desktop are displayed again, the electronic device will apply the new wallpaper on these interfaces, that is, the updated wallpaper will also be displayed on the AOD, lock screen, and desktop, so as to ultimately achieve the purpose of changing the wallpaper.
[0165] Figures 11A-11D The human-computer interaction method for a user to edit a main object of a wallpaper is exemplified.
[0166] First, electronic devices can display Fig.11AThe "Wallpaper Editing" interface shown in the figure is displayed, and preview 23 is displayed in the interface. In an embodiment of the present application, the "Wallpaper Editing" interface can be referred to as the fifth interface. Like the aforementioned preview 17, preview 23 can be used to present the display status of AOD wallpapers, lock screen wallpapers, and desktop wallpapers at interfaces such as AOD, lock screen, and desktop, respectively. Preview 23 can be a video, which can present the dynamic display process of the wallpaper from AOD to lock screen and then to the desktop when the video is played. In addition, in preview 23, the main object of the wallpaper and its background can be displayed separately, such as indicating this separation by mark 24. In this way, the user can select the main object or the background for editing respectively. The separate display is achieved based on the cutout of the video clip that generates the wallpaper for the main object.
[0167] Specifically, the mark 24 may be a highlight boundary line designed along the boundary of the main object, or a semi-transparent layer with some dynamic effect covering the main object, etc. The embodiment of the present application does not limit the interface expression form of the mark 24.
[0168] Users can click Fig.10E Click the "Edit" control in the interface shown to enter Fig.11A The "Wallpaper Editing" interface shown in the figure is not limited thereto, and the entrance to the "Wallpaper Editing" interface can also be set at other locations, which is not limited in the embodiment of the present application.
[0169] like Figure 11A-11B As shown in FIG. 1 , in a state where the main object of the wallpaper and its background are displayed separately, the electronic device can detect the user's operation of selecting the main object, and in response, can provide the "decoration" editing option. The electronic device can further detect the user's operation of clicking the "decoration" editing option, and in response, can Fig. 11C One or more decorations are shown, such as Fig. 11C "Element 1", "Element 2", "Element 3", "Element 4" shown in.
[0170] Subsequently, the electronic device may detect the user's operation of adding a decoration (such as "element 1") to the main object, and in response thereto, the electronic device may generate a new wallpaper, and Fig.11D The updated preview 23 is shown. The main object in the new wallpaper is added with the decoration selected by the user, and the updated preview 23 can be used to show the display status of the new wallpaper in the AOD, lock screen, and desktop. Furthermore, the user can also adjust the position of the decoration, such as moving the "hat" to the top of the character's head.
[0171] Furthermore, in order to simplify editing operations and improve the efficiency of editing main objects, after detecting the user's operation of adding decorations to the main object of a wallpaper (such as an AOD wallpaper, lock screen wallpaper, or desktop wallpaper), the electronic device can synchronously apply the editing to other wallpapers, that is, add decorations to the main objects of other wallpapers in the same way, so that a main object editing operation can take effect on the entire dynamic wallpaper.
[0172] After adding decorations to the main object of the wallpaper, when the AOD, lock screen, and desktop are displayed again, the electronic device will apply the new wallpaper with the decorations added to the main object on these interfaces, that is, the new wallpaper will also be displayed on the AOD, lock screen, and desktop, so as to ultimately achieve the purpose of adding decoration to the main object of the wallpaper.
[0173] It is not limited to adding decorations to the main object. Selecting the main object for editing may also include: replacing the main object, adding filters to the main object, and other editing operations.
[0174] Figure 12A-12E The human-computer interaction method for a user to edit a wallpaper background is exemplified.
[0175] Electronic devices can display Fig. 12A The "Wallpaper Editing" interface shown in the example can be referenced for its implementation Fig.11A The relevant description of the interface shown is not repeated here.
[0176] like Figure 12A-12B As shown in FIG. 1 , in a state where the main object of the wallpaper and its background are displayed separately, the electronic device can detect the user's operation of selecting the wallpaper background, and in response thereto, can provide an editing option of "switch background". The electronic device can further detect the user's operation of clicking the editing option of "switch background", and in response thereto, Fig. 12C As shown in the example, background editing options such as "Switch background color" and "Switch background image" can be provided, and background color options such as "Color 1", "Color 2", "Color 3", "Color 4" and the background image switching option of "Select from gallery" can be provided.
[0177] The electronic device may then detect that the user has selected the option "Select from Gallery" and display Fig. 12C The example shown is a gallery calling interface, which displays photos in the gallery. Fig.12D As shown, the electronic device can further detect the operation of the user selecting photo 26 as the new wallpaper background, and in response thereto, a new wallpaper can be generated based on the new background, and as shown in FIG. Fig.12E The updated preview 23 is shown as an example. The background of the new wallpaper is changed, and the updated preview 23 can be used to show the display status of the new wallpaper in the AOD, lock screen, desktop and other interfaces.
[0178] After changing the wallpaper background, when the AOD, lock screen, and desktop are displayed again, the electronic device will apply the new wallpaper with the changed background on these interfaces, that is, the new wallpaper will also be displayed on the AOD, lock screen, and desktop, so as to ultimately achieve the purpose of changing the wallpaper background.
[0179] In addition, the electronic device can detect the user's operation of adjusting the position of the subject relative to the background, such as Fig.13A In the exemplary example of dragging the subject to the left, a new foreground layer can be generated based on the updated position of the subject, thereby generating a new wallpaper. Fig. 13B As shown in the example, the position of the main object in the new wallpaper is the adjusted position. After the position of the main object in the wallpaper is changed, when the AOD, lock screen, and desktop are displayed again, the electronic device will apply the new wallpaper with the changed position of the main object on these interfaces, that is, the new wallpaper will also be displayed on the AOD, lock screen, and desktop, so as to ultimately achieve the purpose of changing the position of the main object in the wallpaper.
[0180] The electronic device can also detect the user's operation of adjusting the main object or background scaling of the wallpaper, such as Fig.14A As shown in the example, when the main object is selected, drag the slide bar to change the zoom ratio of the main object. Fig. 14B In the exemplary embodiment, when the background is selected, the slider is dragged to change the zoom ratio of the subject. In response, the electronic device can generate a new foreground layer or background layer based on the zoom ratio selected by the user, thereby generating a new wallpaper. Fig.14A , Fig. 14B As shown, the electronic device can also detect the user's operation of adjusting the scaling ratio of the main object and the background at the same time, such as zooming in on the main object while reducing the background to form a movie visual experience. In response to this, the foreground layer and the background layer can be updated at the same time, and a new wallpaper can be generated based on the new foreground layer and the new background layer. After changing the scaling ratio of the main object and the background in the wallpaper, when the AOD, lock screen, and desktop are displayed again, the electronic device will apply the new wallpaper with the changed scaling ratio of the main object and the background at these interfaces, that is, the new wallpaper will also be displayed at the AOD, lock screen, and desktop, so as to ultimately achieve the purpose of changing the scaling ratio of the main object and the background in the wallpaper.
[0181] The embodiment of the present application also provides a video cutout method to cut out the main object in the video clip used to generate the wallpaper, separate the main object and its background, and then insert a layer of non-interactive elements such as time between the foreground layer and the background layer to avoid the main object image being blocked by non-interactive elements such as time, which can better show the style of the main object in the wallpaper.
[0182] The video cutout method provided in the embodiment of the present application can also improve the accuracy of video cutout.
[0183] First, the video cutout process is explained.
[0184] The video contains multiple video frames. The video frame to be cut out is called the target video frame. In this way, the target video frame can be any video frame in the video that needs to be cut out. Video cutout means extracting the object area with object transparency from the video frame. The above objects can be people, animals, vehicles, lane lines, etc.
[0185] In the video cutout process, first, the first target video frame is determined, which may be the first video frame or other video frames in the video, and the determined target video frame is cutout to obtain the object area in the target video frame; then, the next video frame of the target video frame is determined as a new target video frame, and the new target video frame is cutout; in this way, each time the object area in the target video frame is obtained, the next video frame is determined as the new target video frame, until the segmentation result of the last video frame of the video is obtained, and the object cutout of the entire video is completed.
[0186] Next, the application scenarios of the video cutout solution provided in the embodiments of the present application are illustrated with examples.
[0187] 1. Real-time video scene
[0188] In this scenario, the video to be played is cut out to obtain the object area in each video frame in the video, so that when the video is played, only the area content of the object area in each video frame in the video can be played.
[0189] 2. Video editing scenes
[0190] In this scenario, after the video is segmented to obtain the area where the object is located in each video frame in the video, the video frames in the video can be edited by replacing the background, erasing the object, blurring the background, retaining the color, etc. according to the location of the area where the object is located in the video frame, the picture content, etc., so as to obtain a new video. In addition, after the video frames in the video are edited to obtain a new video, other applications can be implemented based on the new video, such as video creation, terminal lock screen, etc.
[0191] In the scenario where the new video obtained by video editing is applied to the terminal lock screen, after the video is segmented to obtain the area where the object is located in each video frame in the video, and the video is edited according to the information of the area where the object is located in the video frame to obtain a new video, a dynamic lock screen wallpaper for the terminal can be generated according to the picture content of each video frame in the new video, so that the dynamic lock screen wallpaper can be displayed when the terminal is in the lock screen state.
[0192] 3. Video surveillance scenarios
[0193] In this scenario, after the monitoring device captures a video of a specific area, it can detect objects in the specific area by cutting out the video.
[0194] Next, the video segmentation solution provided in the embodiment of the present application is described in detail through specific examples.
[0195] In one embodiment of the present application, see Fig.15 , a flow chart of a first video cutout method is provided. In this embodiment, the method includes the following steps S301-S305.
[0196] Step S301: compressing information of a target video frame in a video to obtain a first image feature.
[0197] The target video frame may be any video frame among the video frames included in the video.
[0198] Information compression of the target video frame can be understood as: extracting features of the target video frame to obtain a first image feature whose scale is smaller than that of the target video frame. Feature extraction of the target video frame can extract edge information of the content in the image, and the edge information can reflect the area where the object is located in the video frame.
[0199] In addition, when extracting features from the target video frame, multiple cascade feature extractions can be performed. As the number of feature extractions increases, the scale of the resulting features becomes smaller and smaller. From the perspective of scale, the larger the scale of the first image feature, the more detailed edge information it contains. Excessive detailed edge information is not conducive to determining the area where the object is located in the video frame in some cases; conversely, the smaller the scale of the first feature, the more macro edge information it contains, which is more conducive to determining the area where the object is located in the video frame.
[0200] Furthermore, the dimension of the above-mentioned first image feature can be the same as the dimension of the target video frame, that is, the target video frame is a two-dimensional image, so its dimension is 2, then the first image feature can also be two-dimensional data. In this case, the first image feature can also be considered as a feature map.
[0201] In one embodiment of the present application, when information compression is performed on a target video frame, it may be implemented based on an encoding method, for example, based on an encoding network.
[0202] In another embodiment of the present application, the target video frame may be subjected to information compression by performing a convolution transformation on the target video frame. In the process of performing a convolution transformation on the target video frame, the target video frame may be subjected to multiple convolution transformations, thereby continuously reducing the scale of features obtained by the convolution transformation.
[0203] In addition, the target video frame can also be compressed by combining convolution transformation, linear transformation, batch normalization, nonlinear transformation and other processing. For details, please refer to the following Fig. 22 Steps S301A-S301E in the illustrated embodiment are not described in detail here.
[0204] Step S302: Segment the target video frame based on the first image feature, and obtain the features of the contour mask image of the object obtained in the segmentation process and the first contour mask image of the object in the target video frame.
[0205] Among them, the above-mentioned first contour mask image is obtained by segmenting the target video frame. Therefore, there is a corresponding relationship between the pixel points in the first contour mask image and the pixel points in the target video frame. The first contour mask image can be understood as a binary image indicating the position of the area where the object is located in the target video frame. The area where the object is located indicated by the first contour mask image is determined based on the approximate outline of the object. Therefore, the area where the object is located can be considered as the approximate area where the object is located.
[0206] For example, the first contour mask image may be a mask image containing pixel points with two pixel values of 0 and 1. In the target video frame, the pixel points corresponding to the pixel points with a pixel value of 1 in the first contour mask image may be the pixel points in the area where the object is located, and the pixel points corresponding to the pixel points with a pixel value of 0 in the first contour mask image may be the pixel points in the area other than the area where the object is located.
[0207] Specifically, the target video frame may be segmented by any of the following three implementation methods.
[0208] In the first implementation, based on the first image feature, the subsequent Fig.18 The segmentation method mentioned in the illustrated embodiment segments the target video frame, which will not be described in detail here.
[0209] In the second implementation, based on the first image feature, the target video frame can be segmented using a pre-trained network model, as shown in the following Fig.23 The contour feature generation network and the first image generation network in the video segmentation model in the illustrated embodiment are not described in detail here.
[0210] In a third implementation manner, based on the first image feature, the target video frame may be segmented using an image segmentation algorithm, a video segmentation algorithm, or the like.
[0211] Step S303: reconstructing features based on the first image features and the features of the obtained contour mask image, fusing the reconstructed features with the first hidden state information of the video, and updating the first hidden state information to obtain a fusion result.
[0212] The first hidden state information represents: a fusion feature of a transparency mask image of an edge of an object in a video frame that is cut out before a target video frame.
[0213] The edge of an object in a video frame may reflect the image content details of the object in the video frame. The image content details may include the specific location of the region where the object is located and the transparency of the object image content presented by the pixels in the region where the object is located.
[0214] The transparency mask image of the edge of the object in the video frame can be understood as a mask image indicating the image content details of the object in the video frame, and the pixel value of the pixel point in the transparency mask image can represent the transparency of the object image content presented by the pixel point at the same position in the video frame.
[0215] For example, the pixel value range of the pixel in the transparency mask image can be 0-1. If the pixel value of the pixel in the transparency mask image is 0, it means that the pixel at the same position in the video frame belongs to an area outside the area where the object is located, and the image content presented by the pixel at the same position does not include the image content of the object; if the pixel value of the pixel in the transparency mask image is 0.6, it means that the pixel at the same position in the video frame belongs to the area where the object is located, and the transparency of the image content of the object presented by the pixel at the same position is 0.6; if the pixel value of the pixel in the transparency mask image is 1, it means that the pixel at the same position in the video frame belongs to the area where the object is located, and the image content presented by the pixel at the same position is all the image content of the object.
[0216] The video frames to be cut out before the target video frame mentioned in this step include at least two video frames, and of course, may also be all video frames to be cut out before the target video frame.
[0217] When the target video frame is the first video frame in the video, there is no previous video frame that has been cut out. In this case, the first hidden state information may be preset data, for example, preset all-zero data.
[0218] Specifically, the first hidden state information may be represented in a tensor form or in a matrix form.
[0219] From the description of step S301, it can be seen that the first image feature is a feature that is smaller in scale relative to the target video frame, and the first image feature can reflect the area where the object is located in the target video frame. In order to successfully extract the area where the object is located from the target video frame later, it is necessary to perform feature mapping on the small-scale first image feature, and the ultimate goal is to map it to the target video frame, and then obtain the area where the object is located in the target video frame. In view of this, it is necessary to perform upsampling processing on the first image feature.
[0220] Specifically, feature reconstruction is performed based on the first image feature, and then a feature with an increased scale is reconstructed, and then the reconstructed feature and the first hidden state information are fused to obtain a fusion result. Since the first hidden state information represents the fusion feature of the transparency mask image of the edge of the object in the video frame that is cut out before the target video frame, that is, the first hidden state information can represent the information of the edge of the object in the video frame before the target video frame, and the information of the edge of the object can include the specific location information of the area where the object is located, after the reconstructed feature and the first hidden state information are fused, the obtained fusion result can not only reflect the area where the object is located in the target video frame, but also adjust the area where the object is located in the target video frame in combination with the area where the object is located in the previous video frame, thereby ensuring the smoothness of the area where the object is located between adjacent video frames, or the temporal correlation.
[0221] Because the first hidden state information is also needed to be used when the subsequent video frames are cut out, it needs to be updated based on the information of the object in the target video frame. Specifically, the hidden state information can be updated based on the first image feature, or based on the fusion result. For example, the result of fusion of the reconstructed feature and the first hidden state information can be used as the new first hidden state information.
[0222] Specifically, when reconstructing features based on the first image features, an upsampling algorithm can be used to transform the first image features to obtain reconstructed first image features; the first image features can be deconvolved to obtain reconstructed first image features; the first image features can be reconstructed based on a decoding network to obtain reconstructed first image features. For example, the decoding network can be the decoding part in the U-Net network architecture, or the decoder part in the U2-Net network architecture.
[0223] The reconstructed first image feature and the first hidden state information may be fused by any one of the following two implementations.
[0224] In a first implementation manner, a fusion algorithm, a network, etc. may be used to fuse the reconstructed first image feature and the first hidden state information to obtain a fusion result.
[0225] For example, a long short-term memory (LSTM) network, a gated recurrent unit (GRU), etc. are used to fuse the reconstructed first image features and the first hidden state information to obtain a fusion result.
[0226] In the second implementation, the reconstructed first image feature and the first hidden state information may be directly processed by superposition, concatenation or dot multiplication to obtain a processing result as a fusion result.
[0227] For other implementations of the above step S303, please refer to the subsequent embodiments and will not be described in detail here.
[0228] Step S304: Based on the fusion result, a target transparency mask image of the edge of the object in the target video frame is obtained.
[0229] The target transparency mask image may be understood as a mask image indicating the image content details of the object in the target video frame. The pixel value range of the pixel points in the target transparency mask image may be 0-1, and its scale is the same as that of the target video frame.
[0230] In one implementation of the present application, the above fusion result may include the confidence that each pixel in the target video frame belongs to the object. In this case, after obtaining the above fusion result, the target transparency mask image can be obtained according to the confidence corresponding to each pixel included in the fusion result.
[0231] For example, the confidence corresponding to each pixel in the fusion result can be used as the pixel value of each pixel at the same position in the mask image to obtain the target transparency mask image.
[0232] For another example, a first threshold and a second threshold may be preset, wherein the first threshold is greater than the second threshold. If the confidence corresponding to the pixel included in the fusion result is greater than or equal to the first threshold, it means that the confidence corresponding to the pixel is close to 1, and the confidence that the pixel belongs to the object is high. At this time, the pixel value of the pixel at the same position in the mask image can be determined to be 1; if the confidence corresponding to the pixel included in the fusion result is less than or equal to the second threshold, it means that the confidence corresponding to the pixel is close to 0, and the confidence that the pixel belongs to the object is low. At this time, the pixel value of the pixel at the same position in the mask image can be determined to be 0; if the confidence corresponding to the pixel included in the fusion result is between the second threshold and the first threshold, the confidence corresponding to the pixel can be mapped to the 0-1 interval according to the mapping relationship between the confidence interval from the second threshold to the first threshold and the 0-1 interval, and the mapped value is the pixel value of the pixel at the same position in the mask image.
[0233] In another implementation of the present application, the following Fig.17 In the illustrated embodiment, steps S304A-S304C obtain the target transparency mask image.
[0234] Step S305: performing region matting on the target video frame according to the target transparency mask image and the first contour mask image to obtain a matting result.
[0235] Specifically, the first contour mask image indicates the approximate area where the object is located in the target video frame, and the target transparency mask image indicates the image content details of the object in the target video frame. In this way, based on the target transparency mask image and the first contour mask image, the target video frame can be regionally cut out in combination with the approximate area where the object is located in the target video frame and the image content details to obtain a cutout result.
[0236] In one implementation of the present application, the target transparency mask image and the first contour mask image can be dot-multiplied according to the position of each pixel point to obtain a mask image after dot multiplication, and then the obtained mask image and the target video frame are dot-multiplied again according to the position of each pixel point to obtain a dot multiplication result as the cutout result, thereby realizing regional cutout of the target video frame.
[0237] Also, see Fig.16 , shows a schematic diagram of the process from target video frames to target transparency mask images and first contour mask images, and then to the matting results when multiple video frames are used as target video frames. Fig.16 In the figure, the first row of images are multiple target video frames; the second row of images are first contour mask images corresponding to the multiple target video frames; the third row of images are target transparency mask images corresponding to the multiple target video frames; and the fourth row of images are cutout results corresponding to the multiple target video frames.
[0238] As can be seen from the above, in the scheme provided by this embodiment, when the target video frame is segmented based on the first image feature, the feature of the contour mask image of the object in the segmentation process is obtained, and the feature is representative of the contour of the object in the contour mask image. The first hidden state information represents the fusion feature of the transparency mask image of the edge of the object in the video frame that is cut out before the target video frame. In this way, feature reconstruction is performed based on the first image feature and the feature of the obtained contour mask image, and the reconstructed feature and the first hidden state information are fused. The obtained fusion result not only fuses the contour information of the object in the contour mask image, but also fuses the information of the edge of the object in the video frame that is cut out before the target video frame. Since there is often a time domain correlation between video frames in the video, when the target transparency mask image is obtained based on the fusion result, the information of the object in the video frame with time domain correlation is considered on the basis of the target video frame, thereby improving the accuracy of the obtained target transparency mask image. On this basis, according to the target transparency mask image and the first contour mask image, the target video frame can be accurately cut out. It can be seen that the application of the video cutout scheme provided by the embodiment of the present application can improve the accuracy of video cutout.
[0239] In addition, when obtaining the target transparency mask image, the fusion features of the transparency mask images of the object edges in the video frames that were cut out before the target video frame are considered, that is, the image information of the object edges in these video frames is considered, rather than only the image information of the target video frame itself. This can improve the inter-frame smoothness of the changes in the object edge area in the target transparency mask image corresponding to each video frame in the video, thereby improving the inter-frame smoothness of the changes in the object edge area in the cutout results corresponding to each video frame. Moreover, in the case of a video of a moving object, since this solution considers the information of the cutout video frame when cutting out the target video frame, it can make the cutout result of the target video frame and the cutout result of the cutout video frame smoother.
[0240] Other implementations of obtaining the target transparency mask image in the above step S304 are described below.
[0241] In one embodiment of the present application, see Fig.17 , a flow chart of a second video cutout method is provided. In this embodiment, the above step S304 can be implemented by the following steps S304A-S304C.
[0242] Step S304A: Based on the fusion result, a second contour mask image of the object in the target video frame is obtained.
[0243] The second contour mask image may be a binary image, and its scale is the same as that of the target video frame.
[0244] In one implementation of the present application, it can be seen from the description of the above step S304 that the above fusion result may include the confidence that each pixel in the target video frame belongs to the object. In this case, after obtaining the above fusion result, the fusion result can be binarized based on a preset third threshold to obtain a second contour mask image.
[0245] When performing the above-mentioned binarization processing, the numerical value greater than the third threshold in the fusion result can be set to 0, and the numerical value not greater than the third threshold can be set to 1. Of course, the numerical value less than the third threshold in the fusion result can also be set to 0, and the numerical value not less than the third threshold can be set to 1. The embodiment of the present application is not limited to this.
[0246] Step S304B: Fuse the second contour mask image and the fusion result to obtain the target fusion feature.
[0247] Specifically, the second contour mask image and the fusion result may be fused by any one of the following two implementation methods.
[0248] In the first implementation manner, the second contour mask image and the fusion result can be fused using a fusion algorithm, a network, etc. to obtain a fused target fusion feature.
[0249] In the second implementation, the second contour mask image and the fusion result may be directly processed by superposition, concatenation or dot multiplication to obtain a processing result as the target fusion feature.
[0250] Step S304C: Based on the target fusion features, a target transparency mask image of the edge of the object in the target video frame is obtained.
[0251] The target fusion feature is obtained by fusing the second contour mask image and the fusion result. The second contour mask image indicates the approximate area where the object is located in the target video frame. The fusion result may include the confidence that each pixel in the target video frame belongs to the object. In this way, the target fusion feature obtained by fusing the two may include the confidence that each pixel in the approximate area where the object of the target video frame belongs to the object. Based on the target fusion feature, the target transparency mask image can be obtained with the confidence corresponding to each pixel included in the feature.
[0252] Specifically, the implementation method of obtaining the target transparency mask image based on the target fusion feature can refer to the implementation method of obtaining the target transparency mask image based on the fusion result in the above step S304, which will not be repeated here.
[0253] As can be seen from the above, in the solution provided by this embodiment, a second contour mask image of the object is obtained based on the fusion result, and the second contour mask image can indicate the approximate area where the object is located determined by the approximate contour of the object in the target video frame. In this way, the second contour mask image is fused with the fusion result to obtain a target fusion feature. The target fusion feature only needs to focus on the detailed information of the image in the approximate area where the object is located. Based on the target fusion feature, the target transparency mask image of the edge of the object can be accurately obtained while only focusing on the image content in the approximate area where the object is located. Therefore, video cutout is performed based on the target transparency mask image, which can improve the accuracy of video cutout.
[0254] The first implementation method of segmenting the target video frame mentioned in the above step S302 is described below.
[0255] In one embodiment of the present application, see Fig.18 , provides a flow chart of the first feature processing method, Fig.18 The feature processing process shown includes the segmentation process of step S302 and the reconstruction and fusion process of step S303. The number of feature reconstructions included in the segmentation process is the same as the number of transparency information fusions included in the reconstruction and fusion process.
[0256] exist Fig.18 In the segmentation process, the feature reconstruction is performed twice, and the transparency information fusion is performed twice. In addition, the number of feature reconstructions performed in the segmentation process and the number of transparency information fusions performed in the reconstruction fusion process may be other numbers, such as 3, 4, 5, etc., which are not limited in this embodiment.
[0257] The following is combined with the above Fig.18 , respectively, the segmentation process of step S302 and the reconstruction and fusion process of step S303 are described.
[0258] First, for the segmentation process of the above-mentioned step S302, when the target video frame is segmented based on the first image feature, cascade feature reconstruction can be performed based on the first image feature to obtain features of the contour mask image of the object with successively increasing scales, and based on the features obtained by the last processing, the first contour mask image of the object in the target video frame is obtained.
[0259] Cascade feature reconstruction can be understood as multiple feature reconstructions, and the result of each feature reconstruction is the feature of a contour mask image of a scale. In addition, the object of the first feature reconstruction is the above-mentioned first image feature, and the objects of other feature reconstructions are the features obtained by the previous feature reconstruction.
[0260] exist Fig.18 In the cascade feature reconstruction process, the feature reconstruction object is the first image feature. After the first image feature is reconstructed, the contour feature 1 of the contour mask image with increased scale can be obtained. At this time, the first feature reconstruction process ends.
[0261] In the second feature reconstruction process, the object to be reconstructed is the above-mentioned contour feature 1. After feature reconstruction of contour feature 1, contour feature 2 of the contour mask image with the scale increased again can be obtained. At this time, the second feature reconstruction process ends, and contour feature 2 is the feature obtained by the last processing. In this way, based on contour feature 2, the first contour mask image of the object in the target video frame can be obtained.
[0262] The following describes the implementation method of each feature reconstruction and the implementation method of obtaining the first contour mask image based on the features obtained in the last processing.
[0263] In each feature reconstruction process, feature reconstruction can be achieved through either of the following two implementation methods.
[0264] In the first implementation manner, feature reconstruction can be achieved by the feature reconstruction, feature fusion and update method mentioned in the subsequent embodiments, which will not be described in detail here.
[0265] In the second implementation, an upsampling algorithm, a deconvolution transform, or a feature decoding network may be used to process the object of each feature reconstruction.
[0266] After obtaining the feature obtained by the last feature reconstruction process, the feature may include the confidence that each pixel in the target video frame belongs to the object. In this case, after obtaining the feature, the feature can be binarized based on the preset fourth threshold to obtain the first contour mask image.
[0267] Secondly, for the reconstruction fusion process of step S303, in this embodiment, the first hidden state information includes: a plurality of first sub-hidden state information, each of which represents the fusion characteristics of a transparency mask image of a certain scale. The plurality of first sub-hidden state information may represent the fusion characteristics of transparency mask images of increasing scales. For example, the first hidden state information may include three first sub-hidden state information, and the three first sub-hidden state information may represent the fusion characteristics of transparency mask images of scales of 24*24, 28*28, and 32*32, respectively.
[0268] In this case, when reconstructing, fusing and updating features based on the first image features and the features of the obtained contour mask image, a preset number of transparency information fusions may be performed in the following manner, and the features obtained by the last processing are determined as the fusion results:
[0269] Based on the first target feature and the second target feature in the feature of the obtained contour mask image, feature reconstruction is performed to obtain a scale-enlarged second image feature, the second image feature and the first sub-state information in the first hidden state information are fused and the first sub-state information is updated to obtain a third image feature.
[0270] The preset number is a pre-set number and is the same as the number of reconstructions of the cascade feature reconstruction.
[0271] When information fusion is performed for the first time, the first target feature is the first image feature. When information fusion is performed for other times, the first target feature is the feature obtained by the last information fusion.
[0272] The scale of the first target feature is the same as the scale of the second target feature.
[0273] The scale of the transparency mask image corresponding to the fusion feature represented by the first sub-state information is the same as the scale of the second image feature.
[0274] exist Fig.18In the above-mentioned preset number is 2, that is, two transparency information fusion processes are performed. In the first transparency information fusion process, the first target feature is the first image feature, the second target feature is the above-mentioned contour feature 1, the first image feature and the contour feature 1 have the same scale, and feature reconstruction is performed based on the first target feature and the second target feature, that is, feature reconstruction is performed based on the first image feature and the contour feature 1 to obtain a scale-enlarged second image feature 1, the second image feature 1 and the first sub-state information 1 corresponding to the second image feature 1 are fused and the first sub-state information 1 is updated to obtain the third image feature 1, and at this time, the first transparency information fusion process ends.
[0275] After obtaining the third image feature 1 in the first transparency information fusion process, the second transparency information fusion is performed. In the second transparency information fusion process, the first target feature is the third image feature 1, the second target feature is the contour feature 2, the third image feature 1 and the contour feature 2 have the same scale, feature reconstruction is performed based on the first target feature and the second target feature, that is, feature reconstruction is performed based on the third image feature 1 and the contour feature 2 to obtain a second image feature 2 with an increased scale, the second image feature 2 and the first sub-state information 2 corresponding to the second image feature 2 are fused and the first sub-state information 2 is updated to obtain the third image feature 2, at which point the second transparency information fusion process ends, and the third image feature 2 is the final fusion result.
[0276] The implementation method of each feature reconstruction based on the first target feature and the second target feature can be referred to in the above Fig.15 In the embodiment shown, step S303 is implemented based on the first image features and the features of the contour mask image to perform feature reconstruction; each time the second image features and the first sub-state information are merged and the first sub-state information is updated, see the aforementioned Fig.15 In the illustrated embodiment, the implementation manner of step S303 of fusing the reconstructed features with the first hidden state information and updating the first hidden state information will not be described in detail here.
[0277] As can be seen from the above, in the solution provided by this embodiment, based on the first image feature, the target video frame is segmented in a cascaded feature reconstruction manner, and the cascaded feature reconstruction includes multiple feature reconstruction processes, which can improve the accuracy of the first contour mask image; after obtaining the first image feature, multiple transparency information fusions are performed, and each transparency information fusion process includes three processing processes: feature reconstruction, feature and hidden state information fusion, and updating of hidden state information. This can improve the accuracy of the final fusion result, so that based on the relatively accurate first contour mask image and the fusion result, the target video frame can be regionally cut out, which can improve the accuracy of regional cut out, and then improve the accuracy of video cut out.
[0278] The amount of data of the two features required for feature reconstruction in each transparency information fusion process is usually large, so the amount of calculation for feature reconstruction is also large, and the amount of calculation for performing a preset number of transparency information fusions is also large.
[0279] In view of this, in one embodiment of the present application, when performing feature reconstruction based on the first target feature and the second target feature in the features of the obtained contour mask image, representative features of the edge of the object can be screened from the second target feature in the features of the obtained contour mask image, and feature reconstruction can be performed based on the first target feature and the representative feature to obtain a second image feature with an enlarged scale.
[0280] The representative feature is a feature included in the second target feature that is representative of the edge details of the object in the target video frame.
[0281] In the process of feature reconstruction based on the first image feature, the attributes of the object represented by the features of the reconstructed contour mask image can be determined. In this way, in each transparency information fusion process, after determining the second target feature with the same scale as the first target feature in the features of each contour mask image, the features representative of the edge of the object in the second target feature can be determined according to the attributes of the object represented by the second target feature, thereby extracting the determined features from the second target feature, and the extracted features are the representative features.
[0282] Specifically, after determining that the second target feature has representative features for the edge of the object, the representative features can be extracted from the second target feature using a feature screening algorithm, a feature screening network, or an attention mechanism.
[0283] After the representative features are screened out, the implementation method of feature reconstruction based on the first target feature and the representative features can refer to the implementation method of feature reconstruction based on the first target feature and the second target feature in the aforementioned embodiment, which will not be repeated here.
[0284] See also Fig.19 , Fig.19 exist Fig.18 The specific adjustment content is: in the first transparency information fusion process, when reconstructing features based on the first image feature and the contour feature 1, firstly, the representative feature 1 is screened out from the contour feature 1, and then the feature reconstruction is performed based on the first image feature and the representative feature 1 to obtain the second image feature 1 with an increased scale. In the second transparency information fusion process, when reconstructing features based on the third image feature 1 and the contour feature 2, firstly, the representative feature 2 is screened out from the contour feature 2, and then the feature reconstruction is performed based on the third image feature 1 and the representative feature 2 to obtain the second image feature 2 with an increased scale.
[0285] From the above, it can be seen that in the solution provided by this embodiment, in the process of feature reconstruction based on the first target feature and the second target feature, representative features with smaller data volume are screened out from the second target feature, so that feature reconstruction is performed based on the first target feature and the representative feature. This can reduce the amount of calculation for feature reconstruction, thereby reducing the amount of calculation for performing a preset number of transparency information fusions, improving the efficiency of obtaining fusion results, and further improving the efficiency of video cutouts.
[0286] The following is the above Fig.18 The method of realizing feature reconstruction by feature reconstruction, feature fusion and updating mentioned in the illustrated embodiment is explained.
[0287] In one embodiment of the present application, after obtaining the above-mentioned first image feature, a preset number of contour information fusions may be performed in the following manner to obtain features of contour mask images of objects with successively increasing scales, wherein one contour information fusion process is equivalent to one feature reconstruction process:
[0288] Feature reconstruction is performed based on the third target feature to obtain a fourth image feature with an increased scale, the fourth image feature is fused with the second sub-state information in the second hidden state information and the second sub-state information is updated to obtain the feature of the contour mask image of the object.
[0289] Among them, when the information fusion is performed for the first time, the third target feature is the first image feature, and when the information fusion is performed for other times, the third target feature is the feature obtained by the previous feature reconstruction.
[0290] The second hidden state information representation is: the fused features of the contour mask image of the object in the video frame segmented before the target video frame.
[0291] The second hidden state information includes: a plurality of second sub-hidden state information, each second sub-hidden state information represents a fusion feature of a contour mask image of a certain scale. The plurality of second sub-hidden state information may represent fusion features of contour mask images of successively increasing scales. For example, the second hidden state information may include three second sub-hidden state information, and the three second sub-hidden state information may represent fusion features of contour mask images of successively increasing scales of 24*24, 28*28, and 32*32.
[0292] The scale of the contour mask image corresponding to the fusion feature represented by the second sub-state information is the same as the scale of the fourth image feature. For example, if the second hidden state information includes three second sub-hidden state information, and the three second sub-hidden state information represent the fusion features of the contour mask images with scales of 24*24, 28*28, and 32*32, respectively, it means that the preset number is three, and three fusions are required. The scales of the fourth image features obtained in the three fusion processes are 24*24, 28*28, and 32*32, respectively.
[0293] Below, the above preset number is 3, combined with Fig. 20 , the above contour information fusion process is explained in detail.
[0294] See also Fig. 20 , a flowchart of a cascade feature reconstruction method is provided. Fig. 20 In the process of the first contour information fusion, after obtaining the first image feature, the first contour information fusion is performed. In the process of the first contour information fusion, the third target feature is the first image feature, and feature reconstruction is performed based on the third target feature, that is, feature reconstruction is performed based on the first image feature to obtain a fourth image feature 1 with an increased scale, and the fourth image feature 1 and the second sub-state information 1 corresponding to the fourth image feature 1 are fused and the second sub-state information 1 is updated to obtain a contour feature 3 of the contour mask image of the object. At this time, the first contour information fusion process ends.
[0295] After the contour feature 3 is obtained, the second contour information fusion is performed. In the second contour information fusion process, the third target feature is the contour feature 3, and feature reconstruction is performed based on the third target feature, that is, feature reconstruction is performed based on the contour feature 3 to obtain the fourth image feature 2 with a scale increased again, and the fourth image feature 2 and the second sub-state information 2 corresponding to the fourth image feature 2 are fused and the second sub-state information 2 is updated to obtain the contour feature 4 of the contour mask image of the object. At this time, the second contour information fusion process ends.
[0296] After the contour feature 4 is obtained, the third contour information fusion is performed. In the third contour information fusion process, the third target feature is the contour feature 4, and feature reconstruction is performed based on the third target feature, that is, feature reconstruction is performed based on the contour feature 4 to obtain the fourth image feature 3 with a further increased scale, and the fourth image feature 3 and the second sub-state information 3 corresponding to the fourth image feature 3 are fused and the second sub-state information 3 is updated to obtain the contour feature 5 of the contour mask image of the object. At this time, the second contour information fusion process ends, and the obtained contour feature 5 is the feature of the contour mask image of the object finally obtained.
[0297] The implementation method of feature reconstruction based on the third target feature in each contour information fusion process can be found in the above Fig.15 In the embodiment shown, the implementation method of feature reconstruction based on the first image feature in step S303; the implementation method of fusing and updating the second sub-state information based on the fourth image feature and the second sub-state information can be referred to in the aforementioned Fig.15 The implementation manner of fusing and updating the first hidden state information based on the reconstructed first image feature and the first hidden state information in step S303 in the illustrated embodiment will not be described in detail here.
[0298] As can be seen from the above, in the scheme provided by this embodiment, the second hidden state information represents the fusion features of the contour mask image of the object in the video frame segmented before the target video frame, and each second sub-hidden state information represents the fusion features of the contour mask image of a scale. Therefore, in each contour information fusion process, the fourth image feature obtained by feature reconstruction is fused with the second sub-state information, that is, the information in the fusion features of the contour mask image of a scale of the object in the video frame segmented before the target video frame is fused with the fourth image feature. In this way, the features of the contour mask image of the object finally obtained not only include the information of the object in the target video frame, but also include the information of the object in the video frame segmented before the target video frame, so that the first contour mask image is obtained based on the finally obtained features, which can improve the accuracy of the first contour mask image, and then improve the accuracy of video cutout.
[0299] When reconstructing features during the contour information fusion process, the processing may be based on the third target feature for feature reconstruction, or may be based on the third target feature and other information for feature reconstruction.
[0300] In one embodiment of the present application, the first image feature includes a plurality of first sub-image features. When information compression is performed on a target video frame in a video, cascade information compression may be performed on the target video frame to obtain first sub-image features with successively reduced scales.
[0301] Cascade information compression can be understood as multiple information compressions, each of which results in a first sub-image feature. Furthermore, the object of the first information compression is the target video frame, and the objects of other information compressions are the first sub-image features obtained from the previous information compression.
[0302] The implementation method of each information compression can be found in the above Fig.15 The implementation method of information compression on the target video frame in step S301 is shown.
[0303] For example, each time information is compressed, multiple convolution transformations may be performed on the object to be compressed.
[0304] For example, each time information is compressed, one or more subsequent Fig. 22 The processing flow shown in steps S301A-S301E in the illustrated embodiment implements information compression.
[0305] After obtaining the first sub-image features with successively decreasing scales, contour information fusion may be performed a preset number of times based on the first sub-image features.
[0306] When performing feature reconstruction in the first contour information fusion process, the first sub-image feature with the smallest scale among the first sub-image features may be used as the target feature, and feature reconstruction is performed based on the target feature.
[0307] When performing feature reconstruction in other contour information fusion processes, the feature obtained in the previous contour information fusion can be used as the third target feature, and feature reconstruction is performed based on the third target feature and the first sub-image feature with the same scale as the third target feature to obtain a fourth image feature with an increased scale.
[0308] When reconstructing features based on the third target feature and the first sub-image feature having the same scale as the third target feature, the third target feature and the first sub-image feature having the same scale as the third target feature can be fused into one feature through fusion methods such as superposition and dot product, and then feature reconstruction is performed based on the fused feature.
[0309] The following takes the above preset number of 3 as an example, combined with Fig.21 , the process of the above contour information fusion is explained in detail.
[0310] See also Fig.21 , a flowchart of another cascade feature reconstruction method is provided. Fig.21In the process of first contour information fusion, after obtaining the first sub-image features with successively decreasing scales, the first contour information fusion is performed. In the process of first contour information fusion, the third target feature is the first sub-image feature 1 with the smallest scale among the first sub-image features, and feature reconstruction is performed based on the third target feature, that is, feature reconstruction is performed based on the first sub-image feature 1 to obtain the fourth image feature 4 with an increased scale, and the fourth image feature 4 and the second sub-state information 4 corresponding to the fourth image feature 4 are fused and the second sub-state information 4 is updated to obtain the contour feature 6 of the contour mask image of the object. At this time, the first contour information fusion process ends.
[0311] After obtaining the contour feature 6 in the first contour information fusion process, the second contour information fusion is performed. In the second contour information fusion process, the third target feature is the above-mentioned contour feature 6, and the first sub-image feature with the same scale as the third target feature is the first sub-image feature 2. Feature reconstruction is performed based on the third target feature and the first sub-image feature with the same scale as the third target feature, that is, feature reconstruction is performed based on the contour feature 6 and the first sub-image feature 2 to obtain a fourth image feature 5 with a further increased scale. The fourth image feature 5 and the second sub-state information 5 corresponding to the fourth image feature 5 are fused and the second sub-state information 5 is updated to obtain the contour feature 7 of the contour mask image of the object. At this time, the second contour information fusion process ends.
[0312] After obtaining the contour feature 7 in the second contour information fusion process, the third contour information fusion is performed. In the third contour information fusion process, the third target feature is the above-mentioned contour feature 7, the first sub-image feature with the same scale as the third target feature is the first sub-image feature 3, and feature reconstruction is performed based on the third target feature and the first sub-image feature with the same scale as the third target feature, that is, feature reconstruction is performed based on the contour feature 7 and the first sub-image feature 3 to obtain a fourth image feature 6 with a further increased scale, and the fourth image feature 6 and the second sub-state information 6 corresponding to the fourth image feature 6 are fused and the second sub-state information 6 is updated to obtain the contour feature 8 of the contour mask image of the object. At this time, the third contour information fusion process is completed, and the contour feature 8 obtained in the process is the feature of the contour mask image of the object finally obtained.
[0313] As can be seen from the above, in the solution provided by this embodiment, the target video frame is cascadedly compressed to obtain first sub-image features with successively decreasing scales. In the subsequent contour information fusion processes except the first one, feature reconstruction can be performed based on the third target feature and the first sub-image feature with the same scale as the third target feature. This can improve the accuracy of feature reconstruction, thereby improving the accuracy of the features finally obtained after contour information fusion, and further improving the accuracy of video cutout.
[0314] According to the above content, the number of transparency information fusions and the number of contour information fusions are both preset numbers. The larger the preset number, the more accurate the fusion result obtained by the transparency information fusion and the more accurate the features obtained by the contour information fusion. However, the amount of calculation is also greater.
[0315] In view of this, in one embodiment of the present application, the above preset number is 4, 5 or 6. This can not only improve the accuracy of video cutout, but also avoid excessive calculation, thereby ensuring that video cutout is achieved with higher efficiency, while also saving computing resources of the terminal. Therefore, the solution provided in the embodiment of the present application can be applicable to the terminal, and is friendly to the application of the solution on the terminal, thereby realizing a lightweight application of the video cutout solution in the terminal.
[0316] Since the data volume of the second image feature and the fourth image feature itself is relatively large, the computational complexity of fusing the second image feature with the first sub-state information in the first hidden state information and updating the first sub-state information is usually large, and the computational complexity of fusing the fourth image feature with the second sub-state information in the second hidden state information and updating the second sub-state information is usually also large.
[0317] In view of the above situation, in order to reduce the amount of calculation for fusing the second image feature and the first sub-state information and updating the first sub-state information, in one embodiment of the present application, when fusing the second image feature and the first sub-state information and updating the first sub-state information, the second image feature is segmented to obtain the second sub-image feature and the third sub-image feature; the second sub-image feature and the first sub-state information in the first hidden state information are fused and the first sub-state information is updated to obtain the fourth sub-image feature; and the fourth sub-image feature and the third sub-image feature are spliced to obtain the third image feature.
[0318] The feature can be represented in the form of a matrix or a tensor. Taking a tensor as an example, segmenting the second image feature can be understood as segmenting the feature tensor representing the second image feature into two sub-tensors in any dimensional direction of the feature tensor.
[0319] For example, for a feature tensor with a scale of H*C*W, the feature tensor can be split in the W dimension direction to obtain two sub-feature tensors with scales of H*C*W1 and H*C*W2, where W1+W2=W.
[0320] Specifically, when segmenting the second image feature, the second image feature may be segmented in equal proportion to obtain two sub-features of the same scale, or the second image feature may be segmented in any proportion to obtain two sub-features of different scales. Furthermore, after segmenting the second image feature to obtain two sub-features, any one of the sub-features may be determined as the second sub-image feature, and the other sub-feature may be determined as the third sub-image feature.
[0321] After the second sub-image feature and the third sub-image feature are obtained by segmentation, the second sub-image feature and the first sub-state information in the first hidden state information can be merged and the first sub-state information can be updated to obtain the fourth sub-image feature. The specific implementation method can be referred to above. Fig.15 In the illustrated embodiment, the implementation method of step S303 of fusing the reconstructed features with the first hidden state information and updating the first hidden state information will not be described in detail here.
[0322] The splicing feature can be regarded as the reverse processing of feature segmentation. After the fourth sub-image feature is obtained by fusion, the fourth sub-image feature and the third sub-image feature can be spliced. When splicing these two features, the fourth sub-image feature and the third sub-image feature can be spliced into one feature in the dimensional direction based on the segmentation, that is, in the dimensional direction based on the segmentation, the third sub-image feature is spliced behind the fourth sub-image feature, or the fourth sub-image feature is spliced behind the third sub-image feature. The feature obtained by splicing is the third image feature.
[0323] For example, if the scale of the third sub-image feature is H*C*W3 and the scale of the fourth sub-image feature is H*C*W4, then when splicing the fourth sub-image feature and the third sub-image feature, the fourth sub-image feature and the third sub-image feature can be spliced into a feature with a scale of H*C*W5 in the W dimension direction, where W3+W4=W5.
[0324] As can be seen from the above, in the solution provided by this embodiment, the second image feature is segmented to obtain a second sub-image feature and a third sub-image feature. The data volume of the second sub-image feature and the third sub-image feature are both smaller than the data volume of the second image feature. In this way, the second sub-image feature and the first sub-state information are fused, which can reduce the amount of fusion calculations and improve the fusion efficiency, thereby improving the efficiency of obtaining the third image feature, and further improving the efficiency of video cutout. At the same time, it also saves the computing resources of the terminal, thereby realizing a lightweight application of video cutout solutions in the terminal.
[0325] In order to reduce the amount of calculation for fusing the fourth image feature and the second sub-state information and updating the second sub-state information, in one embodiment of the present application, when fusing the fourth image feature and the second sub-state information and updating the second sub-state information, the fourth image feature is segmented to obtain a fifth sub-image feature and a sixth sub-image feature; the fifth sub-image feature and the second sub-state information in the second hidden state information are fused and the second sub-state information is updated to obtain a seventh sub-image feature; the seventh sub-image feature and the sixth sub-image feature are spliced to obtain the features of the contour mask image of the object.
[0326] For the implementation method of segmenting the fourth image feature, reference may be made to the aforementioned implementation method of segmenting the second image; for the implementation method of fusing the fifth sub-image feature with the second sub-state information and updating the second sub-state information, reference may be made to the aforementioned implementation method of fusing the second sub-image feature with the first sub-state information and updating the first sub-state information; for the implementation method of splicing the seventh sub-image feature with the sixth sub-image feature, reference may be made to the aforementioned implementation method of splicing the fourth sub-image feature with the third sub-image feature, which will not be repeated here.
[0327] As can be seen from the above, in the solution provided by this embodiment, the fourth image feature is segmented to obtain the fifth sub-image feature and the sixth sub-image feature. The data volume of the fifth sub-image feature and the sixth sub-image feature are both smaller than the data volume of the fourth image feature. In this way, the fifth sub-image feature and the first sub-state information are fused, which can reduce the amount of fusion calculation and improve the fusion efficiency, thereby improving the efficiency of obtaining the features of the contour mask image of the object, and further improving the efficiency of video cutout. At the same time, it also saves the computing resources of the terminal, thereby realizing a lightweight application of video cutout solution in the terminal.
[0328] In addition, in order to solve the problem of large amount of calculation in the above two fusion processes, in one embodiment of the present application, when the second image feature and the first sub-state information are fused and the first sub-state information is updated, the second image feature can be segmented to obtain the second sub-image feature and the third sub-image feature; the second sub-image feature and the first sub-state information in the first hidden state information are fused and the first sub-state information is updated to obtain the fourth sub-image feature; the fourth sub-image feature and the third sub-image feature are spliced to obtain the third image feature. And, when the fourth image feature and the second sub-state information are fused and the second sub-state information is updated, the fourth image feature is segmented to obtain the fifth sub-image feature and the sixth sub-image feature; the fifth sub-image feature and the second sub-state information in the second hidden state information are fused and the second sub-state information is updated to obtain the seventh sub-image feature; the seventh sub-image feature and the sixth sub-image feature are spliced to obtain the feature of the contour mask image of the object. In this way, the amount of calculation of the two fusion processes can be reduced, thereby further improving the fusion efficiency, and further improving the efficiency of video cutout, and realizing a lightweight application video cutout solution in the terminal.
[0329] The following describes the implementation method of information compression of the target video frame in the above step S301 by combining convolution transformation, linear transformation, batch normalization processing, nonlinear transformation and other processing.
[0330] In one embodiment of the present application, see Fig. 22 , a flow chart of a third video cutout method is provided. In this embodiment, the above step S301 can be implemented by the following steps S301A-S301E.
[0331] Step S301A: performing a convolution transformation on a target video frame in the video to obtain a fifth image feature.
[0332] Specifically, a pre-set convolution kernel can be used to perform convolution calculation on the target video frame to obtain the fifth image feature, or a trained convolutional neural network can be used to perform convolution transformation on the target video frame to obtain the fifth image feature output by the model.
[0333] Step S301B: performing a linear transformation on the fifth image feature based on the convolution kernel to obtain a sixth image feature.
[0334] Among them, the above convolution kernel is a pre-set convolution kernel.
[0335] Specifically, based on the convolution kernel, the fifth image feature can be linearly transformed by performing a convolution transformation on the fifth image feature. Since the network processor (NPU) in the terminal has a strong computing power for convolution transformation, performing linear transformation by convolution transformation can shorten the time consumption of linear transformation, thereby shortening the time consumption of video cutout and improving the efficiency of video cutout.
[0336] In one embodiment of the present application, the convolution kernel is: a 1x1 convolution kernel. Since the 1x1 convolution kernel itself has a small amount of data, the fifth image feature is linearly transformed based on the 1x1 convolution kernel. On the premise that the fifth image feature can be linearly transformed, the amount of linear transformation calculations can be reduced, and the linear transformation calculation efficiency can be improved, thereby improving the efficiency of video cutout. In addition, the video cutout solution provided in this embodiment is applied to a terminal. In the terminal, the fifth image feature is linearly transformed based on the 1x1 convolution kernel, without occupying more computing resources of the terminal, thereby facilitating the terminal to achieve linear transformation and promoting lightweight video cutout on the terminal side.
[0337] Step S301C: performing batch normalization processing on the sixth image feature to obtain the seventh image feature.
[0338] Specifically, a batch normalization algorithm, model, etc. may be used to perform batch normalization processing on the sixth image feature to obtain the seventh image feature.
[0339] For example, the BatchNorm2d algorithm may be used to perform batch normalization processing on the sixth image feature.
[0340] Step S301D: Perform nonlinear transformation on the seventh image feature to obtain the eighth image feature.
[0341] Specifically, a nonlinear transformation function, algorithm or activation function may be used to perform a nonlinear transformation on the seventh image feature to obtain the eighth image feature.
[0342] For example, the GELU activation function or the RELU activation function can be used to perform a nonlinear transformation on the seventh image feature. When the RELU activation function is used to perform a nonlinear transformation on the seventh image feature, since the quantization effect of processing data using the RELU activation function is better, the RELU activation function is used to perform a nonlinear transformation on the seventh image feature, which can improve the transformation effect of the nonlinear transformation, thereby improving the accuracy of the eighth image feature.
[0343] Step S301E: performing a linear transformation on the eighth image feature based on the convolution kernel to obtain a first feature of the target video frame.
[0344] The implementation method of the linear transformation in this step is the same as the implementation method of the linear transformation in the above step S301B, which will not be repeated here.
[0345] In addition, when obtaining the first feature, the processing flow shown in steps S301A-S301E may be performed once, or multiple times. For example, the processing flow shown in steps S301A-S301E may be performed 4 times, 5 times, or other numbers of times.
[0346] When executing the processing flow shown in steps S301A-S301E multiple times, the input of the first processing flow is the target video frame in the video, and the input of other processing flows is the features output by the previous processing flow. The feature output by the last processing flow is the above-mentioned first feature. In this case, as the above-mentioned processing flow is executed multiple times, the scale of the features output by each processing flow continues to decrease.
[0347] From the above, it can be seen that in the solution provided by this embodiment, when compressing the information of the target video frame, the target video frame is subjected to multiple processing such as convolution transformation, linear transformation, batch normalization processing, nonlinear transformation, etc., so that more accurate information compression of the target video can be achieved, thereby improving the accuracy of the first image feature, and then performing video segmentation based on the first image feature, which can improve the accuracy of video cutout.
[0348] In addition, in the solution provided in the embodiment of the present application, the sixth image feature is first batch standardized and then the seventh image feature obtained through processing is nonlinearly transformed. This can prevent the loss of feature quantization accuracy during information compression, thereby improving the quantization accuracy of information compression and further improving the accuracy of the first image feature and the accuracy of video cutout.
[0349] The solution provided in the embodiment of the present application is applied to the terminal, and processes such as convolution transformation, linear transformation, batch normalization processing, and nonlinear transformation are relatively friendly to the terminal's computing power. Therefore, performing convolution transformation, linear transformation, batch normalization processing, and nonlinear transformation in the terminal can facilitate the terminal to compress information, thereby promoting lightweight video cutout implementation on the terminal side.
[0350] The video cutout solution provided in the embodiment of the present application can also be implemented based on a neural network model. The video cutout solution is described below in conjunction with the neural network model.
[0351] In one embodiment of the present application, the above steps can be implemented using a pre-trained video cutout model.
[0352] See also Fig.23, provides a structural diagram of the first video cutout model, from Fig.23 It can be seen that the video matting model includes an information compression network, a first image generation network, a second image generation network, a result output network, three groups of contour feature generation networks and three groups of transparency feature generation networks. Each group of contour feature generation networks corresponds to the scale of a contour mask image, including a first reconstruction subnetwork and a first fusion subnetwork. Each group of transparency feature generation networks corresponds to the scale of a transparency mask image, including a second reconstruction subnetwork and a second fusion subnetwork.
[0353] Fig.23 The video cutout model is illustrated by taking the number of contour feature generation networks included as three as an example. In addition, the number of contour feature generation networks included in the video cutout model can also be four, five or other numbers. The number of transparency feature generation networks is the same as the number of contour feature generation networks, and this embodiment does not limit this.
[0354] Below Fig.23 The connection relationship between each network in the video cutout model shown is explained.
[0355] The three groups of contour feature generation networks in the video cutout model are contour feature generation network 1, contour feature generation network 2 and contour feature generation network 3, whose scales of the corresponding contour mask images increase successively. The first reconstruction subnetwork contained in each group of contour feature generation networks is connected to the first fusion subnetwork; the three groups of transparency feature generation networks are transparency feature generation network 1, transparency feature generation network 2 and transparency feature generation network 3, whose scales of the corresponding transparency mask images increase successively. The second reconstruction subnetwork contained in each group of transparency feature generation networks is connected to the second fusion subnetwork.
[0356] There are two branches in the video cutout model: the segmentation branch and the cutout branch. The segmentation branch mainly includes three groups of contour feature generation networks and the first image generation network, and the cutout branch mainly includes three groups of transparency feature generation networks and the second image generation network. The first layer of the video cutout model is the information compression network, which is connected to the two branches respectively, and there is a connection relationship between the two branches.
[0357] The following describes the connection relationship between the networks contained in the segmentation branch and the cutout branch, as well as the connection relationship between the two branches.
[0358] First, the connection relationship of the network included in the segmentation branch itself is explained. The information compression network is connected to the first reconstruction subnetwork 1 included in the contour feature generation network 1, the first fusion subnetwork 1 included in the contour feature generation network 1 is connected to the first reconstruction subnetwork 2 included in the contour feature generation network 2, the first fusion subnetwork 2 included in the contour feature generation network 2 is connected to the first reconstruction subnetwork 3 included in the contour feature generation network 3, and the first fusion subnetwork 3 included in the contour feature generation network 3 is connected to the first image generation network.
[0359] Secondly, the connection relationship of the network included in the cutout branch itself is explained. The information compression network is connected to the second reconstruction subnetwork 1 included in the transparency feature generation network 1, the second fusion subnetwork 1 included in the transparency feature generation network 1 is connected to the second reconstruction subnetwork 2 included in the transparency feature generation network 2, the second fusion subnetwork 2 included in the transparency feature generation network 2 is connected to the second reconstruction subnetwork 3 included in the transparency feature generation network 3, and the second fusion subnetwork 3 included in the transparency feature generation network 3 is connected to the second image generation network.
[0360] Finally, the connection relationship between the segmentation branch and the cutout branch is described. The first fusion subnetwork 1 included in the contour feature generation network 1 is connected to the second reconstruction subnetwork 1 included in the transparency feature generation network 1; the first fusion subnetwork 2 included in the contour feature generation network 2 is connected to the second reconstruction subnetwork 2 included in the transparency feature generation network 2; the first fusion subnetwork 3 included in the contour feature generation network 3 is connected to the second reconstruction subnetwork 3 included in the transparency feature generation network 3. In addition, the first image generation network and the second image generation network are respectively connected to the result output network.
[0361] The following describes each network and sub-network in the video cutout model.
[0362] For the information compression network, when information compression is performed on a target video frame, the target video frame is input into the information compression network, and the information compression network performs information compression on the target video frame, thereby obtaining a first feature output by the information compression network.
[0363] The implementation method of the information compression network to compress the target video frame can be found in the above content and will not be repeated here.
[0364] See also Fig.24 , Fig.24 It is a structural diagram of an information compression network. Fig.24 In the information compression network shown, the network layers from top to bottom are: convolutional layer, linear layer 1, batch normalization layer, nonlinear layer and linear layer 2.
[0365] The convolution layer is used to perform a convolution transformation on the target video frame to obtain the fifth image feature.
[0366] The linear layer 1 is used to perform a linear transformation on the fifth image feature based on the convolution kernel to obtain the sixth image feature.
[0367] The batch normalization layer is used to perform batch normalization processing on the sixth image feature to obtain the seventh image feature.
[0368] The nonlinear layer is used to perform a nonlinear transformation on the seventh image feature to obtain the eighth image feature.
[0369] The linear layer 2 is used to perform a linear transformation on the eighth image feature based on the convolution kernel to obtain the first feature.
[0370] The implementation method of data processing by the convolution layer, linear layer 1, batch normalization layer, nonlinear layer and linear layer 2 can be found in the above content and will not be repeated here.
[0371] For the target first reconstruction subnetwork in the target contour feature generation network, when reconstructing features based on the third target feature, the third target feature is input into the target first reconstruction subnetwork, and the target first reconstruction subnetwork performs feature reconstruction based on the third target feature to obtain a fourth image feature with an increased scale output by the target first reconstruction subnetwork. The scale of the contour mask image corresponding to the target contour feature generation network is the same as the scale of the fourth image feature.
[0372] The implementation method of the target first reconstruction sub-network performing feature reconstruction based on the third target feature can be found in the above-mentioned embodiment, which will not be described in detail here.
[0373] In one embodiment of the present application, the first reconstruction sub-network is implemented based on the QARepVGG network structure.
[0374] Since the quantization calculation accuracy of the QARepVGG network is relatively high, implementing the above-mentioned first reconstruction subnetwork based on the QARepVGG network structure can improve the quantization calculation capability of the first reconstruction subnetwork, thereby improving the accuracy of feature reconstruction by the first reconstruction subnetwork, and further improving the accuracy of video cutout.
[0375] In another embodiment of the present application, the first reconstruction subnetwork in the specific contour feature generation network is implemented based on the QARepVGG network structure.
[0376] The specific contour feature generation network is a contour feature generation network whose corresponding contour mask image has a scale smaller than a first preset scale.
[0377] The first preset scale may be a preset scale.
[0378] When constructing the above-mentioned video cutout model, the scale of the contour mask image corresponding to each group of contour feature generation networks can be determined, so that the contour feature generation network whose corresponding contour mask image scale is smaller than the first preset scale can be determined as a specific contour feature generation network, so that when constructing a specific contour feature generation network, the specific contour feature generation network is constructed based on the QARepVGG network structure. For other contour feature generation networks, they can be constructed based on other network structures.
[0379] Since the computational amount of the U-shaped residual block in the specific contour feature generation network constructed based on the QARepVGG network structure increases with the increase of the scale of the contour mask image corresponding to the network, when constructing each contour feature generation network, only the specific contour feature generation network whose corresponding contour mask image has a scale smaller than the first preset scale can be used to implement the first reconstruction subnetwork in the specific contour feature generation network based on the QARepVGG network structure. This can reduce the computational amount of each contour feature generation network, improve the efficiency of obtaining the features of the contour mask image of the object, thereby improving the efficiency of video cutout, and can also lightweight deploy the above-mentioned video cutout model in the terminal.
[0380] For the target first fusion sub-network in the target contour feature generation network, when the fourth image feature and the second sub-state information are fused and the second sub-state information is updated, the fourth image feature can be input into the target first fusion sub-network, and the second sub-hidden state information provided by the target first fusion sub-network is the second sub-state information. In this way, the target first fusion sub-network fuses the fourth image feature and the second sub-hidden state information provided by itself and updates the second sub-hidden state information provided by itself, thereby obtaining the features of the contour mask image of the object output by the target first fusion sub-network.
[0381] The implementation manner in which the target first fusion sub-network fuses the fourth image feature and the second sub-hidden state information and updates the second sub-hidden state information can be referred to the aforementioned embodiment, which will not be repeated here.
[0382] In one embodiment of the present application, the first fusion sub-network is a gated recurrent unit (Gated Recurrent Unit, GRU) or a long short-term memory (Long Short-Term Memory, LSTM) unit.
[0383] Both GRU and LSTM units have information memory functions. Using either of these two units as the first fusion sub-network, the unit itself can store hidden state information of the fusion features of the contour mask image representing the object in the segmented video frame, so that the fourth image feature and the second sub-hidden state information provided by itself can be accurately fused, thereby improving the accuracy of the features of the contour mask image of the object output by the sub-network, thereby improving the accuracy of the video cutout.
[0384] For the above-mentioned first image generation network, when obtaining the first contour mask image based on the features of the contour mask image of the object finally obtained, the features of the contour mask image of the object finally obtained are input into the first image generation network, and the first image generation network generates an image based on the features, thereby obtaining the image output by the network as the above-mentioned first contour mask image.
[0385] The implementation manner in which the first image generation network generates an image based on the feature can be found in the aforementioned embodiment and will not be repeated here.
[0386] For the target second reconstruction subnetwork in the target transparency feature generation network, when performing feature reconstruction based on the first target feature and the second target feature, the first target feature and the second target feature are input into the target second reconstruction subnetwork, and the target second reconstruction subnetwork performs feature reconstruction based on the first target feature and the second target feature, thereby obtaining a scale-enlarged second image feature output by the target second reconstruction subnetwork, wherein the scale of the transparency mask image corresponding to the target transparency feature generation network is the same as the scale of the second image feature.
[0387] The implementation method of the target second reconstruction sub-network performing feature reconstruction based on the first target feature and the second target feature can be found in the above content and will not be repeated here.
[0388] In one embodiment of the present application, the second reconstruction sub-network is implemented based on the QARepVGG network structure.
[0389] Since the quantization calculation accuracy of the QARepVGG network is relatively high, implementing the above-mentioned second reconstruction subnetwork based on the QARepVGG network structure can improve the quantization calculation capability of the second reconstruction subnetwork, thereby improving the accuracy of feature reconstruction based on the first target feature and the second target feature by the second reconstruction subnetwork, and further improving the accuracy of video cutout.
[0390] In another embodiment of the present application, the second reconstruction subnetwork in the specific transparency feature generation network is implemented based on the QARepVGG network structure.
[0391] The specific transparency feature generation network is a transparency feature generation network whose corresponding transparency mask image has a scale smaller than a second preset scale.
[0392] The second preset scale may be a preset scale. The second preset scale may be the same as or different from the first preset scale.
[0393] When constructing the above-mentioned video cutout model, the scale of the transparency mask image corresponding to each group of transparency feature generation networks can be determined, so that the transparency feature generation network whose corresponding transparency mask image scale is smaller than the second preset scale can be determined as a specific transparency feature generation network, so that when constructing a specific transparency feature generation network, the specific transparency feature generation network is constructed based on the QARepVGG network structure. For other transparency feature generation networks, they can be constructed based on other network structures.
[0394] Since the computational amount of the U-shaped residual block in the specific transparency feature generation network constructed based on the QARepVGG network structure increases with the increase of the scale of the transparency mask image corresponding to the network, when constructing each transparency feature generation network, only the specific transparency feature generation network whose scale of the corresponding transparency mask image is smaller than the second preset scale can be used to implement the second reconstruction subnetwork in the specific transparency feature generation network based on the QARepVGG network structure. This can reduce the computational amount of each transparency feature generation network, improve the efficiency of obtaining the fusion result, thereby improving the efficiency of video matting, and can also lightweight deploy the above-mentioned video matting model in the terminal.
[0395] For the target second fusion subnetwork in the target transparency feature generation network, when the second image feature and the first sub-state information are fused and the first sub-state information is updated, the second image feature can be input into the target second fusion subnetwork, and the first sub-hidden state information provided by the target second fusion subnetwork is the first sub-state information. In this way, the target second fusion subnetwork fuses the second image feature and the first sub-hidden state information provided by itself and updates the first sub-hidden state information provided by itself, thereby obtaining the third image feature output by the target second fusion subnetwork.
[0396] The implementation manner in which the target second fusion sub-network fuses the second image feature and the first sub-hidden state information and updates the first sub-hidden state information can be referred to the aforementioned embodiment and will not be repeated here.
[0397] In one embodiment of the present application, the second fusion sub-network is a gated recurrent unit (Gated Recurrent Unit, GRU) or a long short-term memory (Long Short-Term Memory, LSTM) unit.
[0398] Both GRU and LSTM units have information memory functions. Using either of these two units as the second fusion sub-network, the unit itself can store hidden state information of the fusion features of the transparency mask image representing the objects in the segmented video frame, so that the second image features and the first sub-hidden state information provided by itself can be accurately fused, thereby improving the accuracy of the third image features output by the sub-network, thereby improving the accuracy of video cutout.
[0399] For the above-mentioned second image generation network, when the first contour mask image is obtained based on the final fusion result, the fusion result is input into the second image generation network, and the second image generation network generates an image based on the fusion result, thereby obtaining the image output by the network as the above-mentioned target transparency mask image.
[0400] The implementation method of the second image generation network generating an image based on the fusion result can be found in the aforementioned embodiment and will not be repeated here.
[0401] For the result output network, when performing regional cutout, the target transparency mask image and the first contour mask image are input into the result output network, and the result output network performs regional cutout on the target video frame according to the target transparency mask image and the first contour mask image to obtain the cutout result output by the result output network.
[0402] The implementation method of the result output network for performing region cutout on the target video frame can be found in the aforementioned embodiment and will not be described in detail here.
[0403] As can be seen from the above, in the solution provided by this embodiment, various networks and sub-networks contained in the video cutout model are used to perform video cutout. Since the video cutout model is a pre-trained video cutout model, the use of the video cutout model can improve the accuracy of video cutout. Moreover, the video cutout model does not require any interaction with other devices. Therefore, the video cutout model can be deployed in an offline device, which can improve the convenience of video segmentation.
[0404] Since the data volume of the first target feature and the second target feature is usually large, the computational complexity of the target second reconstruction sub-network for feature reconstruction is large, which results in low efficiency of feature reconstruction by the second reconstruction sub-network, and thus low efficiency of video cutout using the video cutout model.
[0405] In view of this, in one embodiment of the present application, see Fig.25 , provides a structural diagram of the second video cutout model, and Fig.23 compared to, Fig.25In the video cutout model shown, each transparency feature generation network also includes a feature screening subnetwork, the first fusion subnetwork 1 is connected to the second reconstruction subnetwork 1 through the feature screening subnetwork 1, the first fusion subnetwork 2 is connected to the second reconstruction subnetwork 2 through the feature screening subnetwork 2, and the first fusion subnetwork 3 is connected to the second reconstruction subnetwork 3 through the feature screening subnetwork 3.
[0406] For the target feature screening subnetwork in the target transparency feature generation network, before the feature output by the target first fusion subnetwork is input into the target second reconstruction subnetwork, the feature can be input into the target feature screening subnetwork, and the target feature screening subnetwork screens out target screening features that are representative of the edge contour of the object from the feature, and inputs the target screening features into the target second reconstruction subnetwork in the target transparency feature generation network, and the target second reconstruction subnetwork performs feature reconstruction based on the target screening features.
[0407] Specifically, the implementation method of the target feature screening subnetwork screening features, and the implementation method of the target second reconstruction subnetwork reconstructing features based on the target screening features and the first target features can be referred to the aforementioned embodiments, which will not be repeated here.
[0408] From the above, it can be seen that in the solution provided by this embodiment, by adding a feature screening subnetwork to the transparency feature generation network, the computational complexity of the second reconstruction subnetwork in the transparency feature generation network for feature reconstruction can be reduced, and the efficiency of the second reconstruction subnetwork for feature reconstruction can be improved, thereby improving the efficiency of the video cutout model for video cutout.
[0409] In one embodiment of the present application, see Fig.26 , provides a flowchart of the fourth video cutout method, Fig.26 In the figure, the video cutout model processes the video frame 1 and the video frame 2 contained in the video in sequence. When the video cutout model processes the video frame 1, the video frame 1 is processed by each network and sub-network in the model respectively, and the cutout result 1 corresponding to the video frame 1 is obtained, wherein each first fusion sub-network and second fusion sub-network in the model outputs information to other network layers on the one hand, and updates the hidden state information contained in itself on the other hand, which is used for the model to fuse with the input features when processing the next frame of the video frame 2. When the video cutout model processes the video frame 2, the video frame 2 is processed by each network and sub-network in the model respectively, and the cutout result 2 corresponding to the video frame 2 is obtained.
[0410] In one embodiment of the present application, when the video cutout model includes a large number of contour feature generation networks and transparency feature generation networks, the amount of computation required for the video cutout model to process video frames is also large. In view of this, the first fusion subnetwork in the last group or groups of contour feature generation networks included in the video cutout model can be removed, and the second fusion subnetwork in the last group or groups of transparency feature generation networks included in the video cutout model can also be removed, thereby reducing the amount of computation required for the video cutout model to process video frames, and the video cutout model can also be lightweight deployed in the terminal.
[0411] Take the first fusion sub-network in the network generated by removing the last set of contour features as an example, see Fig. 27 , provides a structural diagram of the third video cutout model, and Fig.23 Compared with the video cutout model shown, Fig. 27 The last set of contour feature generation networks 3 in the video cutout model shown includes only the first reconstruction subnetwork 3, and the output result of the first reconstruction subnetwork 3 is the feature of the contour mask image of the object finally obtained.
[0412] In one embodiment of the present application, see Fig.28 , provides a structural diagram of the fourth video cutout model, Fig.26 The video cutout model shown includes multiple layers of cascaded information compression networks, the output result of each layer of information compression network is a first sub-feature, and one layer of information compression network is connected to the first reconstruction sub-network in a group of contour feature generation networks, and the last layer of information compression network is connected to the first group of contour feature generation networks. The scale of the first sub-feature output by the connected information compression network is the same as the scale of the third target feature to be processed by the contour feature generation network. The first sub-feature output by the last layer of information compression network serves as the third target feature to be processed by the first group of contour feature generation networks, and the first reconstruction sub-network in other groups of contour feature generation networks performs feature reconstruction based on the third target feature and the first sub-feature output by the information compression network connected to the network, thereby improving the accuracy of feature reconstruction and improving the accuracy of video cutout.
[0413] The training process of the above video cutout model is described below.
[0414] In one embodiment of the present application, see Fig.29 , a flow chart of the first model training method is provided. In this embodiment, the above method includes the following steps S1701-S1706.
[0415] Step S1701: input the first sample video frame in the sample video into the initial model of the video matting model for processing, and obtain the first sample contour mask image of the object in the first sample video frame output by the first image generation network in the initial model.
[0416] The sample video may be any video obtained through the Internet, a video library or other channels. In addition, after obtaining the video through the Internet, a video library or other channels, multiple videos may be spliced into one video to obtain the spliced video as the sample video.
[0417] The first sample contour mask image has the same scale as the first sample video frame, and the pixel values of the pixels in the first sample contour mask image represent the confidence that the pixel at the same position in the first sample video frame predicted by the model belongs to the area where the object is located.
[0418] The above initial model is used to process the video frames input to the model according to the untrained model parameters configured by itself. In the process of the model processing the video frames, the image output by the first image generation network in the model can be obtained.
[0419] Specifically, after the first sample video frame is input into the initial model, the initial model can process the first sample video frame according to the model parameters configured by itself, and obtain the image output by the first image generation network in the model as the first sample contour mask image of the object in the first sample video frame.
[0420] Step S1702: Obtain a first difference between the annotated mask image corresponding to the first sample video frame and the annotated mask image corresponding to the second sample video frame.
[0421] The second sample video frame is: a video frame in the sample video that is before the first sample video frame and is separated by a preset number of frames.
[0422] The above-mentioned preset frame number is a pre-set frame number, for example, 3 frames, 4 frames or other numerical frame numbers.
[0423] The first sample video frame may be any video frame that is a preset number of video frames or that is after a preset number of video frames in the sample video.
[0424] The first difference may be calculated by the terminal itself, or may be calculated by other devices, and the terminal device then obtains the calculated first difference from the other devices.
[0425] The following describes an implementation method in which a terminal or other device calculates the first difference.
[0426] The terminal or other device can obtain the annotated mask image corresponding to each sample video frame in the sample video. The annotated mask image corresponding to the sample video frame can be understood as the actual mask image of the object in the sample video frame. In this way, after determining the second sample video frame according to the frame number of the first sample video frame and the preset frame number, the annotated mask image corresponding to the first sample video frame and the annotated mask image corresponding to the second sample video frame can be obtained from the annotated mask images corresponding to each obtained sample video frame, thereby calculating the first difference between the two obtained annotated mask images.
[0427] When calculating the first difference between two annotated mask images, in one implementation method, the pixel values of the pixels at the same position of the two images can be subtracted, and the number of results that are not "0" in the calculation results of each pixel point can be counted as the above-mentioned first difference, or the ratio of the number of results that are not "0" to the total number of pixels of the annotated mask image can be counted as the above-mentioned first difference; in another implementation method, the similarity of the two images can be calculated, and the calculated similarity can be subtracted from 1 to obtain the calculation result as the above-mentioned first difference.
[0428] Step S1703: obtaining a second difference between the first sample contour mask image and the second sample contour mask image, wherein the second sample contour mask image is: a mask image output by the first image generation network when the initial model processes the second sample video frame.
[0429] Specifically, similar to the aforementioned video cutout process, during the model training process, each sample video frame in the sample video can be input into the model frame by frame to obtain a sample contour mask image of the object in each sample video frame output by the first image generation network in the model. After obtaining the first sample contour mask image, a second sample video frame with a preset number of frames detected with the first sample video frame can be determined in the video frame that is before the first sample video frame and has been processed by the model, and a second sample contour mask image output by the first image generation network when the model processes the second sample video frame can be obtained, and a second difference between the first sample contour mask image and the second sample contour mask image can be calculated.
[0430] The implementation method of calculating the above-mentioned second difference is the same as the implementation method of calculating the first difference in the aforementioned step S1702, and will not be repeated here.
[0431] Step S1704: Obtain a third difference between the first sample transparency mask image and the second sample transparency mask image.
[0432] The first sample transparency mask image is: a mask image output by the second image generation network when the initial model processes the first sample video frame.
[0433] The second sample transparency mask image is: the mask image output by the second image generation network when the initial model processes the second sample video frame.
[0434] Specifically, after the first sample video frame is input into the initial model, the initial model can process the first sample video frame according to the model parameters configured by itself, and obtain the image output by the second image generation network in the model as the above-mentioned first sample transparency mask image; after the second sample video frame is input into the initial model, the initial model can process the second sample video frame according to the model parameters configured by itself, and obtain the image output by the second image generation network in the model as the above-mentioned second sample transparency mask image, so that after obtaining the first sample transparency mask image and the second sample transparency mask image, the third difference between the first sample transparency mask image and the second sample transparency mask image can be calculated.
[0435] The implementation method of calculating the third difference is the same as the implementation method of calculating the first difference in the aforementioned step S1702, and will not be repeated here.
[0436] Step S1705: Calculate the training loss based on the first difference, the second difference and the third difference.
[0437] The training loss is calculated based on the first difference, the second difference, and the third difference. The training loss can be calculated using a loss function, an algorithm, etc.
[0438] Step S1706: Based on the training loss, adjust the model parameters of the initial model to obtain a video cutout model.
[0439] Specifically, based on the training loss, the model parameters of the initial model can be adjusted by any of the following three implementation methods.
[0440] In the first implementation method, for each model parameter in the initial model, the correspondence between the training loss and the adjustment range of the model parameter can be set in advance. After the training loss is calculated, the actual adjustment range of the model parameter can be calculated according to the correspondence, and the model parameter can be adjusted according to the actual adjustment range.
[0441] In the second implementation method, the initial model usually needs to be trained using a large amount of sample data. During the training process, the training loss needs to be continuously calculated, and the model parameters of the initial model need to be continuously adjusted based on the training loss. In view of this, after calculating the training loss, the difference in training loss changes can be determined based on the training loss and the previously calculated training loss, and then the model parameters of the initial model can be adjusted based on the difference in changes.
[0442] In the third implementation method, based on the training loss, the model parameters of the initial model can be adjusted using model parameter adjustment algorithms, functions, etc.
[0443] As can be seen from the above, in the solution provided by the present embodiment, since there is often a time domain correlation between the first sample video frame and the second sample video frame spaced by a preset number of frames, a first difference between the annotated mask image corresponding to the first sample video frame and the annotated mask image corresponding to the second sample video frame, a second difference between the first sample contour mask image and the second sample contour mask image, and a third difference between the first sample transparency mask image and the second sample transparency mask image are obtained. The training loss is calculated based on the first difference, the second difference, and the third difference. When the model parameters of the initial model are adjusted based on the training loss, the initial model can learn the time domain correlation between different video frames of the video, thereby improving the accuracy of the trained model, and then using the model to perform video cutout, the accuracy of video cutout can be improved.
[0444] When obtaining the second difference, in addition to the method mentioned in step S1703, the second difference can also be obtained by Fig.30 The method mentioned in step S1703A in the illustrated embodiment is used for obtaining.
[0445] In one embodiment of the present application, see Fig.30 , a flowchart of the second model training method is provided.
[0446] In this embodiment, the first sample contour mask image includes: a first mask sub-image identifying an area where the object is located in the first sample video frame and a second mask sub-image identifying an area outside the object in the first sample video frame.
[0447] The pixel values of the pixels in the first mask sub-image represent the confidence that the pixel at the same position in the first sample video frame predicted by the model belongs to the area where the object is located, and the pixel values of the pixels in the second mask sub-image represent the confidence that the pixel at the same position in the first sample video frame predicted by the model belongs to the area outside the object.
[0448] like Fig.31 As shown, Fig.31 A mask image provided in an embodiment of the present application is a first sample mask image.
[0449] Fig.31 The mask image shown includes two sub-images, namely: a first mask sub-image identifying the area where the object is located in the first sample video frame and a second mask sub-image identifying the area outside the object in the first sample video frame.
[0450] The second sample contour mask image includes: a third mask sub-image identifying the area where the object is located in the second sample video frame and a fourth mask sub-image identifying the area outside the object in the second sample video frame.
[0451] The pixel values of the pixels in the third mask sub-image represent the confidence that the pixel at the same position in the second sample video frame predicted by the model belongs to the area where the object is located, and the pixel values of the pixels in the fourth mask sub-image represent the confidence that the pixel at the same position in the second sample video frame predicted by the model belongs to the area outside the object.
[0452] In this case, the above step S1703 can be implemented by the following step S1703A.
[0453] Step S1703A: Obtain a difference between the first mask sub-image and the third mask sub-image, and obtain a difference between the second mask sub-image and the fourth mask sub-image, to obtain a second difference including the obtained differences.
[0454] The implementation manner of obtaining the difference between the first mask sub-image and the third mask sub-image and the difference between the second mask sub-image and the fourth mask sub-image is the same as the implementation manner of obtaining the first difference or the second difference described above, and will not be repeated here.
[0455] After obtaining the two differences, namely the difference between the first mask sub-image and the third mask sub-image and the difference between the second mask sub-image and the fourth mask sub-image, the two differences can be accumulated to obtain a second difference including the two differences, or the average of the two differences can be used as the second difference, or the larger difference between the two differences can be determined as the second difference, etc.
[0456] From the above, it can be seen that in the solution provided by this embodiment, since the area in the video frame is composed of two types of areas, namely the area where the object is located and the area outside the object, the greater the difference in the area where the object is located in different video frames, the greater the difference in the area outside the object in different video frames. It can be seen that the difference in the area outside the object can also reflect the difference in the area where the object is located. Therefore, the second difference is obtained based on the two differences, namely the difference between the first mask sub-image and the third mask sub-image and the difference between the second mask sub-image and the fourth mask sub-image. The second difference is calculated comprehensively from two different angles. This can improve the accuracy of the second difference, thereby improving the accuracy of model training, and improving the accuracy of video cutout using the model.
[0457] When the second fusion subnetwork in the above-mentioned video cutout model fuses the features and the hidden state information provided by itself, it can ensure that the features of the transparency mask image corresponding to the segmented video frame of the video are considered in the process of cutout of the target video frame, that is, the time domain continuity between the video frames is guaranteed. However, in this case, if the second image generation network in the model has no hard restriction that the output image is a binary image, there may be semi-transparent areas in the image finally output by the second image generation network, and the pixel value range of the pixel points in the target transparency mask image output by the second image generation network is 0-1, and semi-transparent areas may appear in the target transparency mask image output by itself. In this way, it is difficult to determine the reason for the semi-transparent area in the target transparency mask image, and it is also difficult to train the networks and sub-networks in the cutout branch of the model.
[0458] In view of this, in one embodiment of the present application, see Fig.32 , a schematic diagram of the structure of a second image generation network is provided, Fig.32 In the paper, the second image generation network includes a hard segmentation subnetwork, a result fusion subnetwork and an image generation subnetwork.
[0459] Among them, the input of the hard segmentation sub-network is the fusion result of the output of the second fusion sub-network in the last set of transparency feature generation network.
[0460] The hard segmentation subnetwork is used to obtain a second contour mask image of the object in the target video frame based on the fusion result.
[0461] The result fusion subnetwork is used to fuse the second contour mask image and the fusion result to obtain the target fusion feature.
[0462] The image generation subnetwork is used to obtain a target transparency mask image of the object edge in the target video frame based on the target fusion features.
[0463] In the process of model training, not only can the initial model be trained based on the above-mentioned first difference, second difference and third difference, but also the mask images corresponding to different sample video frames output by the hard segmentation sub-network can be obtained, and the fourth difference between the obtained mask images can be calculated, so as to train the initial model according to the fourth difference. Since the hard segmentation sub-network outputs a contour mask image belonging to a binary image, the possibility of a semi-transparent area appearing in the mask image output by the second image generation network itself can be eliminated during the training process, so that the training of the initial model can be accurately and quickly realized.
[0464] like Fig.33 As shown, Fig.33 The left side shows the final cutout result when the fourth difference is not used for model training. Fig.33The right side shows the final cutout result when the fourth difference is used for model training.
[0465] Next, the electronic device provided by the embodiment of the present application is introduced.
[0466] Electronic equipment can be equipped Or portable terminal devices with other operating systems, such as mobile phones, tablet computers, desktop computers, laptop computers, handheld computers, notebook computers, ultra-mobile personal computers (UMPC), netbooks, as well as cellular phones, personal digital assistants (PDA), augmented reality (AR) devices, virtual reality (VR) devices, artificial intelligence (AI) devices, wearable devices, vehicle-mounted devices, smart home devices and / or smart city devices, etc.
[0467] Fig.34 The electronic device 100 provided by the embodiment of the present application is exemplarily shown. The electronic device 100 may support AOD, and its screen has pixel-level light-emitting capability and supports lighting only some pixels of the screen, such as an organic light-emitting diode (OLED) screen.
[0468] The electronic device 100 may include a processor 110, a display screen 120, a camera 130, an internal memory 140, a SIM (Subscriber Identification Module) card interface 150, a USB (Universal Serial Bus) interface 160, a charging management module 170, a power management module 171, a battery 172, a sensor module 180, a mobile communication module 190, a wireless communication module 200, an antenna 1 and an antenna 2, etc. The sensor module 180 may include a pressure sensor 180A, a fingerprint sensor 180B, a touch sensor 180C, an ambient light sensor 180D, etc.
[0469] The structure illustrated in the embodiment of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown in the figure, or combine some components, or split some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0470] The processor 110 may include one or more processing units, for example: the processor 110 may include a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processor (NPU), etc. Among them, different processing units may be independent components or integrated into one or more processors. In some embodiments, the electronic device 100 may also include one or more processors 110. Among them, the controller may generate an operation control signal according to the instruction opcode and the timing signal to complete the control of fetching and executing instructions. In some other embodiments, a memory may also be provided in the processor 110 for storing instructions and data. Exemplarily, the memory in the processor 110 may be a cache memory. The memory may store instructions or data that the processor 110 has just used or circulated. If the processor 110 needs to use the instruction or data again, it may be directly called from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the electronic device 100 in processing data or executing instructions.
[0471] In some embodiments, the processor 110 may include one or more interfaces. The interface may include an Inter-Integrated Circuit (I2C) interface, an Inter-Integrated Circuit Sound (I2S) interface, a Pulse Code Modulation (PCM) interface, a Universal Asynchronous Receiver / Transmitter (UART) interface, a Mobile Industry Processor Interface (MIPI), a General-Purpose Input / Output (GPIO) interface, a SIM card interface and / or a USB interface, etc. Among them, the USB interface 160 is an interface that complies with the USB standard specification, and can specifically be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 160 can be used to connect a charger to charge the electronic device 100, and can also be used to transmit data between the electronic device 100 and a peripheral device. The USB interface 160 can also be used to connect headphones to play audio through the headphones.
[0472] The interface connection relationship between the modules illustrated in the embodiment of the present application is for illustrative purposes only and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0473] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 190, the wireless communication module 200, the modem processor, the baseband processor, and the like.
[0474] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve the utilization of antennas. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.
[0475] The electronic device 100 implements the display function through a GPU, a display screen 120, and an application processor. The GPU is a microprocessor for image processing, which connects the display screen 120 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.
[0476] The display screen 120 is used to display images, videos, etc. The display screen 120 includes a display panel, which may be an organic light emitting diode (OLED), an active-matrix organic light emitting diode, or an active-matrix organic light emitting diode (AMOLED) or other display panel with pixel-level light emitting capability.
[0477] The electronic device 100 can realize the shooting function through the ISP, the camera 130, the video codec, the GPU, the display screen 120 and the application processor, etc., wherein the camera 130 includes a front camera and a rear camera.
[0478] The ISP is used to process the data fed back by the camera 130. For example, when shooting, the shutter is opened, and light is transmitted to the camera photosensitive element through the lens. The light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can perform algorithm optimization on the noise, brightness and color of the image. The ISP can also optimize the exposure and color temperature of the shooting scene and other parameters. In some embodiments, the ISP can be set in the camera 130.
[0479] The camera 130 is used to take photos or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal oxide semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard red green blue (RGB), YUV or other format. In some embodiments, the electronic device 100 may include 1 or N cameras 130, where N is a positive integer greater than 1.
[0480] The digital signal processor is used to process digital signals, and can process not only digital image signals but also other digital signals. For example, when the electronic device 100 is selecting a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.
[0481] Video codecs are used to compress or decompress digital videos. The electronic device 100 may support one or more video codecs. Thus, the electronic device 100 may play or record videos in a variety of coding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.
[0482] NPU is a neural network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission mode between neurons in the human brain, it can quickly process input information and can also continuously self-learn. NPU can realize applications such as intelligent cognition of electronic device 100, such as image recognition, face recognition, voice recognition, text understanding, etc.
[0483] The internal memory 140 can be used to store one or more computer programs, which include instructions. The processor 110 can execute the above instructions stored in the internal memory 140, so that the electronic device 100 performs the video cutout method provided in some embodiments of the present application, as well as various applications and data processing. The internal memory 140 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system; the program storage area can also store one or more applications (such as a gallery, contacts, etc.). The data storage area can store data (such as photos, contacts, etc.) created during the use of the electronic device 100. In addition, the internal memory 140 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more disk storage components, flash memory components, universal flash memory (Universal Flash Storage, UFS), etc. In some embodiments, the processor 110 can execute the video cutout method provided in the embodiment of the present application, as well as other applications and data processing by running instructions stored in the internal memory 140, and / or instructions stored in a memory provided in the processor 110.
[0484] The internal memory 140 can be used to store program codes or program instructions for implementing the wallpaper display method and video cutout method provided in the embodiments of the present application. The processor 110 can call the program codes or program instructions of the wallpaper display method and video cutout method stored in the internal memory 140 to execute the wallpaper display method and video cutout method of the embodiments of the present application.
[0485] The sensor module 180 may include a pressure sensor 180A, a fingerprint sensor 180B, a touch sensor 180C, an ambient light sensor 180D, and the like.
[0486] The pressure sensor 180A is used to sense the pressure signal and can convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 180A can be set on the display screen 120. There are many types of pressure sensors 180A, for example, they can be resistive pressure sensors, inductive pressure sensors or capacitive pressure sensors. The capacitive pressure sensor can be a parallel plate including at least two conductive materials. When the force acts on the pressure sensor 180A, the capacitance between the electrodes changes, and the electronic device 100 determines the intensity of the pressure according to the change in capacitance. When the touch operation acts on the display screen 120, the electronic device 100 detects the touch operation according to the pressure sensor 180A. The electronic device 100 can also calculate the position of the touch according to the detection signal of the pressure sensor 180A. In some embodiments, touch operations acting on the same touch position but with different touch operation intensities can correspond to different operation instructions. For example: when a touch operation with a touch operation intensity less than the first pressure threshold acts on a short message application icon, an instruction to view the short message is executed; when a touch operation with a touch operation intensity greater than or equal to the first pressure threshold acts on a short message application icon, an instruction to create a new short message is executed.
[0487] The fingerprint sensor 180B is used to collect fingerprints. The electronic device 100 can use the collected fingerprint characteristics to realize functions such as unlocking, accessing application locks, taking photos, and answering calls.
[0488] The touch sensor 180C is also called a touch control device. The touch sensor 180C can be arranged on the display screen 120, and the touch sensor 180C and the display screen 120 form a touch screen, which is also called a touch control screen. The touch sensor 180C is used to detect a touch operation acting on or near it. The touch sensor 180C can pass the detected touch operation to the application processor to determine the type of touch event. A visual output related to the touch operation can be provided through the display screen 120. In other embodiments, the touch sensor 180C can also be arranged on the surface of the electronic device 100, and is arranged at a different position from the display screen 120.
[0489] The ambient light sensor 180D is used to sense the brightness of the ambient light. The electronic device 100 can adaptively adjust the brightness of the display screen 120 according to the perceived brightness of the ambient light. The ambient light sensor 180D can also be used to automatically adjust the white balance when shooting. The ambient light sensor 180D can also transmit the environmental information of the device to the GPU.
[0490] The ambient light sensor 180D is also used to obtain the brightness, light ratio, color temperature, etc. of the environment in which the camera 130 captures images.
[0491] Fig.35A software system architecture used by the electronic device 100 is shown. The software system architecture may use a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture.
[0492] The layered architecture divides the terminal's software system into several layers, each with clear roles and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the software system can be divided into five layers, namely, application layer (applications), application framework layer (application framework), system library, hardware abstraction layer (HAL) and kernel layer (kernel).
[0493] The application layer can include a series of application packages. The application layer runs applications by calling the application programming interface (API) provided by the application framework layer. Fig.35 As shown, the application package may include applications such as browser, gallery, music, and video. It is understandable that the port of each of the above applications may be used to receive data.
[0494] The application framework layer provides API and programming framework for the applications in the application layer. The application framework layer includes some predefined functions. Fig.35 As shown, the application framework layer may include a window manager, a content provider, a view system, a resource manager, a notification manager, and a DHCP (Dynamic Host Configuration Protocol) module, etc.
[0495] The system library can include multiple functional modules, such as surface manager, 3D graphics processing library, 2D graphics engine and file library.
[0496] The hardware abstraction layer may include multiple library modules, such as display library modules and motor library modules. The terminal system may load the corresponding library modules for the device hardware, thereby enabling the application framework layer to access the device hardware.
[0497] The kernel layer is a layer between hardware and software. The kernel layer is used to drive the hardware so that the hardware works. The kernel layer at least includes display driver, audio driver, sensor driver and motor driver, etc., which are not limited by the present application embodiment. It can be understood that display driver, audio driver, sensor driver and motor driver, etc. can all be regarded as a driver node. Each of the above-mentioned driver nodes includes an interface that can be used to receive data.
[0498] The term "user interface (UI)" in the specification, claims and drawings of this application refers to the medium interface for interaction and information exchange between an application or operating system and a user, which realizes the conversion between the internal form of information and the form acceptable to the user. The user interface of an application is source code written in a specific computer language such as Java and extensible markup language (XML). The interface source code is parsed and rendered on the terminal device, and finally presented as content that the user can recognize, such as pictures, text, buttons and other controls. Controls, also known as widgets, are the basic elements of the user interface. Typical controls include toolbars, menu bars, text boxes, buttons, scroll bars, pictures and text. The properties and contents of controls in the interface are defined by tags or nodes, such as XML through <textview> 、 <imgview> 、
[0499] <videoview>The controls contained in the interface are specified by nodes such as <head> and <body>. A node corresponds to a control or attribute in the interface. After being parsed and rendered, the node is presented as user-visible content. In addition, many applications, such as hybrid applications, usually also contain web pages in their interfaces. A web page, also known as a page, can be understood as a special control embedded in the application interface. A web page is source code written in a specific computer language, such as hypertext markup language (HTML), cascading style sheets (CSS), JavaScript (JS), etc. The web page source code can be loaded and displayed as user-recognizable content by a browser or a web page display component with similar functions to a browser. The specific content contained in a web page is also defined by tags or nodes in the web page source code. For example, HTML is defined by <body>. 、 、 <video> 、 <canvas>To define the elements and attributes of a web page.
[0500] The most common form of user interface is graphical user interface (GUI), which refers to a user interface related to computer operation that is displayed in a graphical manner. It can be an icon, window, control or other interface element displayed on the display screen of an electronic device, where a control can include icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, widgets and other visual interface elements.
[0501] Each step in the above method embodiment provided by the present application can be completed by an integrated logic circuit of hardware in a processor or by instructions in the form of software. The method steps disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware processor, or by a combination of hardware and software modules in a processor.
[0502] The present application also provides an electronic device, which may include: a memory and a processor, wherein the memory may be used to store a computer program; and the processor may be used to call the computer program in the memory so that the electronic device executes the method in any one of the above embodiments.
[0503] The present application also provides a chip system, such as Fig.36 As shown, the chip system may include a processor 2201 for implementing the functions involved in the method performed by the electronic device in any of the above embodiments. The chip system may also include a memory, the memory is used to store program instructions and data, and the memory is located inside or outside the processor. The chip system may also include an input and output interface, and the input and output interface can be used as a bridge connecting the processor and peripheral input and output devices. Specifically, when the peripheral input device collects input data, such as the touch screen collects the user's touch operation, the audio input circuit collects audio data, the camera collects image data, etc., the input and output interface can transmit the input data to the processor for processing; when the processor completes the processing of the input data and generates a processing result, the input and output interface can transmit the processing result to the peripheral output device for output, such as transmitting the processing result to the display screen for display, transmitting the processing result to the audio output circuit for voice output, etc.
[0504] The chip system may be composed of the chip, or may include the chip and other discrete devices.
[0505] Optionally, the processor in the chip system may be one or more. The processor may be implemented by hardware or by software. When implemented by hardware, the processor may be a logic circuit, an integrated circuit, etc. When implemented by software, the processor may be a general-purpose processor implemented by reading software code stored in a memory.
[0506] Optionally, the memory in the chip system may also be one or more. The memory may be integrated with the processor or may be separately arranged with the processor, which is not limited in the embodiments of the present application. Exemplarily, the memory may be a non-transient processor, such as a read-only memory ROM, which may be integrated with the processor on the same chip or may be arranged on different chips respectively. The embodiments of the present application do not specifically limit the type of memory and the arrangement of the memory and the processor.
[0507] Exemplarily, the chip system can be a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on chip (SoC), a central processor unit (CPU), a network processor (NP), a digital signal processor (DSP), a microcontroller unit (MCU), a programmable logic device (PLD) or other integrated chips.
[0508] The present application also provides a computer program product, which includes: a computer program (also referred to as code, or instruction), which enables a computer to execute the method executed by the electronic device in any of the above embodiments when the computer program is executed.
[0509] The present application also provides a computer-readable storage medium, which stores a computer program (also referred to as code or instruction). When the computer program is executed, the computer executes the method executed by the electronic device in any of the above embodiments.
[0510] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the process or function according to the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, a computer, a server or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or a data center that includes one or more available media integrations. Available media can be magnetic media, (e.g., floppy disk, hard disk, tape), optical media (e.g., DVD), or semiconductor media (e.g., solid-state drive Solid State Disk), etc.
[0511] Those skilled in the art can understand that to implement all or part of the processes in the above-mentioned embodiments, the processes can be completed by computer programs to instruct related hardware, and the programs can be stored in computer-readable storage media. When the programs are executed, they can include the processes of the above-mentioned method embodiments. The aforementioned storage media include: ROM or random access memory RAM, magnetic disk or optical disk and other media that can store program codes.
[0512] In short, the above are only embodiments of the technical solution of the present invention, and are not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made according to the disclosure of the present invention shall be included in the protection scope of the present invention.< / canvas> < / video> < / videoview> < / imgview> < / textview>
Claims
1. A wallpaper display method, characterized in that: include: An electronic device in a locked state or unlocked state detects a first operation, and in response to the first operation, the electronic device displays a first interface, wherein the first interface includes an AOD wallpaper; When the electronic device displays the first interface, in response to a second operation, the electronic device displays a second interface, where the second interface includes a lock screen wallpaper; In response to a third operation for unlocking, the electronic device displays a third interface, wherein the third interface includes a desktop wallpaper; Among them, the AOD wallpaper, the lock screen wallpaper and the desktop wallpaper respectively correspond to the first video frame, the second video frame and the third video frame in the first video, the second video frame is after the first video frame, and the third video frame is after the second video frame.
2. The method according to claim 1, characterized in that The first video frame, the second video frame, and the third video frame are all one frame, the frame interval between the first video frame and the second video frame is greater than the first frame interval, and the frame interval between the second video frame and the third video frame is greater than the second frame interval.
3. The method according to claim 1, characterized in that The first video frame, the second video frame, and the third video frame are all multiple frames, the start frame in the second video frame is the end frame in the first video frame, and the start frame in the third video frame is the end frame in the second video frame.
4. The method according to any one of claims 1 to 3, characterized in that The first video is selected from a gallery of videos.
5. The method according to any one of claims 1 to 4, characterized in that The first video is selected from videos provided by a settings application.
6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: The electronic device displays a fourth interface, wherein the fourth interface includes a video frame sequence of the second video; The first video is selected from the video frame sequence, wherein the first video is a partial video clip in the second video or the entire second video.
7. The method according to claim 6, characterized in that There is a selection box on the video frame sequence, and a video segment in the selection box is the first video; the user operation of selecting the video segment is specifically an operation of sliding the selection box along the video frame sequence.
8. The method according to claim 6 or 7, characterized in that The fourth interface also includes a first preview, which is used to display the display status of the AOD wallpaper, the lock screen wallpaper and the desktop wallpaper at the first interface, the second interface and the third interface respectively.
9. The method according to claim 8, characterized in that The first preview is a video, and when played, the first preview presents a dynamic display process of switching from the AOD wallpaper to the lock screen wallpaper and then to the desktop wallpaper.
10. The method according to any one of claims 1 to 9, characterized in that The method further includes: the electronic device displays a fifth interface, and the main object and the background of the wallpaper in the fifth interface are allowed to be selected separately by the user.
11. The method according to claim 10, characterized in that Before the electronic device displays the fifth interface, the method further includes: the electronic device cuts out the main object in the first video to separate the main object from the background.
12. The method according to claim 11, characterized in that The method further comprises: In response to a user operation of selecting the subject object for subject object editing, the electronic device updates the image of the subject object, and refreshes the AOD wallpaper, the lock screen wallpaper, and the desktop wallpaper based on the updated subject object image.
13. The method according to claim 11 or 12, characterized in that Also includes: In response to a user operation of selecting a background of the wallpaper for background editing, the electronic device updates an image of the background, and refreshes the AOD wallpaper, the lock screen wallpaper, and the desktop wallpaper based on the updated background image.
14. An electronic device, characterized in that: The method comprises one or more processors and one or more memories; wherein the one or more memories are coupled to the one or more processors, and the one or more memories are used to store computer program codes, wherein the computer program codes include computer instructions, and when the one or more processors execute the computer instructions, the method as described in any one of claims 1 to 13 is executed.
15. A chip system, the chip system is applied to electronic equipment, the chip system comprises one or more processors, characterized in that: The processor is configured to call computer instructions so as to execute the method according to any one of claims 1 to 13.
16. A computer-readable storage medium, comprising a computer-executable program, characterized in that: When the computer executable program is run on an electronic device, the method according to any one of claims 1 to 13 is executed.
Citation Information
Cited By
Wallpaper display method and electronic device
EP4787825A1
Wallpaper display method and electronic device
WO2025092127A1