Wallpaper display method and electronic device

By generating coherent dynamic wallpapers between the AOD, lock screen and desktop interfaces of electronic devices, the problem of fragmented wallpaper display is solved, the user experience is improved, personalized settings are supported, and a natural interface transition is achieved.

WO2025092127A9PCT designated stage expired Publication Date: 2025-10-16HONOR DEVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/112315
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-02
Filing Date
2024-08-15
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

In the prior art, wallpaper display effects are fragmented between different interfaces of electronic devices, lacking coherence and dynamism, resulting in poor user experience.

Method used

By detecting user operations to switch between different interfaces of electronic devices, a coherent dynamic wallpaper is generated from AOD to lock screen to desktop. The connection of video frames and image processing technology are used to ensure the natural transition of wallpaper between different interfaces, and users can customize and edit wallpaper content.

Benefits of technology

The system realizes the coherent and dynamic display of wallpapers between electronic device interfaces, improves the user experience, and provides the convenience of generating personalized wallpapers and a natural transition effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024112315_16102025_PF_FP_ABST
    Figure CN2024112315_16102025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a wallpaper display method and an electronic device. In the method, the electronic device may display an AOD interface, a lock screen interface and an unlock screen interface in sequence, and display a wallpaper on the AOD interface, the lock screen interface and a desktop, wherein the wallpaper can be generated on the basis of a video, and video frames used by the wallpaper on the AOD interface, the lock screen interface, and the desktop are arranged in sequence, forming a coherent and dynamic wallpaper playback screen as a user switches from the AOD interface to the lock screen interface and then to the desktop. Moreover, the wallpaper is well integrated with interfaces such as AOD, lock screen, and desktop, ensuring that the subject in the wallpaper is not shielded by non-interactive elements such as time on these interfaces. In addition, the electronic device can update the wallpaper on the basis of user-customized or edited operations, so that the user can set a personalized wallpaper, and user operations are simple.
Need to check novelty before this filing date? Find Prior Art

Description

Wallpaper display method and electronic device

[0001] The present application claims priority to the Chinese patent application No. 202311458292.9, filed on November 2, 2023, and entitled "Wallpaper display method and electronic device", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] The present application relates to the field of electronic technology, and in particular to a wallpaper display method and an electronic device. BACKGROUND

[0003] When using an electronic device such as a mobile phone, wallpaper is the interface that users see most frequently and can also reflect the personality of the user. Currently, wallpaper is mainly applied to user interfaces such as always on display (AOD), lock screen, and desktop. Users often like to set wallpaper in multiple places such as AOD, lock screen, and desktop to pursue aesthetics and individuality.

[0004] SUMMARY

[0005] Embodiments of the present application provide a wallpaper display method and an electronic device, which can generate a coherent and dynamic wallpaper from AOD to lock screen to desktop based on user-defined operations, meet the user's demand for customizing wallpaper, and provide a rich and colorful space for users to create wallpaper.

[0006] In a first aspect, embodiments of the present application provide a wallpaper display method, which can include: an electronic device in a lock screen state or an unlocked state detecting a first operation, in response to the first operation, the electronic device displaying a first interface, the first interface including an AOD wallpaper. When the electronic device displays the first interface, in response to a second operation, the electronic device displays a second interface, the second interface including a lock screen wallpaper. In response to a third operation for unlocking, the electronic device displays a third interface, the third interface including a desktop wallpaper.

[0007] Wherein, the AOD wallpaper, the lock screen wallpaper and the desktop wallpaper can correspond to a first video frame, a second video frame and a third video frame in a first video respectively, the second video frame is after the first video frame, and the third video frame is after the second video frame.

[0008] In the first aspect, the first interface, the second interface, and the third interface can be an AOD interface, a lock screen interface, and a desktop interface, respectively.

[0009] In the first aspect, the first operation can be a screen-off operation, for example, an operation of pressing the power key to turn off the screen in the electronic device in a lock screen state or an unlock state. In the lock screen state or the unlock state, the screen of the electronic device is in a lighted state, where the lock screen interface is displayed on the screen in the lock screen state, and the desktop interface or the user interface of the application opened by the user is displayed on the screen in the unlock state. Not limited to being triggered by the screen-off operation, if the user does not operate the mobile phone for a long time in the lock screen state or the unlock state, the electronic device can also display the AOD interface.

[0010] In the first aspect, the second operation can be a user operation of turning on the screen, for example, an operation of tapping the screen, an operation of pressing the power key, and the like, but is not limited thereto. The screen can also be turned on by a timer in the AOD interface, and the user can set the time for turning on the screen. When the time comes, the electronic device turns on the screen. The embodiments of the present application do not limit the operation of turning on the screen in the AOD interface.

[0011] By implementing the method provided in the first aspect, the AOD wallpaper, the lock screen wallpaper, and the desktop wallpaper are linked together from AOD to the lock screen to the desktop, and they are no longer isolated from each other.

[0012] In combination with the first aspect, in some embodiments, the first video frame, the second video frame, and the third video frame are each one frame, the frame interval between the first video frame and the second video frame is greater than the first frame interval, and the frame interval between the second video frame and the third video frame is greater than the second frame interval. For example, the first video frame, the second video frame, and the third video frame can be extracted from the first video at a certain frame interval (e.g., more than 1 second). In this way, the pictures of the multiple video frames can be different, and present a dynamic change effect when played continuously.

[0013] In combination with the first aspect, in some embodiments, the first video frame, the second video frame, and the third video frame are each multiple frames, the starting frame in the second video frame is the ending frame in the first video frame, and the starting frame in the third video frame is the ending frame in the second video frame. In other words, the video frames used by the AOD wallpaper, the lock screen wallpaper, and the desktop wallpaper need to be connected at the beginning and the end according to the interface switching time, and a dynamic effect is achieved from AOD to the lock screen to the desktop as a whole.

[0014] In combination with the first aspect, in some embodiments, the wallpaper can be well integrated with the AOD, the lock screen, the desktop, and the like, and the main object (e.g., a person) in the wallpaper will not be blocked by the non-interactive elements such as time in the AOD, the lock screen, the desktop, and the like. The interactive elements in the AOD, the lock screen, or the desktop, such as the fingerprint unlocking key in the AOD interface, the password keyboard in the lock screen interface, and the application icon in the desktop, can be arranged above the main object in the wallpaper, so as to facilitate the user to operate these interactive elements and ensure the implementation of the interface interaction function.

[0015] In some embodiments in combination with the first aspect, the AOD wallpaper, lock screen wallpaper and desktop wallpaper can directly use the original video frames in the first video, or use image frames after image processing such as cropping and filter on the original video frames.

[0016] In some embodiments in combination with the first aspect, the first video can be selected from the gallery video. For example, as shown in FIG. 3B, the user can select the video clip 14 as the first video for generating the dynamic wallpaper. The first video can be selected from the videos in the gallery, or selected from the live photos in the gallery. The live photo is also a kind of video. In this way, the user's demand for customizing the wallpaper can be met.

[0017] In some embodiments in combination with the first aspect, the first video is selected from the videos provided by the settings application.

[0018] In some embodiments, the method provided by the first aspect is not limited to the scenario from AOD to lock screen to desktop, but can also be applied to other adjacent interface transition scenarios, such as AOD to desktop to some application interface, which especially occurs when the user does not set the lock screen, and such as desktop to application main interface to application secondary interface. Of course, these application interfaces can support wallpaper display, rather than application interfaces such as game display interface, video playing interface, etc. that are not suitable for wallpaper display. For example, the chat interface of “WeChat” can set chat wallpaper. That is, the electronic device displays a plurality of user interfaces of different levels in sequence, and the electronic device also displays a wallpaper at the plurality of user interfaces, the wallpaper is generated based on a first video, and video frames adopted by the wallpaper at the plurality of user interfaces of different levels are arranged in sequence; wherein the main object in the wallpaper is displayed above the non-interactive elements in the user interface, and below the interactive elements in the user interface. In this way, when the user performs multi-level interface transition, a coherent and dynamic wallpaper display effect is formed as a whole.

[0019] In some embodiments in combination with the first aspect, the wallpaper display method can further include: displaying, by the electronic device, a fourth interface, the fourth interface including a video frame sequence of a second video. Selecting, from the video frame sequence, a first video, wherein the first video is a partial video clip in the second video or the entire second video. Wherein the fourth interface can be, for example, the “video wallpaper preview” shown in FIG. 10E, the second video can be the video 16, and the video frame sequence of the second video can be the video frame sequence 18.

[0020] In some embodiments, there can be a selection box on the video frame sequence, such as the selection box 19 in FIG. 10E, and the video in the selection box is the first video. The user operation of selecting the video clip is specifically an operation of sliding the selection box along the video frame sequence.

[0021] In some embodiments, the fourth interface further includes a first preview for showing a display state of the AOD wallpaper, the lock screen wallpaper and the desktop wallpaper at the first interface, the second interface and the third interface respectively. The first preview can be the preview 17 in FIG. 10E. The first preview can be a video, and the first preview presents a dynamic display process of switching from the AOD wallpaper to the lock screen wallpaper and then to the desktop wallpaper when playing.

[0022] In combination with the first aspect, in some embodiments, the wallpaper display method can further include: displaying, by the electronic device, a fifth interface in which the subject and the background of the wallpaper are allowed to be individually selected by the user. The fifth interface can be, for example, the "wallpaper editing" interface shown in FIG. 11A. In some embodiments, the wallpaper display method can further include: in response to a user operation of selecting the subject for subject editing, updating, by the electronic device, the image of the subject, and refreshing the AOD wallpaper, the lock screen wallpaper and the desktop wallpaper based on the updated subject image. In some embodiments, the wallpaper display method can further include: in response to a user operation of selecting the background of the wallpaper for background editing, updating, by the electronic device, the image of the background, and refreshing the AOD wallpaper, the lock screen wallpaper and the desktop wallpaper based on the updated background image.

[0023] In some embodiments, before the electronic device displays the fifth interface, the wallpaper display method can further include: performing, by the electronic device, matting on the first video for the subject, and separating the subject and the background.

[0024] In combination with the first aspect, in some embodiments, the electronic device performs matting on the first video for the subject, and specifically includes: performing information compression on a target video frame in the video to obtain first image features; the video is a video clip; based on the first image features, performing segmentation on the target video frame to obtain features of a contour mask image of an object obtained in the segmentation process and a first contour mask image of the object in the target video frame; the object is the subject; based on the first image features and the obtained features of the contour mask image, performing feature reconstruction, fusing and updating first hidden state information of the video obtained by the reconstruction to obtain a fusion result, wherein the first hidden state information represents: a fusion feature of a transparency mask image of an object edge in a video frame before the target video frame for matting; based on the fusion result, obtaining a target transparency mask image of the object edge in the target video frame; and based on the target transparency mask image and the first contour mask image, performing region matting on the target video frame to obtain a matting result.

[0025] In some embodiments of the first aspect, based on the fusion result, a target transparency mask image of the object edge in the target video frame is obtained, including: based on the fusion result, a second contour mask image of the object in the target video frame is obtained; the second contour mask image and the fusion result are fused to obtain a target fusion feature; and based on the target fusion feature, the target transparency mask image of the object edge in the target video frame is obtained.

[0026] In some embodiments of the first aspect, based on the first image feature, the target video frame is segmented to obtain a feature of a contour mask image of the object obtained in the segmentation process and a first contour mask image of the object in the target video frame, including: based on the first image feature, cascade feature reconstruction is performed to obtain features of contour mask images of the object with increasing scales, and based on the feature obtained in the last processing, the first contour mask image of the object in the target video frame is obtained; the first hidden state information includes a plurality of first sub-hidden state information, and each first sub-hidden state information represents a fusion feature of a transparency mask image of a certain scale;

[0027] The reconstructed feature and the first hidden state information of the video are fused and the first hidden state information is updated based on the first image feature and the obtained feature of the contour mask image to obtain a fusion result, including: performing a preset number of times of transparency information fusion in the following manner: based on the first target feature and a second target feature in the obtained feature of the contour mask image, feature reconstruction is performed to obtain a second image feature with an increased scale, wherein the first target feature is the first image feature when the information fusion is performed for the first time, and the first target feature is the feature obtained by performing the information fusion last time when the information fusion is performed for other times, and the scale of the first target feature is the same as the scale of the second target feature; the second image feature and a first sub-state information in the first hidden state information are fused and the first sub-state information is updated to obtain a third image feature, wherein the scale of the transparency mask image corresponding to the fusion feature represented by the first sub-state information is the same as the scale of the second image feature.

[0028] In some embodiments of the first aspect, based on the first target feature and a second target feature in the obtained feature of the contour mask image, feature reconstruction is performed to obtain a second image feature with an increased scale, including: from the second target feature in the obtained feature of the contour mask image, a representative feature of the object edge is screened; based on the first target feature and the representative feature, feature reconstruction is performed to obtain the second image feature with an increased scale.

[0029] In some embodiments of the first aspect, the first image feature comprises a plurality of first sub-image features. The information compression of the target video frame in the video is performed to obtain the first image feature, comprising: performing cascaded information compression of the target video frame in the video to obtain each first sub-image feature with a scale successively decreasing; and the third target feature is the first sub-image feature with the smallest scale when performing the feature reconstruction for the first time. The feature reconstruction is performed based on the third target feature to obtain a fourth image feature with a scale successively increasing, comprising: performing the feature reconstruction based on the third target feature and the first sub-image feature with the same scale as the third target feature to obtain the fourth image feature with the scale successively increasing when performing the feature reconstruction for other times.

[0030] In some embodiments of the first aspect, the first image feature comprises a plurality of first sub-image features. The information compression of the target video frame in the video is performed to obtain the first image feature, comprising: performing cascaded information compression of the target video frame in the video to obtain each first sub-image feature with a scale successively decreasing; and the third target feature is the first sub-image feature with the smallest scale when performing the feature reconstruction for the first time. The feature reconstruction is performed based on the third target feature to obtain a fourth image feature with a scale successively increasing, comprising: performing the feature reconstruction based on the third target feature and the first sub-image feature with the same scale as the third target feature to obtain the fourth image feature with the scale successively increasing when performing the feature reconstruction for other times.

[0031] In some embodiments of the first aspect, the fusion of the second image feature and the first sub-state information in the first hidden state information and the update of the first sub-state information are performed to obtain the third image feature, comprising: segmenting the second image feature to obtain a second sub-image feature and a third sub-image feature; fusing the second sub-image feature and the first sub-state information in the first hidden state information and updating the first sub-state information to obtain a fourth sub-image feature; and splicing the fourth sub-image feature and the third sub-image feature to obtain the third image feature.

[0032] Fusing the fourth image feature and the second sub-state information in the second hidden state information and updating the second sub-state information to obtain a feature of the object's contour mask image, comprising: segmenting the fourth image feature to obtain a fifth sub-image feature and a sixth sub-image feature; fusing the fifth sub-image feature and the second sub-state information in the second hidden state information and updating the second sub-state information to obtain a seventh sub-image feature; and splicing the seventh sub-image feature and the sixth sub-image feature to obtain the feature of the object's contour mask image.

[0033] In combination with the first aspect, in some embodiments, the target video frame in the video is information compressed to obtain a first feature, comprising: inputting the target video frame in the video into a pre-trained video matting model information compression network to obtain a first image feature output by the information compression network, wherein the video matting model further comprises: a first image generation network, a second image generation network, a result output network, a plurality of groups of the same number of contour feature generation networks, and a plurality of groups of the same number of transparency feature generation networks, each group of contour feature generation networks corresponds to a scale of one contour mask image, and comprises a first reconstruction sub-network and a first fusion sub-network, and each group of transparency feature generation networks corresponds to a scale of one transparency mask image, and comprises a second reconstruction sub-network and a second fusion sub-network;

[0034] Based on the first target feature and the second target feature in the obtained feature of the contour mask image, a feature reconstruction is performed to obtain a second image feature with an increased scale, comprising: inputting the first target feature and the second target feature in the obtained feature of the contour mask image into a target second reconstruction sub-network in a target transparency feature generation network to obtain a second image feature with an increased scale output by the target second reconstruction sub-network, wherein the target transparency feature generation network corresponds to a transparency mask image with the same scale as the second image feature;

[0035] Fusing the second image feature and the first sub-state information in the first hidden state information and updating the first sub-state information to obtain a third image feature, comprising: inputting the second image feature into a target second fusion sub-network in the target transparency feature generation network to enable the target second fusion sub-network to fuse the second image feature and the first sub-state information provided by itself and update the first sub-state information to obtain a third image feature output by the target second fusion sub-network;

[0036] Based on the fusion result, a target transparency mask image of an object edge in the target video frame is obtained, comprising: inputting the fusion result into the second image generation network to obtain a target transparency mask image of the object in the target video frame output by the second image generation network;

[0037] The feature reconstruction is performed based on the third target feature, and fourth image features with increased scales are obtained, including: inputting the third target feature into a target first reconstruction subnetwork in a target contour feature generation network to obtain fourth image features with increased scales output by the target first reconstruction subnetwork, wherein a scale of a contour mask image corresponding to the target contour feature generation network is the same as a scale of the fourth image features;

[0038] The fourth image features and second sub-state information in the second hidden state information are fused and the second sub-state information is updated to obtain features of the contour mask image of the object, including: inputting the fourth image features into a target first fusion subnetwork in the target contour feature generation network, so that the target first fusion subnetwork fuses the fourth image features and the second sub-state information provided by itself and updates the second sub-state information to obtain features of the contour mask image of the object output by the target first fusion subnetwork;

[0039] The features obtained through the last processing are used to obtain a first contour mask image of the object in the target video frame, including: inputting the features obtained through the last processing into a first image generation network to obtain a first contour mask image of the object in the target video frame output by the first image generation network;

[0040] The target video frame is regionally matting based on the target transparency mask image and the first contour mask image to obtain a matting result, including: inputting the target transparency mask image and the first contour mask image into a result output network, so that the result output network performs regional matting on the target video frame based on the obtained images to obtain a matting result output by the result output network.

[0041] In combination with the first aspect, in some embodiments, the transparency feature generation network further includes a feature screening subnetwork;

[0042] Before inputting the first target feature and the second target feature in the features of the obtained contour mask image into a target second reconstruction subnetwork in the target transparency feature generation network, the method further includes: inputting the second target feature in the features of the obtained contour mask image into a target feature screening subnetwork in the target transparency feature generation network to obtain target screening features in the second target feature that are characteristic of the edge contour of the object output by the target feature screening subnetwork;

[0043] Inputting the first target feature and the second target feature in the features of the obtained contour mask image into the target second reconstruction subnetwork in the target transparency feature generation network includes: inputting the first target feature and the target screening features into the target second reconstruction subnetwork in the target transparency feature generation network.

[0044] In some embodiments, the first fusion subnetwork is a Gated Recurrent Unit (GRU) or a Long Short-Term Memory (LSTM) unit; and / or, the second fusion subnetwork is a Gated Recurrent Unit (GRU) or a Long Short-Term Memory (LSTM) unit; and / or, the first reconstruction subnetwork is implemented based on a QARepVGG network structure, or the first reconstruction subnetwork in a specific contour feature generation network is implemented based on a QARepVGG network structure, where the specific contour feature generation network is a contour feature generation network with a scale smaller than a first preset scale for the corresponding contour mask image; and / or, the second reconstruction subnetwork is implemented based on a QARepVGG network structure, or the second reconstruction subnetwork in a specific transparency feature generation network is implemented based on a QARepVGG network structure, where the specific transparency feature generation network is a transparency feature generation network with a scale smaller than a second preset scale for the corresponding transparency mask image.

[0045] In some embodiments, the video matting model is trained in the following manner.

[0046] The first sample video frame in the sample video is input into an initial model of the video matting model for processing to obtain a first sample contour mask image of an object in the first sample video frame output by a first image generation network in the initial model;

[0047] A first difference between a labeled mask image corresponding to the first sample video frame and a labeled mask image corresponding to a second sample video frame is obtained, where the second sample video frame is a video frame in the sample video that is before the first sample video frame and is separated by a preset number of frames;

[0048] A second difference between the first sample contour mask image and a second sample contour mask image is obtained, where the second sample contour mask image is a mask image output by the first image generation network when the initial model processes the second sample video frame;

[0049] A third difference between a first sample transparency mask image and a second sample transparency mask image is obtained, where the first sample transparency mask image is a mask image output by the second image generation network when the initial model processes the first sample video frame, and the second sample transparency mask image is a mask image output by the second image generation network when the initial model processes the second sample video frame;

[0050] The training loss is calculated based on the first difference, the second difference, and the third difference;

[0051] Based on the training loss, model parameter adjustment is performed on the initial model to obtain the video matting model.

[0052] In some embodiments of the first aspect, the first sample contour mask image comprises a first mask sub-image identifying a region in the first sample video frame where the object is located and a second mask sub-image identifying a region in the first sample video frame where the object is not located; the second sample contour mask image comprises a third mask sub-image identifying a region in the second sample video frame where the object is located and a fourth mask sub-image identifying a region in the second sample video frame where the object is not located; and obtaining the second difference between the first sample contour mask image and the second sample contour mask image comprises obtaining a difference between the first mask sub-image and the third mask sub-image and obtaining a difference between the second mask sub-image and the fourth mask sub-image, to obtain the second difference comprising the obtained differences.

[0053] In some embodiments of the first aspect, the information compression is performed on the target video frame in the video to obtain the first feature, comprising: performing convolutional transformation on the target video frame in the video to obtain a fifth image feature; performing linear transformation on the fifth image feature based on a convolution kernel to obtain a sixth image feature; performing batch normalization processing on the sixth image feature to obtain a seventh image feature; performing nonlinear transformation on the seventh image feature to obtain an eighth image feature; and performing linear transformation on the eighth image feature based on a convolution kernel to obtain the first feature of the target video frame.

[0054] In some embodiments of the first aspect, the convolution kernel is a 1x1 convolution kernel; and / or the nonlinear transformation on the seventh image feature to obtain the eighth image feature comprises nonlinear transformation on the seventh image feature based on a RELU activation function to obtain the eighth image feature.

[0055] In a second aspect, the present application provides an electronic device, comprising a processor and a memory; wherein the memory is coupled to the processor, and the memory is configured to store computer program codes, the computer program codes comprising computer instructions, which, when executed by the processor, cause the electronic device to perform the method described in the first aspect and any possible implementation manner of the first aspect.

[0056] In a third aspect, the present application provides a chip system applied to an electronic device, the chip system comprising one or more processors configured to invoke computer instructions to cause the electronic device to perform the method described in the first aspect and any possible implementation manner of the first aspect.

[0057] In a fourth aspect, the present application provides a computer readable storage medium comprising computer executable programs, which, when executed on an electronic device, cause the electronic device to perform the method described in the first aspect and any possible implementation manner of the first aspect. BRIEF DESCRIPTION OF DRAWINGS

[0058] FIG. 1 shows a setting interface of screen-off display provided by an embodiment of the present application;

[0059] FIG. 2 shows the overall flow of a wallpaper display method provided by an embodiment of the present application;

[0060] FIG. 3A exemplarily shows a wallpaper provided by an embodiment of the present application;

[0061] FIG. 3B shows a wallpaper provided by an embodiment of the present application generated based on a video clip;

[0062] FIG. 4A shows a way of sequentially arranging video frames adopted by AOD wallpaper, lock screen wallpaper and desktop wallpaper respectively;

[0063] FIG. 4B shows another way of sequentially arranging video frames adopted by AOD wallpaper, lock screen wallpaper and desktop wallpaper respectively;

[0064] FIG. 5 shows a method of fusing wallpaper and user interface provided by an embodiment of the present application;

[0065] FIG. 6 shows a method flow of editing wallpaper background (e.g. changing background) provided by an embodiment of the present application;

[0066] FIG. 7 shows the display of new wallpaper after background change at AOD, lock screen or desktop interface, etc;

[0067] FIG. 8 shows a method flow of editing foreground (e.g. adding decorations) provided by an embodiment of the present application;

[0068] FIG. 9 shows the display of new wallpaper with decorations added to the subject at AOD, lock screen or desktop interface, etc;

[0069] FIGS. 10A-10H exemplarily show human-computer interaction for user to select video clip to generate wallpaper;

[0070] FIGS. 11A-11D exemplarily show human-computer interaction for user to edit subject of wallpaper;

[0071] FIGS. 12A-12E exemplarily show human-computer interaction for user to edit background of wallpaper;

[0072] FIGS. 13A-13B exemplarily show human-computer interaction for user to change position of subject;

[0073] FIGS. 14A-14B exemplarily show human-computer interaction for user to change scaling ratio of subject and background of wallpaper;

[0074] FIG. 15 shows a first video cutout method provided by an embodiment of the present application;

[0075] FIG. 16 shows an image change provided by an embodiment of the present application;

[0076] FIG. 17 shows a second video matting method provided by an embodiment of the present application;

[0077] FIG. 18 shows a first feature processing method provided by an embodiment of the present application;

[0078] FIG. 19 shows a second feature processing method provided by an embodiment of the present application;

[0079] FIG. 20 shows a first cascaded feature reconstruction method provided by an embodiment of the present application;

[0080] FIG. 21 shows a second cascaded feature reconstruction method provided by an embodiment of the present application;

[0081] FIG. 22 shows a third video matting method provided by an embodiment of the present application;

[0082] FIG. 23 shows a structure of a first video matting model provided by an embodiment of the present application;

[0083] FIG. 24 shows a structure of an information compression network provided by an embodiment of the present application;

[0084] FIG. 25 shows a structure of a second video matting model provided by an embodiment of the present application;

[0085] FIG. 26 shows a fourth video matting method provided by an embodiment of the present application;

[0086] FIG. 27 shows a structure of a third video matting model provided by an embodiment of the present application;

[0087] FIG. 28 shows a structure of a fourth video matting model provided by an embodiment of the present application;

[0088] FIG. 29 shows a first model training method provided by an embodiment of the present application;

[0089] FIG. 30 shows a second model training method provided by an embodiment of the present application;

[0090] FIG. 31 shows a mask provided by an embodiment of the present application;

[0091] FIG. 32 shows a structure of a second image generation network provided by an embodiment of the present application;

[0092] FIG. 33 shows a comparison of matting results provided by an embodiment of the present application;

[0093] FIG. 34 shows an electronic device provided by an embodiment of the present application;

[0094] FIG. 35 shows a software system architecture applied to an electronic device according to an embodiment of the present application;

[0095] FIG. 36 shows a chip system according to an embodiment of the present application. DETAILED DESCRIPTION

[0096] The terms used in the following embodiments of the present application are only for the purpose of describing particular embodiments and are not intended to be limiting of the present application.

[0097] As described in the background section, an electronic device such as a mobile phone can provide a wallpaper function to support a user to set a wallpaper at AOD, lock screen, desktop, etc. One wallpaper function can support a user to independently set a wallpaper at AOD, lock screen, desktop, etc. respectively, but the AOD wallpaper, lock screen wallpaper, desktop wallpaper, etc. are relatively fragmented and do not provide a coherent display effect of the wallpaper. Another wallpaper function, also known as "super wallpaper" or "from AOD to lock screen to desktop", can provide a coherent and dynamic display effect of the wallpaper from AOD to lock screen to desktop, but the wallpaper transition effect is relatively harsh and the user experience is limited.

[0098] In this document, the AOD interface is an interface displayed by the electronic device when the screen is off, which can be used to display information such as time, date, short message and incoming call reminders, etc. It can save user operations and is very intuitive and convenient. In the embodiments of the present application, the AOD interface can also include an AOD wallpaper. In order to provide the AOD interface, the screen of the electronic device needs to have a pixel-level light-emitting capability to support only lighting part of the pixels of the screen, for example, an organic light-emitting diode (OLED) screen. The lock screen interface is an interface displayed by the electronic device when the screen is on but not unlocked, which can include a system status bar, time, date, weather, lock screen logo, unlock prompt, and icons of shortcut functions (such as flashlight, camera, etc.) that do not require user authorization, etc. The desktop is an interface displayed by the electronic device when the screen is on and unlocked, which can also be referred to as the main interface, which can include desktop icons of various application programs, a system status bar, etc. The interface displayed by the electronic device when the screen is on and unlocked is not limited to the desktop, but can also be other interfaces, such as the user interface of an application program, the user interface displayed in the last unlocked state (before the lock screen). In this document, the interface displayed when the screen is on and unlocked is collectively referred to as the unlocked interface.

[0099] As shown in FIG. 1, the electronic device can provide an AOD function switch in a user interface of a settings application, such as a "Always On Display" switch in an "Always On Display" interface of the settings application. Embodiments of the present application do not limit the interface form of the AOD function switch, whether it can also appear in other application interfaces, and the like. The user can select whether to use the AOD function by operating the switch. When the AOD function switch is turned on, the electronic device can display the AOD interface when the screen is turned off, and otherwise not display the AOD interface. As shown in FIG. 1, when the AOD function switch is turned on, the electronic device can also display "full screen" and "partial" screen-off modes in the "Always On Display" interface. In the "full screen" screen-off mode, the entire screen area displays the AOD interface in a dark tone effect to save screen power consumption. In the "partial" screen-off mode, only a partial screen is lit to display the AOD interface. In some embodiments, in the "full screen" screen-off mode, the AOD wallpaper in the AOD interface can be changed by setting the lock screen wallpaper, i.e., the picture content of the AOD wallpaper is the same as the picture content of the lock screen wallpaper.

[0100] Embodiments of the present application provide a wallpaper display method that can achieve a dynamic wallpaper display effect from a coherent AOD to a lock screen to a desktop and the like, and does not need to additionally design a transition or transition mode from AOD to lock screen, lock screen to desktop, and the like. The transition effect between them is provided by the video, which is more vivid and natural. Moreover, the wallpaper is well integrated with the AOD, lock screen, desktop and the like, and the main object in the wallpaper is not blocked by non-interactive elements such as time in the interface. In addition, the wallpaper can be generated based on user-defined or editing operations, which facilitates the user to set a personalized wallpaper, and the user operation is simple, and one editing operation can take effect on the entire dynamic wallpaper.

[0101] FIG. 2 shows the overall flow of the wallpaper display method provided by embodiments of the present application. As shown in FIG. 2, the method can include the following steps:

[0102] S11, the electronic device in a lock screen state or an unlocked state detects a first operation, and in response to the first operation, the electronic device displays an AOD interface. The AOD interface includes an AOD wallpaper. In addition to the AOD wallpaper, the AOD interface can also display date, time, weather and the like.

[0103] The first operation can be a screen-off operation, for example, an operation of pressing the power key to turn off the screen in the electronic device in a lock screen state or an unlock state. In the lock screen state or the unlock state, the screen of the electronic device is in a lighted state, wherein the lock screen state displays a lock screen interface on the screen, and the unlock state displays a desktop interface or a user interface of an application opened by the user on the screen. Not limited to being triggered by the screen-off operation, if the user does not operate the mobile phone for a long time in the lock screen state or the unlock state, the electronic device can also display the AOD interface.

[0104] S12, when the electronic device displays the AOD interface, the electronic device detects a second operation, such as a user operation of turning on the screen.

[0105] Specifically, the user operation of turning on the screen in the AOD interface can be an operation of tapping the screen, an operation of pressing the power key, and the like, but is not limited thereto. The screen can also be turned on in the AOD interface in a timed manner, and the user can set the time for turning on the screen, and when the time comes, the electronic device turns on the screen. The present application embodiment does not limit the operation of turning on the screen in the AOD interface.

[0106] S13, in response to the second operation, the electronic device displays a lock screen interface, and the lock screen interface includes a lock screen wallpaper.

[0107] S14, the electronic device detects a third operation for unlocking, and the third operation successfully unlocks the screen.

[0108] S15, in response to the third operation, the electronic device displays a desktop interface, and the desktop interface includes a first desktop wallpaper.

[0109] The AOD wallpaper, the lock screen wallpaper, and the desktop wallpaper can respectively adopt a first video frame, a second video frame, and a third video frame in a video segment, and the second video frame is after the first video frame, and the third video frame is after the second video frame, so as to form a coherent and dynamic wallpaper playing picture as the user switches from the AOD interface to the lock screen interface and then to the desktop.

[0110] In the present application embodiment, the AOD interface, the lock screen interface, and the desktop interface can be respectively referred to as a first interface, a second interface, and a third interface; and the video segment (or video segment) used to generate the dynamic wallpaper can be referred to as a first video.

[0111] The AOD wallpaper, the lock screen wallpaper, and the desktop wallpaper can directly adopt the original video frame in the first video, or can use an image frame after image processing such as cropping and filter on the original video frame.

[0112] FIG. 3A illustrates a wallpaper provided by an embodiment of the present application. As shown in FIG. 3A, (a), (b), and (c) respectively show the display of AOD wallpaper 11, lock screen wallpaper 12, and desktop wallpaper 13. AOD wallpaper 11, lock screen wallpaper 12, and desktop wallpaper 13 are sequentially linked to form a dynamic wallpaper. As shown in FIG. 3A, from AOD to lock screen to desktop, AOD wallpaper 11, lock screen wallpaper 12, and desktop wallpaper 13 are linked together, and they are no longer isolated from each other.

[0113] The video clip used to generate the wallpaper can be selected by the user. For example, as shown in FIG. 3B, the user can select video clip 14 to generate the wallpaper, and video clip 14 is from video 15. Video clip 14 can also be selected from a live photo.

[0114] The method of human-computer interaction for the user to select a video clip to generate a wallpaper will be described in detail in the embodiments of FIG. 10A-FIG. 10H, which will not be expanded here.

[0115] Without being limited to AOD to lock screen to desktop, the wallpaper generated by an embodiment of the present application can also be applied to other adjacent interface transition scenarios, such as AOD to desktop to some application interface, which particularly occurs when the user does not set the lock screen, and desktop to application main interface to application secondary interface. Of course, these application interfaces can support wallpaper display, rather than game display interface, video playback interface, and other application interfaces that are not suitable for wallpaper display. For example, the chat interface of “WeChat” can set a chat wallpaper. That is, the electronic device can display multiple user interfaces of different levels in sequence, and display a wallpaper at the multiple user interfaces; the wallpaper is generated based on a video clip selected by the user, and the video frames used by the wallpaper at the multiple user interfaces are sequentially arranged to form a coherent and dynamic wallpaper display effect as a whole when the user performs multi-level interface transition.

[0116] The wallpaper display scheme provided by an embodiment of the present application will be further described below from the aspects of wallpaper linkage, main object highlighting, background editing, and main object editing.

[0117] Wallpaper Linkage

[0118] Since the video frames used by the wallpapers at the AOD interface, the lock screen interface, and the desktop are sequentially arranged, an embodiment of the present application does not need to additionally design the transition mode or transition mode from AOD to lock screen, lock screen to desktop for the wallpaper, and the transition effect between them is self-provided by the video clip, which is more vivid and natural.

[0119] As an example, the video frames used by the wallpapers at the AOD interface, the lock screen interface, and the desktop in the present application can be sequentially arranged in the following ways:

[0120] First implementation

[0121] The AOD wallpaper 11, the lock screen wallpaper 12, and the desktop wallpaper 13 can all be static wallpapers. That is, the video frame adopted by the AOD wallpaper, the video frame adopted by the lock screen wallpaper, and the video frame adopted by the desktop wallpaper are all one frame.

[0122] As shown in FIG. 4A, during the duration (t1 to t2) of the AOD interface, the video frame adopted by the AOD wallpaper 11 is always frame “1”, such as the AOD wallpaper 11 specifically adopting only the face image in frame “1”; during the duration (t2 to t3) of the lock screen interface, the video frame adopted by the lock screen wallpaper 12 is always frame “2”; and during the duration (t3 to t4) of the desktop, the video frame adopted by the desktop wallpaper 13 is always frame “3”. Frame “1”, frame “2”, and frame “3” are several video frames arranged in sequence in a video clip, wherein frame “1” precedes frame “2”, and frame “2” precedes frame “3”.

[0123] The frame interval between the first video frame and the second video frame can be greater than the first frame interval, and the frame interval between the second video frame and the third video frame can be greater than the second frame interval. For example, the first video frame, the second video frame, and the third video frame can be extracted from the first video at a certain frame interval (such as 1 second or more). In this way, the pictures of the multiple video frames can be different, and a dynamic change effect can be presented when they are played in sequence.

[0124] With the first implementation manner, although the wallpaper seen by the user when the user stays at a certain place such as the AOD, the lock screen, or the desktop is static, the AOD wallpaper 11, the lock screen wallpaper 12, and the desktop wallpaper 13 are sequentially linked and displayed when the user switches from the AOD to the lock screen and then to the desktop, which can generate a coherent dynamic effect as a whole.

[0125] Moreover, when the user switches the interface in reverse, that is, from the desktop to the lock screen and then to the AOD interface, a coherent dynamic effect can also be generated as a whole in the embodiments of the present application.

[0126] Second implementation manner

[0127] The AOD wallpaper 11, the lock screen wallpaper 12, and the desktop wallpaper 13 can all be static wallpapers. That is, the video frame adopted by the AOD wallpaper, the video frame adopted by the lock screen wallpaper, and the video frame adopted by the desktop wallpaper are all one frame.

[0128] When the user stays at a certain place such as AOD, lock screen or desktop, the electronic device plays the dynamic wallpaper at the place, and the wallpaper seen by the user is dynamic. To achieve the coherent dynamic effect from AOD to lock screen to desktop as a whole, the video frames adopted by the AOD wallpaper 11, the lock screen wallpaper 12 and the desktop wallpaper 13 need to be connected at the beginning and the end according to the interface switching time. That is, when switching from AOD to the lock screen interface, the starting frame adopted by the lock screen wallpaper 12 can be the next frame of the ending frame adopted by the AOD wallpaper 11; when switching from the lock screen interface to the desktop interface, the starting frame adopted by the desktop wallpaper 13 can be the next frame of the ending frame adopted by the lock screen wallpaper 12. In other words, among two adjacent user interfaces, the starting video frame adopted by the wallpaper at the next level user interface can be the next frame of the ending video frame adopted by the wallpaper at the previous level user interface.

[0129] For example, it is assumed that the AOD wallpaper 11, the lock screen wallpaper 12 and the desktop wallpaper 13 are all generated by frame “1”, frame “2” and frame “3”. As shown in FIG. 4B, several stages and time of a round of wallpaper display from AOD to lock screen to desktop can be as follows:

[0130] t1 to t2: play the AOD wallpaper 11. As shown in FIG. 4B, if the duration of t1 to t2 is long, the user stays at the AOD interface for a long time and does not press the power key to light the screen for a long time, the AOD wallpaper 11 can be played in a loop.

[0131] t2: it is detected that the user switches from the AOD interface to the lock screen interface, the playing of the AOD wallpaper 11 is ended, and the playing of the lock screen wallpaper 12 is started. The starting frame (such as frame “3”) adopted by the lock screen wallpaper 12 can be the next frame of the ending frame (such as frame “2”) adopted by the AOD wallpaper 11.

[0132] t2 to t3: play the lock screen wallpaper 12. As shown in FIG. 4B, if the duration of t2 to t3 is short, the user quickly unlocks to enter the desktop, and the lock screen wallpaper 12 is not played in a loop; otherwise, the lock screen wallpaper 12 is played in a loop.

[0133] t3: it is detected that the user switches from the lock screen interface to the desktop, the playing of the lock screen wallpaper 12 is ended, and the playing of the desktop wallpaper 13 is started. The starting frame (such as frame “3”) adopted by the desktop wallpaper 13 is the next frame of the ending frame (such as frame “2”) adopted by the lock screen wallpaper 12.

[0134] t3 to t4: play the desktop wallpaper 13. As shown in FIG. 4B, if the duration of t1 to t2 is long, the user stays at the desktop for a long time and does not press the power key to turn off the screen for a long time, the AOD wallpaper 11 can be played in a loop.

[0135] t4: the user is detected to switch from the desktop to the AOD interface, the playing of the desktop wallpaper 13 is ended, the playing of the AOD wallpaper 11 is started, and from this, the next round of wallpaper display is entered. The starting frame (such as frame "1") of the next round of AOD wallpaper 11 is the next frame of the ending frame (such as frame "3") of the previous round of desktop wallpaper 13. Of course, the user can also switch from other application interfaces (such as a video playing interface) to the AOD interface, at which time the starting frame of the next round of AOD wallpaper 11 can be the first frame in the video clip by default, that is, frame "1".

[0136] Moreover, in the embodiments of the present application, when the user switches the interface in reverse, that is, switches from the desktop to the lock screen and then to the AOD interface, a coherent dynamic effect is also generated as a whole. Specifically, the starting and ending frames can be connected according to the interface switching time. That is: when switching from the desktop to the lock screen, the starting frame of the lock screen wallpaper can be the next frame of the ending frame of the desktop wallpaper; when switching from the lock screen to the AOD, the starting frame of the AOD wallpaper can be the next frame of the ending frame of the lock screen wallpaper.

[0137] As shown in FIG. 4B, the frame rates of the AOD wallpaper 11, the lock screen wallpaper 12 and the desktop wallpaper 13 can be the same, such as changing the video frame 2 times per 1 second. However, they can also be different. For example, the frame rate of the AOD dynamic wallpaper can be larger, such as changing the video frame 5 times per 1 second, to reduce the risk of damage to the fixed area of the screen due to long-time lighting. Originally, the AOD wallpaper 11 is dynamic compared to the AOD wallpaper 11 being static has already reduced this risk. In FIG. 4A and FIG. 4B, "T" represents the screen refresh rate. FIG. 4A and FIG. 4B are only examples, and the wallpaper frame rate can actually be greater than T, such as changing the video frame of the wallpaper once every 10 T; moreover, the duration of the AOD can be much greater than several T; based on the human reaction delay and operation delay, the duration of the lock screen wallpaper and the desktop wallpaper can also be much greater than several T.

[0138] Main object highlighting

[0139] In the embodiments of the present application, the wallpaper can be well integrated with the AOD, lock screen, desktop and other interfaces, and the main object (such as a person) in the wallpaper will not be blocked by the time and other non-interactive elements in the AOD, lock screen, desktop and other interfaces.

[0140] FIG. 5 shows a method for fusing wallpaper and user interface according to an embodiment of the present application. As shown in FIG. 5, the electronic device can first perform matting for the subject in a video segment used to generate the wallpaper, separate the subject and the background thereof, and obtain a foreground layer and a background layer. The foreground layer can be a layer of the subject (e.g., a person), and the background layer can be a layer of image content other than the subject. Then, the electronic device can insert a layer of non-interactive elements such as time between the foreground layer and the background layer, synthesize these layers, and render an AOD, lock screen, or desktop interface. In this way, the subject image is displayed above the non-interactive elements such as time, is not blocked by the non-interactive elements such as time, and can better highlight the appearance of the subject in the wallpaper.

[0141] The layer of interactive elements in the AOD, lock screen, or desktop interface, such as the layer of fingerprint unlocking keys in the AOD interface, the layer of password keyboard in the lock screen interface, and the layer of application icons in the desktop, can be arranged above the foreground layer, so as to facilitate user operation of the interactive elements and ensure implementation of the interface interaction function. That is, the layer of the AOD, lock screen, or desktop interface can be divided into the layer of non-interactive elements and the layer of interactive elements.

[0142] How to perform matting for the subject in a video segment will be described in subsequent embodiments.

[0143] In addition to the subject not being blocked by the non-interactive elements such as time, the electronic device can further provide a background editing function and a foreground editing function based on the matting.

[0144] Background editing

[0145] FIG. 6 shows a method for editing the background of the wallpaper (e.g., changing the background) according to an embodiment of the present application. As shown in FIG. 6, after obtaining the foreground layer and the background layer of the wallpaper through matting, the electronic device can replace the background layer, then insert a layer of non-interactive elements such as time between the foreground layer and the new background layer, synthesize these layers, and render an AOD, lock screen, or desktop interface for changing the background of the wallpaper. Thereafter, as shown in FIG. 7, the wallpaper displayed at the AOD, lock screen, or desktop interface is changed in the background.

[0146] Further, in order to simplify the editing operation and improve the background editing efficiency, after detecting that the user edits the wallpaper at a certain user interface (such as the lock screen interface), the electronic device can apply the editing to the wallpaper at other user interfaces, that is, the same editing operation is performed on the wallpaper background at other user interfaces, so that the editing operation is effective on the entire wallpaper. Here, the other user interfaces refer to the user interfaces (such as the lock screen interface) other than the user interface receiving the editing operation in the multi-level user interfaces (such as the AOD interface, the lock screen interface, and the desktop) capable of displaying the wallpaper. In other words, the editing operation performed by the user can be the operation of selecting the wallpaper background at a certain user interface in the multi-level user interfaces for background editing; in addition to updating the wallpaper background at the certain user interface, the electronic device can also update the background image at other user interfaces. That is, in response to the operation of the user editing the wallpaper background at a certain user interface, the electronic device can update the wallpaper background image at the multi-level user interfaces.

[0147] In addition to replacing the background, editing the background can also include changing the zoom ratio of the background, blurring the background, softening the background, changing the background color, brightness, and the like. The electronic device can perform corresponding image processing on the background layer alone, such as image zooming processing, image blurring processing, image softening processing, adjusting image brightness, changing image color, and the like, to update the background layer, and then obtain a new wallpaper and apply the new wallpaper at the AOD interface, the lock screen interface, and the desktop. Here, the specific meaning of application can be that the background layer, the foreground layer of the new wallpaper, and the non-interactive element layer and the interactive element layer of the AOD interface, the lock screen interface, and the desktop are image-composed, and a picture of the new wallpaper displayed at these interfaces is rendered.

[0148] The following embodiments of FIGS. 12A-12E will exemplarily introduce a human-computer interaction method for user editing of the wallpaper background, which will not be expanded here.

[0149] Subject matter editing

[0150] FIG. 8 shows a method flow of editing the foreground (such as adding decorations) provided by an embodiment of the present application. As shown in FIG. 8, after obtaining the foreground layer and the background layer of the wallpaper through matting, the electronic device can add the image of the decoration (such as a hat) to the foreground layer to obtain a new foreground layer, then insert the layer of time and other non-interactive elements between the new foreground layer and the background layer, compose these layers, and render the AOD, lock screen, or desktop interface with the wallpaper subject matter decorated. Thereafter, as shown in FIG. 9, the wallpaper displayed at the AOD, lock screen, or desktop interface realizes the decoration of the subject matter. The specific implementation of adding decorations to the subject matter can be layer composition, such as composing the decoration layer and the foreground layer to obtain a new foreground layer.

[0151] Further, in order to simplify the editing operation and improve the efficiency of the subject editing, after detecting that the user edits the subject of the wallpaper at a certain user interface (such as the lock screen interface), the electronic device can apply the editing to the wallpaper at other user interfaces, that is, the same editing operation is performed on the foreground of the wallpaper at other user interfaces, so as to implement the editing operation on the entire wallpaper. Here, the other user interfaces refer to the user interfaces (such as the AOD interface, the lock screen interface and the desktop) other than the user interface (such as the lock screen interface) receiving the editing operation, which can display the wallpaper. In other words, the editing operation performed by the user can be the operation of selecting the subject of the wallpaper for subject editing at a certain user interface in the multi-level user interface. In addition to updating the subject of the wallpaper at the certain user interface, the electronic device can also update the subject image at other user interfaces. That is, in response to the operation of the user editing the subject of the wallpaper at a certain user interface, the electronic device can update the subject image of the wallpaper at the multi-level user interface.

[0152] In addition to adding decorations, the editing foreground can also include changing the zoom ratio of the subject, changing the position of the subject, changing the color and brightness of the subject, and the like. The electronic device can perform corresponding image processing on the foreground layer, such as image zooming processing, changing the position of the subject image in the foreground layer, adjusting the image brightness, changing the image color, and the like, to update the foreground layer containing the subject image, and then obtain a new wallpaper and apply the new wallpaper at the AOD interface, the lock screen interface and the desktop.

[0153] The following embodiments of FIGS. 11A-11D will exemplarily introduce a human-computer interaction method for user editing of the subject. Here, it is not expanded.

[0154] The electronic device can also edit the wallpaper as a whole, such as adding a filter and changing the zoom ratio of the subject and its background. Editing as a whole means that the editing effect is effective on the entire wallpaper, and is not limited to the foreground or the background. For example, when synthesizing the layers, a filter layer is added above the wallpaper background layer, the non-interactive element layer of the lock screen interface, the wallpaper foreground layer and the interactive element layer of the lock screen interface, so that the filter effect is effective on the entire lock screen wallpaper. For another example, the subject and its background are zoomed according to the selected zoom ratio to obtain a new foreground layer and a background layer, and then the lock screen interface is synthesized with the layers, so that the subject and its background in the lock screen wallpaper are zoomed synchronously.

[0155] Based on the above-mentioned wallpaper display scheme, next, a series of human-computer interaction methods provided by the embodiments of the present application are introduced.

[0156] FIGS. 10A-10H exemplarily show a human-computer interaction method for user selection of a video segment to generate a wallpaper.

[0157] First, the electronic device can display the "desktop and wallpaper" setting interface shown in FIG. 10A, and provide a "wallpaper" setting option in the interface. After detecting that the user selects the "wallpaper" 101 setting option, the electronic device can display the "wallpaper" setting interface shown in FIG. 10B, and provide a "video wallpaper" 102 wallpaper setting option in the interface. In this document, "video wallpaper" is only a name, which refers to a wallpaper setting function of generating wallpaper by using video, and the naming is not limited. After detecting that the user selects the "video wallpaper" 102 option, the electronic device can display the "video wallpaper" setting interface shown in FIG. 10C, and display some recommended videos in the interface. The recommended videos can come from the gallery, or from the network, or from some pre-stored videos locally. In order to enrich the user's selection, the electronic device can also provide a "select from gallery" 103 gallery entry in the "video wallpaper" setting interface shown in FIG. 10C. The gallery entry can be a control (such as a button), a link, or other types of interface interaction elements in technical implementation. After detecting that the user clicks the gallery entry, the electronic device can display the interface shown in FIG. 10D, and display the videos in the gallery in the interface. In order to further facilitate the user to select, the videos in the gallery in the interface shown in FIG. 10D can be classified, such as "person", "building".

[0158] The user can select a video 16 from multiple videos for generating wallpaper. Once the operation of the user selecting the video 16 is detected, the electronic device can display the "video wallpaper preview" interface shown in FIG. 10E.

[0159] As shown in FIG. 10E, the "video wallpaper preview" interface can display a preview 17, which can be used to show the display state of the AOD wallpaper, the lock screen wallpaper, and the desktop wallpaper in the AOD, the lock screen, and the desktop interfaces respectively. The preview can be a video, which can present a dynamic display process of the wallpaper from the AOD to the lock screen to the desktop when the video is played. The video can be composed of multiple image frames, in which, the front image frame can present the display state of the wallpaper in the AOD interface, the middle image frame can present the display state of the wallpaper in the lock screen interface, and the rear image frame can present the display state of the wallpaper in the desktop.

[0160] Further, the preview can also present a reverse dynamic display process of the wallpaper from the desktop to the lock screen to the AOD.

[0161] As shown in FIG. 10E, the "video wallpaper preview" interface can further include: the video frame sequence 18 of the selected video 16, a selection box 19. The video frames in the video frame sequence can be presented in the form of video frame thumbnails. The video segment in the selection box 19 is the selected first video, which can be part or all of the video 16. In embodiments of the present application, the "video wallpaper preview" interface can be referred to as the fourth interface, and the video 16 can be referred to as the second video. The user can slide the selection box 19 along the video frame sequence 18 to change the selected video segment. The length of the selection box 19 can be fixed, such as a fixed selection of a 240-frame-long or 4-second-long video segment. The length of the selected video segment can also be variable. The selected video segment is at least longer than the sum of the frame interval 1 and the frame interval 2, so that three video frames with a large interval can be selected from the video segment for the AOD wallpaper, the lock screen wallpaper, and the desktop wallpaper, respectively, so that the pictures of the multiple video frames are different and present a dynamic change effect when played continuously. The frame interval 1 is the frame interval between the video frames used by the AOD wallpaper and the lock screen wallpaper, respectively, and the frame interval 2 is the frame interval between the video frames used by the lock screen wallpaper and the desktop wallpaper, respectively. Without being limited to the selection box 19, the interface interaction element for selecting a video segment from the video frame sequence 18 can also be of other styles, or no special control needs to be designed for the user to select a video segment, and the user can directly select a plurality of consecutive video frames in the video frame sequence 18 to specify a video segment.

[0162] The above describes that the first video can be selected by the user from the gallery video. Without being limited thereto, the settings application can also provide a video for generating a wallpaper, and the first video can be selected by the user from the videos provided by the settings application. The first video can also be provided by the system without user operation.

[0163] As shown in FIG. 10E, the "video wallpaper preview" interface can further include: a cancel control 20, a complete control 21, a "depth-of-field effect" control, an "edit" control, and the like. The cancel control 20 can be used by the user to cancel the generation of the wallpaper by the video segment in the selection box 19; the complete control 21 can be used by the user to confirm the generation of the wallpaper by the video segment in the selection box 19; the "depth-of-field effect" control can be used by the user to adjust the depth of field of the wallpaper. The "edit" control can be an entry for wallpaper editing. Through the entry, the user can perform a series of wallpaper editing operations, such as adding decorations to the subject, changing the background, and the like. The wallpaper editing will be introduced in the following embodiments of FIGS. 11A-11D and 12A-12E, which will not be expanded here.

[0164] After the electronic device detects that the user changes the video segment for generating the wallpaper, such as changing the selected video segment by sliding the selection box 19, the electronic device can generate a new wallpaper based on the new video segment and update the preview 17. The updated preview 17 can be used to show the display state of the new wallpaper at the AOD, lock screen, desktop, and the like. Specifically, when the updated preview 17 is played, the dynamic display process of the new wallpaper from the AOD to the lock screen and then to the desktop can be shown. In this way, the user can conveniently know the dynamic display effect of the new wallpaper in time, and it is convenient for the user to make a wallpaper that the user likes.

[0165] Not limited to user selection, the video segment for generating the wallpaper can also be default, such as the first three video frames of the default selected video constituting the video segment, or the video segment for generating the wallpaper can be constituted by several video frames before and after the highlight video frame (including the highlight frame itself). As for how to identify the highlight video frame from the video, the embodiments of the present application do not make any limitation.

[0166] After the video segment is changed and the wallpaper is updated, when the AOD, lock screen, and desktop are displayed again, the electronic device will apply the new wallpaper at these interfaces, that is, the updated wallpaper will also be displayed at the AOD, lock screen, and desktop, so as to finally achieve the purpose of changing the wallpaper.

[0167] FIGS. 11A-11D exemplarily show a human-computer interaction method for a user to edit the subject matter of the wallpaper.

[0168] First, the electronic device can display the “wallpaper editing” interface shown in FIG. 11A and display the preview 23 in the interface. In the embodiments of the present application, the “wallpaper editing” interface can be referred to as the fifth interface. Like the aforementioned preview 17, the preview 23 can be used to present the display state of the AOD wallpaper, lock screen wallpaper, and desktop wallpaper at the AOD, lock screen, desktop, and the like, respectively. The preview 23 can be a video, and when the video is played, the dynamic display process of the wallpaper from the AOD to the lock screen and then to the desktop can be presented. In addition, in the preview 23, the subject matter of the wallpaper and its background can be displayed separately, such as indicated by the mark 24. In this way, the user can select the subject matter or the background for editing respectively. The separate display is based on the implementation of the subject matter cutout on the video segment for generating the wallpaper.

[0169] Specifically, the mark 24 can be a highlight boundary line designed along the boundary of the subject matter, or a semi-transparent layer with certain dynamic effects covering the subject matter, and the like. The embodiments of the present application do not make any limitation on the interface manifestation form of the mark 24.

[0170] The user can click the “edit” control in the interface shown in FIG. 10E to enter the “wallpaper editing” interface shown in FIG. 11A. Not limited to this, the entrance of the “wallpaper editing” interface can also be set at other positions, and the embodiments of the present application do not make any limitation thereon.

[0171] As shown in FIGS. 11A-11B, in a state where the subject and the background of the wallpaper are displayed separately, the electronic device can detect an operation of the user selecting the subject, and in response, provide an editing option of “decorate”. The electronic device can further detect an operation of the user clicking the “decorate” editing option, and in response, display one or more decorations, as shown in FIG. 11C, such as “Element 1”, “Element 2”, “Element 3”, and “Element 4” in FIG. 11C.

[0172] Subsequently, the electronic device can detect an operation of the user adding a decoration (such as “Element 1”) to the subject, and in response, generate a new wallpaper and update the preview 23 as shown in FIG. 11D. The subject in the new wallpaper is added with the decoration selected by the user, and the updated preview 23 can be used to show the display state of the new wallpaper at AOD, lock screen, and desktop. Further, the user can also adjust the position of the decoration, such as moving the “hat” to the top of the head of the character.

[0173] Further, to simplify the editing operation and improve the efficiency of subject editing, after detecting an operation of the user adding a decoration to the subject of a wallpaper (such as an AOD wallpaper or a lock screen wallpaper or a desktop wallpaper), the electronic device can apply the editing to other wallpapers, i.e., add the decoration to the subjects of other wallpapers as well, so that one subject editing operation takes effect on the entire dynamic wallpaper.

[0174] After adding a decoration to the subject of a wallpaper, when the AOD, lock screen, and desktop are displayed again, the electronic device will apply the new wallpaper with the subject added with the decoration to these interfaces, i.e., also display the new wallpaper at AOD, lock screen, and desktop, to ultimately achieve the purpose of adding a decoration to the subject of the wallpaper.

[0175] Without being limited to adding a decoration to the subject, selecting the subject for subject editing can also include editing operations such as replacing the subject, adding a filter to the subject, etc.

[0176] FIGS. 12A-12E exemplarily show a human-computer interaction method for a user to edit the background of a wallpaper.

[0177] The electronic device can display a “wallpaper editing” interface as exemplarily shown in FIG. 12A, and the implementation thereof can refer to the related description of the interface shown in FIG. 11A, which will not be repeated here.

[0178] As shown in FIGS. 12A-12B, in a state where the subject of the wallpaper and its background are displayed separately, the electronic device can detect an operation of the user selecting the wallpaper background, in response to which, the editing option of “Switch Background” can be provided. The electronic device can further detect an operation of the user clicking the “Switch Background” editing option, in response to which, as exemplarily shown in FIG. 12C, the background editing options of “Switch Background Color”, “Switch Background Picture”, etc. can be provided, and the background color options of “Color 1”, “Color 2”, “Color 3”, “Color 4”, etc. can be provided, as well as the background picture switching option of “Select from Gallery”.

[0179] Subsequently, the electronic device can detect an operation of the user selecting the option of “Select from Gallery”, and display the gallery calling interface exemplarily shown in FIG. 12C, in which the photos in the gallery are displayed. As shown in FIG. 12D, the electronic device can further detect an operation of the user selecting photo 26 as the new background of the wallpaper, in response to which, a new wallpaper can be generated based on the new background, and the preview 23 can be updated as exemplarily shown in FIG. 12E. The background of the new wallpaper is changed, and the updated preview 23 can be used to display the display state of the new wallpaper at the interfaces of AOD, lock screen, desktop, etc.

[0180] After changing the background of the wallpaper, when the AOD, lock screen, desktop, etc. are displayed again, the electronic device will apply the new wallpaper with the changed background at these interfaces, i.e. the new wallpaper will also be displayed at the AOD, lock screen, desktop, etc. to finally achieve the purpose of changing the background of the wallpaper.

[0181] In addition, the electronic device can detect an operation of the user adjusting the position of the subject relative to the background, such as the leftward dragging of the subject exemplarily shown in FIG. 13A, in response to which, a new foreground layer can be generated based on the subject after the position is updated, and a new wallpaper is generated. As exemplarily shown in FIG. 13B, the position of the subject in the new wallpaper is the adjusted position. After changing the position of the subject in the wallpaper, when the AOD, lock screen, desktop, etc. are displayed again, the electronic device will apply the new wallpaper with the changed position of the subject at these interfaces, i.e. the new wallpaper will also be displayed at the AOD, lock screen, desktop, etc. to finally achieve the purpose of changing the position of the subject in the wallpaper.

[0182] The electronic device can also detect an operation of the user adjusting the zoom scale of the subject matter or the background of the wallpaper, as exemplarily shown in FIG. 14A, dragging the slide bar while the subject matter is selected to change the zoom scale of the subject matter, and as exemplarily shown in FIG. 14B, dragging the slide bar while the background is selected to change the zoom scale of the subject matter. In response thereto, the electronic device can generate a new foreground layer or a new background layer based on the zoom scale selected by the user, and further generate a new wallpaper. Without being limited to the examples shown in FIGS. 14A and 14B, the electronic device can also detect an operation of the user adjusting the zoom scales of the subject matter and the background at the same time, such as zooming in the subject matter while zooming out the background to form a movie visual effect, and in response thereto, the foreground layer and the background layer can be updated at the same time, and a new wallpaper can be generated based on the new foreground layer and the new background layer. After the zoom scales of the subject matter and the background in the wallpaper are changed, when the AOD, the lock screen, or the desktop is displayed again, the electronic device will apply the new wallpaper with the changed zoom scales of the subject matter and the background at these interfaces, that is, the new wallpaper will also be displayed at the AOD, the lock screen, or the desktop, so as to ultimately achieve the purpose of changing the zoom scales of the subject matter and the background in the wallpaper.

[0183] The embodiments of the present application also provide a video matting method to complete matting for the subject matter for a video segment used to generate a wallpaper, separate the subject matter and the background thereof, and then insert a layer of time and other non-interactive elements between the foreground layer and the background layer, so as to avoid the subject matter image from being blocked by the time and other non-interactive elements, and better highlight the appearance of the subject matter in the wallpaper.

[0184] The video matting method provided by the embodiments of the present application can also improve the accuracy of video matting.

[0185] Firstly, the video matting process is described.

[0186] The video contains multiple video frames, and a video frame to be subjected to matting is referred to as a target video frame. In this way, the target video frame can be any video frame in the video that needs to be subjected to matting. Video matting refers to extracting an object region in which an object is located and has object transparency from a video frame. The object can be a person, an animal, a vehicle, a lane line, and the like.

[0187] In the video matting process, firstly, a first target video frame is determined. The first target video frame can be the first video frame or other video frame in the video. The target video frame determined is subjected to matting to obtain an object region in the target video frame. Then, a next video frame of the target video frame is determined as a new target video frame, and the new target video frame is subjected to matting. In this way, after the object region in the target video frame is obtained each time, the next video frame is determined as the new target video frame, until the segmentation result of the last video frame in the video is obtained, and then the matting for the object for the entire video is completed.

[0188] Next, the application scenario of the video matting scheme provided by the embodiments of the present application is exemplified.

[0189] 1. Real-time video scene

[0190] In this scenario, video matting is performed on the video to be played to obtain the object region in each video frame in the video. In this way, when the video is played, only the region content of the object region in each video frame in the video can be played.

[0191] 2. Video editing scene

[0192] In this scenario, after the video is segmented to obtain the region where the object is located in each video frame in the video, the video frame in the video can be edited by replacing the background, erasing the object, blurring the background, preserving the color, and the like according to the position of the object region in the video frame, the picture content and the like, so as to obtain a new video. In addition, after the video frame in the video is edited to obtain a new video, other applications such as video creation, terminal lock screen and the like can be realized based on the new video.

[0193] In the scenario of applying the new video obtained by video editing to the terminal lock screen, after the video is segmented to obtain the region where the object is located in each video frame in the video, and the video is edited according to the information of the object region in the video frame to obtain a new video, the dynamic lock screen wallpaper of the terminal can be generated according to the picture content of each video frame in the new video, so as to display the dynamic lock screen wallpaper when the terminal is in the lock screen state.

[0194] 3. Video monitoring scene

[0195] In this scenario, after the monitoring device collects the video of a specific region, the object in the specific region can be detected by performing video matting on the video.

[0196] Next, the video segmentation scheme provided by the embodiments of the present application is described in detail through specific embodiments.

[0197] In one embodiment of the present application, referring to FIG. 15, a flowchart of a first video matting method is provided. In this embodiment, the above method includes the following steps S301-S305.

[0198] Step S301: Compress the information of the target video frame in the video to obtain the first image feature.

[0199] The target video frame can be any video frame in the video frames contained in the video.

[0200] The information compression on the target video frame can be understood as: feature extraction is performed on the target video frame to obtain a first image feature with a scale smaller than that of the target video frame. The feature extraction on the target video frame can extract edge information of the content in the image, and the edge information can reflect the region where the object in the video frame is located.

[0201] In addition, when the feature extraction is performed on the target video frame, multiple feature extractions in cascade can be performed, and as the number of feature extractions increases, the scale of the obtained feature becomes smaller and smaller. In terms of scale, the larger the scale of the first image feature is, the more detailed edge information it contains. Too much detailed edge information is not conducive to determining the region where the object in the video frame is located in some cases. Conversely, the smaller the scale of the first feature is, the more macro edge information it contains, which is more conducive to determining the region where the object in the video frame is located.

[0202] Furthermore, the dimension of the first image feature can be the same as that of the target video frame, that is, the target video frame is a two-dimensional image, and thus its dimension is 2. In this case, the first image feature can also be two-dimensional data, and in this case, the first image feature can also be considered as a feature map.

[0203] In an embodiment of the present application, when the information compression is performed on the target video frame, the implementation can be based on an encoding mode, for example, based on an encoding network.

[0204] In another embodiment of the present application, the information compression on the target video frame can be performed by a convolutional transformation on the target video frame. In the process of the convolutional transformation on the target video frame, multiple convolutional transformations can be performed on the target video frame to continuously reduce the scale of the feature obtained by the convolutional transformation.

[0205] In addition, the information compression on the target video frame can also be combined with convolutional transformation, linear transformation, batch normalization processing, and nonlinear transformation, and specific details can be referred to steps S301A-S301E in the subsequent embodiment shown in FIG. 22, which will not be described here.

[0206] Step S302: based on the first image feature, the target video frame is segmented to obtain a feature of a contour mask image of an object obtained in the segmentation process and a first contour mask image of the object in the target video frame.

[0207] The first contour mask image is obtained by segmenting the target video frame, and thus there is a corresponding relationship between the pixel points in the first contour mask image and the pixel points in the target video frame. The first contour mask image can be understood as a binary image indicating the position of the region where the object in the target video frame is located. The region where the object is located indicated by the first contour mask image is determined according to the approximate contour of the object, and thus the region where the object is located can be considered as the approximate region where the object is located.

[0208] For example, the first contour mask image can be a mask image containing pixels with two pixel values of 0 and 1. In the target video frame, the pixel corresponding to the pixel with a pixel value of 1 in the first contour mask image can be a pixel in the region where the object is located, and the pixel corresponding to the pixel with a pixel value of 0 in the first contour mask image can be a pixel in the region other than the region where the object is located.

[0209] Specifically, the target video frame can be segmented in any one of the following three implementation manners.

[0210] In the first implementation manner, based on the first image feature, the target video frame can be segmented by the segmentation manner mentioned in the subsequent embodiment shown in FIG. 18, which will not be described here.

[0211] In the second implementation manner, based on the first image feature, the target video frame can be segmented by using a pre-trained network model, such as the contour feature generation network and the first image generation network in the video segmentation model in the subsequent embodiment shown in FIG. 23, which will not be described here.

[0212] In the third implementation manner, based on the first image feature, the target video frame can be segmented by using an image segmentation algorithm or a video segmentation algorithm.

[0213] Step S303: Based on the first image feature and the feature of the obtained contour mask image, perform feature reconstruction, fuse and update the first hidden state information of the reconstructed feature and the video, and obtain a fusion result.

[0214] The first hidden state information represents the fusion feature of the transparency mask image of the object edge in the video frame before matting.

[0215] The object edge in the video frame can reflect the image content details of the object in the video frame, and the image content details can include the specific position of the region where the object is located and the transparency of the object image content presented by the pixel points in the region where the object is located.

[0216] The transparency mask image of the object edge in the video frame can be understood as a mask image indicating the image content details of the object in the video frame, and the pixel value of the pixel in the transparency mask image can represent the transparency of the object image content presented by the pixel at the same position in the video frame.

[0217] For example, the pixel value range of a pixel in the transparency mask image can be 0-1. If the pixel value of a pixel in the transparency mask image is 0, it means that the pixel with the same position in the video frame belongs to a region other than the region where the object is located, and the image content presented by the pixel with the same position does not contain the image content of the object. If the pixel value of a pixel in the transparency mask image is 0.6, it means that the pixel with the same position in the video frame belongs to the region where the object is located, and the transparency of the object image content presented by the pixel with the same position is 0.6. If the pixel value of a pixel in the transparency mask image is 1, it means that the pixel with the same position in the video frame belongs to the region where the object is located, and the image content presented by the pixel with the same position is all the image content of the object.

[0218] The video frames before the target video frame for which matting is performed mentioned in the present step include at least two video frames, and of course, all video frames before the target video frame for which matting is performed.

[0219] In the case where the target video frame is the first video frame in the video, the target video frame has no video frame before it for which matting is performed. In this case, the first hidden state information can be preset data, for example, preset all-zero data.

[0220] Specifically, the first hidden state information can be represented in the form of a tensor or in the form of a matrix.

[0221] As can be seen from the description of step S301, the first image feature is a feature with a smaller scale relative to the target video frame, and the first image feature can reflect the region where the object is located in the target video frame. In order to subsequently successfully extract the region where the object is located from the target video frame, it is necessary to perform feature mapping on the small-scale first image feature, and the ultimate goal is to map to the target video frame, thereby obtaining the region where the object is located in the target video frame. In view of this, it is necessary to perform up-sampling processing on the first image feature.

[0222] Specifically, the feature reconstruction is performed based on the first image feature, and then a feature with increased scale is reconstructed, and the reconstructed feature and the first hidden state information are fused to obtain a fusion result. Since the first hidden state information represents the fusion feature of the transparency mask image of the object edge in the video frame before the target video frame, that is, the first hidden state information can represent the information of the object edge in the video frame before the target video frame, and the information of the object edge can include the specific position information of the region where the object is located, so after the reconstructed feature and the first hidden state information are fused, the obtained fusion result can not only reflect the region where the object is located in the target video frame, but also adjust the region where the object is located in the target video frame in combination with the region where the object is located in the previous video frame, thereby ensuring the smoothness or time correlation between adjacent video frames.

[0223] Since the first hidden state information needs to be used when performing matting in subsequent video frames, it needs to be updated based on the information of the object in the target video frame. Specifically, the first hidden state information can be updated based on the first image feature, or the first hidden state information can be updated based on the fusion result, for example, the result of fusing the reconstructed feature and the first hidden state information can be used as the new first hidden state information.

[0224] Specifically, when the feature reconstruction is performed based on the first image feature, an upsampling algorithm can be used to transform the first image feature to obtain the reconstructed first image feature; the first image feature can be deconvoluted to obtain the reconstructed first image feature; or the first image feature can be reconstructed based on a decoding network to obtain the reconstructed first image feature, for example, the decoding network can be a decoding part in a U-Net network architecture, or a decoder part in a U2-Net network architecture.

[0225] The reconstructed first image feature and the first hidden state information can be fused in any one of the following two implementation manners.

[0226] In the first implementation manner, the reconstructed first image feature and the first hidden state information can be fused by using a fusion algorithm, a network, or the like to obtain a fusion result.

[0227] For example, the reconstructed first image feature and the first hidden state information can be fused by using a Long Short-Term Memory (LSTM) network, a Gated Recurrent Unit (GRU), or the like to obtain a fusion result.

[0228] In the second implementation manner, the first image feature and the first hidden state information after being reconstructed can be directly superimposed, spliced, or point-multiplied, and a processing result is obtained as the fusion result.

[0229] Other implementation manners of the step S303 can be referred to subsequent embodiments, which are not described here.

[0230] The step S304: obtaining a target transparency mask image of the object edge in the target video frame based on the fusion result.

[0231] The target transparency mask image can be understood as a mask image indicating image content details of the object in the target video frame, and the pixel value range of the pixel point in the target transparency mask image can be 0-1, and the scale is the same as that of the target video frame.

[0232] In an implementation manner of the present application, the fusion result can include the confidence that each pixel point in the target video frame belongs to the object. In this case, after obtaining the fusion result, the target transparency mask image can be obtained according to the confidence of each pixel point included in the fusion result.

[0233] For example, the confidence of each pixel point in the fusion result can be taken as the pixel value of each pixel point at the same position in the mask image to obtain the target transparency mask image.

[0234] For another example, a first threshold and a second threshold can be set in advance, wherein the first threshold is greater than the second threshold. If the confidence of the pixel point included in the fusion result is greater than or equal to the first threshold, it means that the confidence of the pixel point is close to 1, and the confidence that the pixel point belongs to the object is high, and at this time, the pixel value of the pixel point at the same position in the mask image can be determined as 1; if the confidence of the pixel point included in the fusion result is less than or equal to the second threshold, it means that the confidence of the pixel point is close to 0, and the confidence that the pixel point belongs to the object is low, and at this time, the pixel value of the pixel point at the same position in the mask image can be determined as 0; if the confidence of the pixel point included in the fusion result is between the second threshold and the first threshold, the confidence of the pixel point can be mapped to the 0-1 interval according to the mapping relationship between the confidence interval from the second threshold to the first threshold and the 0-1 interval, and the obtained mapped value is the pixel value of the pixel point at the same position in the mask image.

[0235] In another implementation manner of the present application, the target transparency mask image can be obtained through steps S304A-S304C in the embodiment shown in FIG. 17.

[0236] Step S305: performing region matting on the target video frame according to the target transparency mask image and the first contour mask image to obtain a matting result.

[0237] Specifically, the first contour mask image indicates the approximate region where the object is located in the target video frame, and the target transparency mask image indicates the image content details of the object in the target video frame. Thus, according to the target transparency mask image and the first contour mask image, the region matting can be performed on the target video frame in combination with the approximate region where the object is located and the image content details of the object in the target video frame, to obtain the matting result.

[0238] In an implementation manner of the present application, the target transparency mask image and the first contour mask image can be point multiplied according to the positions of the respective pixel points to obtain a point multiplied mask image, and then the obtained mask image and the target video frame can be point multiplied again according to the positions of the respective pixel points to obtain a point multiplication result as the matting result, so as to realize the region matting on the target video frame.

[0239] In addition, referring to FIG. 16, a schematic diagram is shown when multiple video frames are taken as target video frames, from the target video frames to the target transparency mask image and the first contour mask image, and then to the matting result. In FIG. 16, the first row of images are the multiple target video frames; the second row of images are the first contour mask images corresponding to the multiple target video frames; the third row of images are the target transparency mask images corresponding to the multiple target video frames; and the fourth row of images are the matting results corresponding to the multiple target video frames.

[0240] As can be seen from the above, in the scheme provided by the present embodiment, when the target video frame is segmented based on the first image feature, the feature of the contour mask image of the object in the segmentation process is obtained, the feature is characteristic of the contour of the object in the contour mask image, the first hidden state information represents the fusion feature of the transparency mask image of the object edge in the video frame which is matting before the target video frame, so that the feature reconstruction is performed based on the first image feature and the obtained feature of the contour mask image, and the fusion is performed on the reconstructed feature and the first hidden state information. The fusion result not only fuses the contour information of the object in the contour mask image, but also fuses the information of the object edge in the video frame which is matting before the target video frame. Since the video frames in the video often have time domain correlation, when the target transparency mask image is obtained based on the fusion result, the information of the object in the video frame which has time domain correlation is considered on the basis of the target video frame, so as to improve the accuracy of the obtained target transparency mask image. On this basis, according to the target transparency mask image and the first contour mask image, the target video frame can be accurately region matting. It can be seen that the video matting scheme provided by the present embodiment can improve the accuracy of video matting.

[0241] In addition, the fusion feature of the transparency mask image of the object edge in the video frame in which the matting is performed before the target transparency mask image is obtained considering the target video frame is that the image information of the object edge in the video frame is considered, rather than only the image information of the target video frame itself, so that the inter-frame smoothness of the object edge region change in the target transparency mask image corresponding to each video frame in the video is improved, thereby improving the inter-frame smoothness of the object edge region change in the matting result corresponding to each video frame. In the case where the video is a moving object video, since the information of the video frame in which the matting is performed is considered when the matting is performed on the target video frame, the smoothness between the matting result of the target video frame and the matting result of the video frame in which the matting is performed is higher.

[0242] The following describes another implementation of obtaining the target transparency mask image in step S304.

[0243] In one embodiment of the present application, referring to FIG. 17, a flowchart of a second video matting method is provided. In the embodiment, step S304 can be implemented by steps S304A-S304C.

[0244] Step S304A: obtaining a second contour mask image of the object in the target video frame based on the fusion result.

[0245] The second contour mask image can be a binary image, and the size of the second contour mask image is the same as the size of the target video frame.

[0246] In one implementation of the present application, as can be known from the description of step S304, the fusion result can include the confidence that each pixel point in the target video frame belongs to the object. In this case, after the fusion result is obtained, the fusion result can be binarized based on a third threshold value to obtain the second contour mask image.

[0247] When the binarization is performed, the value greater than the third threshold value in the fusion result can be set to 0, and the value not greater than the third threshold value can be set to 1. Of course, the value less than the third threshold value in the fusion result can be set to 0, and the value not less than the third threshold value can be set to 1. The embodiments of the present application do not limit this.

[0248] Step S304B: fusing the second contour mask image and the fusion result to obtain a target fusion feature.

[0249] Specifically, the second contour mask image and the fusion result can be fused by any one of the following two implementation modes.

[0250] In the first implementation, the second contour mask image and the fusion result can be fused by using a fusion algorithm, a network, or the like to obtain a target fusion feature.

[0251] In the second implementation, the second contour mask image and the fusion result can be directly superimposed, spliced, or subjected to point multiplication or the like to obtain a processing result as the target fusion feature.

[0252] Step S304C: obtaining a target transparency mask image of the object edge in the target video frame based on the target fusion feature.

[0253] The target fusion feature is obtained by fusing the second contour mask image and the fusion result. The second contour mask image indicates the approximate location of the object in the target video frame, and the fusion result can include the confidence that each pixel point in the target video frame belongs to the object. Thus, the target fusion feature obtained by fusing the two can include the confidence that each pixel point in the approximate location of the object in the target video frame belongs to the object, so that based on the target fusion feature, the target transparency mask image can be obtained based on the confidence of each pixel point included in the feature.

[0254] Specifically, the implementation of obtaining the target transparency mask image based on the target fusion feature can refer to the implementation of obtaining the target transparency mask image based on the fusion result in step S304 described above, which will not be repeated here.

[0255] As can be seen from the above, in the scheme provided by the embodiment, the second contour mask image of the object is obtained based on the fusion result. The second contour mask image can indicate the approximate location of the object in the target video frame determined by the approximate contour of the object. Thus, the second contour mask image is fused with the fusion result to obtain the target fusion feature. The target fusion feature only needs to focus on the detailed information of the image in the approximate location of the object, so that based on the target fusion feature, the target transparency mask image of the object edge can be accurately obtained by only focusing on the image content in the approximate location of the object. Thus, video matting is performed based on the target transparency mask image, which can improve the accuracy of video matting.

[0256] The first implementation of segmenting the target video frame mentioned in step S302 will be described below.

[0257] In an embodiment of the present application, referring to FIG. 18, a flowchart of the first feature processing method is provided. The feature processing process shown in FIG. 18 includes the segmentation process of step S302 and the reconstruction fusion process of step S303. The number of times of feature reconstruction included in the segmentation process is the same as the number of times of transparency information fusion included in the reconstruction fusion process.

[0258] In FIG. 18, the segmentation process includes twice feature reconstruction, and the reconstruction fusion process includes twice transparency information fusion. In addition, the number of times of feature reconstruction included in the segmentation process and the number of times of transparency information fusion included in the reconstruction fusion process can also be other numbers, such as 3, 4, 5, etc., which are not limited in the embodiment.

[0259] The segmentation process of step S302 and the reconstruction fusion process of step S303 will be described below in combination with FIG. 18.

[0260] Firstly, for the segmentation process of step S302, when the target video frame is segmented based on the first image feature, the cascaded feature reconstruction can be performed based on the first image feature to obtain the features of the contour mask images with the scales of the object increasing in turn, and the first contour mask image of the object in the target video frame is obtained based on the features obtained by the last processing.

[0261] The cascaded feature reconstruction can be understood as multiple times of feature reconstruction, and the result of each time of feature reconstruction is the features of a contour mask image with a scale, and the object of the first time of feature reconstruction is the first image feature, and the object of other times of feature reconstruction is the features obtained by the last time of feature reconstruction.

[0262] In FIG. 18, the cascaded feature reconstruction process includes twice feature reconstruction. In the first time of feature reconstruction, the object of feature reconstruction is the first image feature, and after the feature reconstruction is performed on the first image feature, the contour feature 1 of the contour mask image with the scale increasing can be obtained, and at this time, the first time of feature reconstruction process ends.

[0263] In the second time of feature reconstruction, the object of feature reconstruction is the contour feature 1, and after the feature reconstruction is performed on the contour feature 1, the contour feature 2 of the contour mask image with the scale increasing again can be obtained, and at this time, the second time of feature reconstruction process ends, and the contour feature 2 is the features obtained by the last processing, so that the first contour mask image of the object in the target video frame can be obtained based on the contour feature 2.

[0264] The implementation of each time of feature reconstruction and the implementation of obtaining the first contour mask image based on the features obtained by the last processing will be described below.

[0265] In each time of feature reconstruction, the feature reconstruction can be implemented by any one of the following two implementation manners.

[0266] In the first implementation manner, the feature reconstruction can be implemented in the manner of feature reconstruction, feature fusion and updating mentioned in the subsequent embodiments, which is not described here.

[0267] In the second implementation, the object of each feature reconstruction can be processed by using an up-sampling algorithm, a deconvolution transform, or a feature decoding network.

[0268] After obtaining the feature obtained by the last feature reconstruction processing, the feature can include the confidence of each pixel point in the target video frame belonging to the object. In this case, after obtaining the feature, the feature can be binarized based on a fourth preset threshold to obtain a first contour mask image.

[0269] Secondly, for the reconstruction fusion process of step S303, in the embodiment, the first hidden state information includes a plurality of first sub-hidden state information, and each first sub-hidden state information represents the fusion feature of a transparency mask image of one scale. The plurality of first sub-hidden state information can represent the fusion features of transparency mask images of scales that increase in turn. For example, the first hidden state information can include three first sub-hidden state information, and the three first sub-hidden state information can represent the fusion features of transparency mask images of scales that are 24*24, 28*28, and 32*32 in turn.

[0270] In this case, when feature reconstruction, fusion, and updating are performed based on the first image feature and the obtained feature of the contour mask image, the following method can be used to perform a preset number of times of transparency information fusion, and the feature obtained by the last processing is determined as the fusion result:

[0271] Based on the first target feature and the second target feature in the obtained feature of the contour mask image, feature reconstruction is performed to obtain a second image feature of a scale that increases, and the second image feature and the first sub-state information in the first hidden state information are fused and the first sub-state information is updated to obtain a third image feature.

[0272] The preset number is a preset number and is the same as the number of reconstructions of the cascaded feature reconstruction.

[0273] The first target feature is the first image feature when the information fusion is performed for the first time, and the first target feature is the feature obtained by the last information fusion when the information fusion is performed for other times.

[0274] The scale of the first target feature is the same as the scale of the second target feature.

[0275] The scale of the transparency mask image corresponding to the fusion feature represented by the first sub-state information is the same as the scale of the second image feature.

[0276] In FIG. 18, the preset number is 2, that is, two transparency information fusion processes are performed. In the first transparency information fusion process, the first target feature is the first image feature, the second target feature is the contour feature 1, the first image feature and the contour feature 1 have the same scale, feature reconstruction is performed based on the first target feature and the second target feature, that is, feature reconstruction is performed based on the first image feature and the contour feature 1, the second image feature 1 with increased scale is obtained, the second image feature 1 and the first sub-state information 1 corresponding to the second image feature 1 are fused and the first sub-state information 1 is updated, and the third image feature 1 is obtained. At this time, the first transparency information fusion process ends.

[0277] After the third image feature 1 in the first transparency information fusion process is obtained, the second transparency information fusion is performed. In the second transparency information fusion process, the first target feature is the third image feature 1, the second target feature is the contour feature 2, the third image feature 1 and the contour feature 2 have the same scale, feature reconstruction is performed based on the first target feature and the second target feature, that is, feature reconstruction is performed based on the third image feature 1 and the contour feature 2, the second image feature 2 with further increased scale is obtained, the second image feature 2 and the first sub-state information 2 corresponding to the second image feature 2 are fused and the first sub-state information 2 is updated, and the third image feature 2 is obtained. At this time, the second transparency information fusion process ends, and the third image feature 2 is the final obtained fusion result.

[0278] The implementation manner of each time of performing feature reconstruction based on the first target feature and the second target feature can refer to the implementation manner of performing feature reconstruction based on the first image feature and the contour mask image in step S303 of the embodiment shown in FIG. 15; and the implementation manner of each time of fusing the second image feature and the first sub-state information and updating the first sub-state information can refer to the implementation manner of fusing the reconstructed feature and the first hidden state information and updating the first hidden state information in step S303 of the embodiment shown in FIG. 15, which will not be described herein again.

[0279] As can be seen from the above, in the scheme provided in this embodiment, the target video frame is segmented in a cascaded feature reconstruction manner based on the first image feature, the cascaded feature reconstruction contains multiple feature reconstruction processes, which can improve the accuracy of the first contour mask image; multiple transparency information fusion processes are performed after the first image feature is obtained, each time of transparency information fusion process contains three processing processes of feature reconstruction, feature and hidden state information fusion, and hidden state information updating, which can improve the accuracy of the finally obtained fusion result, so that the target video frame is regionally matting based on the relatively accurate first contour mask image and the fusion result, the accuracy of the regional matting can be improved, and the accuracy of the video matting is further improved.

[0280] The data amount of the two features required for feature reconstruction in each transparency information fusion process is usually large, so the calculation amount of feature reconstruction is large, and the calculation amount of preset number of times of transparency information fusion is also large.

[0281] In view of this, in an embodiment of the present application, when feature reconstruction is performed based on the first target feature and the second target feature in the feature of the obtained contour mask image, a characteristic feature of the object edge can be screened from the second target feature in the feature of the obtained contour mask image, and feature reconstruction is performed based on the first target feature and the characteristic feature to obtain a second image feature with increased scale.

[0282] The characteristic feature is a feature included in the second target feature and having representation of edge details of the object in the target video frame.

[0283] In the process of feature reconstruction based on the first image feature, the attribute of the object represented by the feature of the reconstructed contour mask image can be determined, so that after the second target feature with the same scale as the first target feature is determined in the feature of each contour mask image in each transparency information fusion process, the feature in the second target feature having representation of the object edge can be determined according to the attribute of the object represented by the second target feature, so that the determined feature is extracted from the second target feature, and the extracted feature is the characteristic feature.

[0284] Specifically, after the feature in the second target feature having representation of the object edge is determined, the characteristic feature can be extracted from the second target feature by using a feature screening algorithm, a feature screening network or an attention mechanism.

[0285] After the characteristic feature is screened, the implementation manner of feature reconstruction based on the first target feature and the characteristic feature can refer to the implementation manner of feature reconstruction based on the first target feature and the second target feature in the foregoing embodiments, which will not be described herein.

[0286] Referring to FIG. 19, FIG. 19 is adjusted on the basis of FIG. 18. The specific adjustment content is that in the first transparency information fusion process, when feature reconstruction is performed based on the first image feature and the contour feature 1, the characteristic feature 1 is first screened from the contour feature 1, and then feature reconstruction is performed based on the first image feature and the characteristic feature 1 to obtain the second image feature 1 with increased scale. In the second transparency information fusion process, when feature reconstruction is performed based on the third image feature 1 and the contour feature 2, the characteristic feature 2 is first screened from the contour feature 2, and then feature reconstruction is performed based on the third image feature 1 and the characteristic feature 2 to obtain the second image feature 2 with further increased scale.

[0287] From the above, in the scheme provided by the embodiment, in the process of feature reconstruction based on the first target feature and the second target feature, the representative feature with small data amount is screened out from the second target feature, so that the feature reconstruction is performed based on the first target feature and the representative feature, the calculation amount of feature reconstruction can be reduced, the calculation amount of the preset number of times of transparency information fusion can be reduced, the efficiency of obtaining the fusion result is improved, and then the efficiency of video matting can be improved.

[0288] The way of feature reconstruction realized by feature reconstruction, feature fusion and updating in the embodiment shown in FIG. 18 is described below.

[0289] In an embodiment of the present application, after the first image feature is obtained, the following way can be used to perform the preset number of times of contour information fusion to obtain the features of the contour mask images of the object with the scales increasing in turn, wherein one contour information fusion processing is equivalent to one feature reconstruction processing:

[0290] The feature reconstruction is performed based on the third target feature to obtain the fourth image feature with the scale increasing, the fourth image feature and the second sub-state information in the second hidden state information are fused and the second sub-state information is updated to obtain the features of the contour mask images of the object.

[0291] Wherein, the third target feature is the first image feature when the information fusion is performed for the first time, and the third target feature is the feature obtained by the last feature reconstruction when the information fusion is performed for other times.

[0292] The second hidden state information represents the fusion features of the contour mask images of the object in the video frames segmented before the target video frame.

[0293] The second hidden state information includes a plurality of second sub-hidden state information, and each second sub-hidden state information represents the fusion features of the contour mask images of one scale. The plurality of second sub-hidden state information can represent the fusion features of the contour mask images with the scales increasing in turn, for example, the second hidden state information can include three second sub-hidden state information, and the three second sub-hidden state information can represent the fusion features of the contour mask images with the scales of 24*24, 28*28 and 32*32 in turn.

[0294] The scale of the contour mask image corresponding to the fusion feature represented by the second sub-state information is the same as the scale of the fourth image feature. For example, if the second hidden state information includes three second sub-hidden state information, which represents the fusion features of contour mask images with scales of 24*24, 28*28, and 32*32 in turn, it indicates that the preset number is three, and three fusions are required. The scales of the fourth image features obtained in the three fusion processes are 24*24, 28*28, and 32*32 in turn.

[0295] The above-mentioned preset number is 3, and the contour information fusion process is described in detail below in conjunction with FIG. 20.

[0296] Referring to FIG. 20, a flowchart of a cascaded feature reconstruction method is provided. In FIG. 20, after the first image feature is obtained, the first contour information fusion is performed. In the first contour information fusion process, the third target feature is the first image feature, and feature reconstruction is performed based on the third target feature, that is, feature reconstruction is performed based on the first image feature to obtain a fourth image feature 1 with an increased scale. The fourth image feature 1 and the second sub-state information 1 corresponding to the fourth image feature 1 are fused and the second sub-state information 1 is updated to obtain a contour feature 3 of the contour mask image of the object. At this time, the first contour information fusion process ends.

[0297] After obtaining the contour feature 3, the second contour information fusion is performed. In the second contour information fusion process, the third target feature is the contour feature 3, and feature reconstruction is performed based on the third target feature, that is, feature reconstruction is performed based on the contour feature 3 to obtain a fourth image feature 2 with an increased scale. The fourth image feature 2 and the second sub-state information 2 corresponding to the fourth image feature 2 are fused and the second sub-state information 2 is updated to obtain a contour feature 4 of the contour mask image of the object. At this time, the second contour information fusion process ends.

[0298] After obtaining the contour feature 4, the third contour information fusion is performed. In the third contour information fusion process, the third target feature is the contour feature 4, and feature reconstruction is performed based on the third target feature, that is, feature reconstruction is performed based on the contour feature 4 to obtain a fourth image feature 3 with an increased scale. The fourth image feature 3 and the second sub-state information 3 corresponding to the fourth image feature 3 are fused and the second sub-state information 3 is updated to obtain a contour feature 5 of the contour mask image of the object. At this time, the second contour information fusion process ends, and the obtained contour feature 5 is the final feature of the contour mask image of the object.

[0299] The implementation manner of the feature reconstruction based on the third target feature in each contour information fusion process can refer to the implementation manner of the feature reconstruction based on the first image feature in step S303 of the embodiment shown in FIG. 15; and the implementation manner of the fusion and update of the second sub-state information based on the fourth image feature and the second sub-state information can refer to the implementation manner of the fusion and update of the first hidden state information based on the reconstructed first image feature and the first hidden state information in step S303 of the embodiment shown in FIG. 15, which will not be described herein again.

[0300] As can be seen from the above, in the scheme provided in the embodiment, the second hidden state information represents the fusion feature of the object contour mask image in the video frame segmented before the target video frame, and each second sub-state information represents the fusion feature of the contour mask image of one scale, so that the fusion of the fourth image feature obtained by the feature reconstruction and the second sub-state information in each contour information fusion process is the fusion of the information in the fusion feature of the contour mask image of one scale of the object in the video frame segmented before the target video frame and the fourth image feature. In this way, the feature of the object contour mask image obtained finally not only contains the information of the object in the target video frame, but also contains the information of the object in the video frame segmented before the target video frame, so that the first contour mask image can be obtained based on the finally obtained feature, the accuracy of the first contour mask image can be improved, and the accuracy of the video matting can be improved.

[0301] When the feature reconstruction is performed in the contour information fusion process, the processing can be based on the feature reconstruction of the third target feature, or the feature reconstruction can be combined with the third target feature and other information.

[0302] In an embodiment of the present application, the first image feature includes a plurality of first sub-image features. When the information compression is performed on the target video frame in the video, the cascade information compression can be performed on the target video frame to obtain the first sub-image features with scales decreasing in turn.

[0303] The cascade information compression can be understood as a plurality of information compressions, and the result of each information compression is a first sub-image feature. The object of the first information compression is the target video frame, and the object of the other information compression is the first sub-image feature obtained by the last information compression.

[0304] The implementation manner of each information compression can refer to the implementation manner of the information compression on the target video frame in step S301 shown in FIG. 15.

[0305] For example, the object of the information compression can be subjected to a plurality of convolution transformations in each information compression.

[0306] For example, each time information compression is performed, the process flow shown in steps S301A-S301E in the embodiment shown in FIG. 22 can be performed one or more times to achieve information compression.

[0307] After obtaining the first sub-image features with gradually reduced scales, the first sub-image features can be used to perform a preset number of times of contour information fusion.

[0308] When feature reconstruction is performed in the first contour information fusion process, the first sub-image feature with the smallest scale among the first sub-image features can be used as a target feature, and feature reconstruction can be performed based on the target feature.

[0309] When feature reconstruction is performed in the other contour information fusion processes, the feature obtained in the last contour information fusion can be used as a third target feature, and feature reconstruction can be performed based on the third target feature and the first sub-image feature with the same scale as the third target feature to obtain a fourth image feature with an increased scale.

[0310] When feature reconstruction is performed based on the third target feature and the first sub-image feature with the same scale as the third target feature, the third target feature and the first sub-image feature with the same scale as the third target feature can be fused into one feature through superposition, point multiplication, or other fusion methods, and feature reconstruction can be performed based on the fused feature.

[0311] The following describes the contour information fusion process with the above-described preset number being 3 as an example and in combination with FIG. 21.

[0312] Referring to FIG. 21, a flowchart of another cascaded feature reconstruction method is provided. In FIG. 21, after obtaining the first sub-image features with gradually reduced scales, the first contour information fusion is performed. In the first contour information fusion process, the third target feature is the first sub-image feature 1 with the smallest scale among the first sub-image features, and feature reconstruction is performed based on the third target feature, that is, feature reconstruction is performed based on the first sub-image feature 1 to obtain a fourth image feature 4 with an increased scale, the fourth image feature 4 and the second sub-state information 4 corresponding to the fourth image feature 4 are fused and the second sub-state information 4 is updated to obtain the contour feature 6 of the contour mask image of the object, and at this time, the first contour information fusion process ends.

[0313] After the contour feature 6 in the first contour information fusion process is obtained, a second contour information fusion is performed. In the second contour information fusion process, the third target feature is the contour feature 6, the first sub-image feature with the same scale as the third target feature is the first sub-image feature 2, feature reconstruction is performed based on the third target feature and the first sub-image feature with the same scale as the third target feature, that is, feature reconstruction is performed based on the contour feature 6 and the first sub-image feature 2, the fourth image feature 5 with a further increased scale is obtained, the fourth image feature 5 and the second sub-state information 5 corresponding to the fourth image feature 5 are fused and the second sub-state information 5 is updated, the contour feature 7 of the object's contour mask image is obtained, and at this time, the second contour information fusion process ends.

[0314] After the contour feature 7 in the second contour information fusion process is obtained, a third contour information fusion is performed. In the third contour information fusion process, the third target feature is the contour feature 7, the first sub-image feature with the same scale as the third target feature is the first sub-image feature 3, feature reconstruction is performed based on the third target feature and the first sub-image feature with the same scale as the third target feature, that is, feature reconstruction is performed based on the contour feature 7 and the first sub-image feature 3, the fourth image feature 6 with a further increased scale is obtained, the fourth image feature 6 and the second sub-state information 6 corresponding to the fourth image feature 6 are fused and the second sub-state information 6 is updated, the contour feature 8 of the object's contour mask image is obtained, and at this time, the third contour information fusion process ends, and the contour feature 8 obtained in the process is the final obtained feature of the object's contour mask image.

[0315] As can be seen from the above, in the scheme provided by the embodiment, the target video frame is subjected to cascade information compression to obtain first sub-image features with scales decreasing in turn, in the subsequent contour information fusion processes except the first time, feature reconstruction can be performed based on the third target feature and the first sub-image feature with the same scale as the third target feature, which can improve the accuracy of feature reconstruction, thereby improving the accuracy of the finally obtained feature after contour information fusion, and further improving the accuracy of video matting.

[0316] According to the above content, it can be known that the number of times of transparency information fusion and the number of times of contour information fusion are both preset numbers, the greater the preset number is, the more accurate the fusion result obtained by transparency information fusion is, and the more accurate the feature obtained by contour information fusion is, however, the greater the calculation amount is.

[0317] Therefore, in one embodiment of the present application, the preset number is 4, 5 or 6. In this way, the accuracy of video matting can be improved, and the calculation amount can be avoided from being too large, so as to ensure that the video matting is implemented with high efficiency, and the calculation resources of the terminal are saved. Therefore, the scheme provided in the embodiments of the present application can be applied to the terminal, and is friendly to the terminal, so that the video matting scheme can be applied to the terminal in a lightweight manner.

[0318] Since the data amount of the second image feature and the fourth image feature is large, the calculation amount of fusing and updating the first sub-state information in the first hidden state information and the second image feature is usually large, and the calculation amount of fusing and updating the second sub-state information in the second hidden state information and the fourth image feature is usually large.

[0319] In view of the above, in order to reduce the calculation amount of fusing and updating the first sub-state information in the first hidden state information and the second image feature, in one embodiment of the present application, when the first sub-state information in the first hidden state information and the second image feature are fused and updated, the second image feature is split to obtain a second sub-image feature and a third sub-image feature; the second sub-image feature and the first sub-state information in the first hidden state information are fused and updated to obtain a fourth sub-image feature; and the fourth sub-image feature and the third sub-image feature are spliced to obtain the third image feature.

[0320] For example, the second image feature can be split in the W dimension direction to obtain two sub-feature tensors with scales of H*C*W1 and H*C*W2, where W1+W2=W.

[0321] For example, for a feature tensor with a scale of H*C*W, the feature tensor can be split in the W dimension direction to obtain two sub-feature tensors with scales of H*C*W1 and H*C*W2, where W1+W2=W.

[0322] Specifically, when the second image feature is split, the second image feature can be split in a proportional manner to obtain two sub-features with the same scale, or the second image feature can be split in an arbitrary proportional manner to obtain two sub-features with different scales. After the second image feature is split to obtain two sub-features, any one of the two sub-features can be determined as the second sub-image feature, and the other sub-feature can be determined as the third sub-image feature.

[0323] After the second sub-image feature and the third sub-image feature are obtained by segmentation, the second sub-image feature and the first sub-state information in the first hidden state information can be fused and the first sub-state information can be updated to obtain a fourth sub-image feature. For details, refer to the implementation manner of fusing and updating the first hidden state information in the implementation example shown in FIG. 15.

[0324] The spliced feature can be regarded as an inverse process of feature segmentation. After the fourth sub-image feature is obtained by fusion, the fourth sub-image feature and the third sub-image feature can be spliced. When splicing the two features, the fourth sub-image feature and the third sub-image feature can be spliced into one feature in the dimension direction according to the segmentation, that is, the third sub-image feature is spliced behind the fourth sub-image feature or the fourth sub-image feature is spliced behind the third sub-image feature in the dimension direction according to the segmentation. The feature obtained by splicing is the third image feature.

[0325] For example, if the scale of the third sub-image feature is H*C*W3 and the scale of the fourth sub-image feature is H*C*W4, when splicing the fourth sub-image feature and the third sub-image feature, the fourth sub-image feature and the third sub-image feature can be spliced into one feature with a scale of H*C*W5 in the W dimension direction, where W3+W4=W5.

[0326] As can be seen from the above, in the scheme provided in the embodiment, the second image feature is segmented to obtain a second sub-image feature and a third sub-image feature. The data amount of the second sub-image feature and the third sub-image feature is smaller than that of the second image feature. Therefore, the second sub-image feature and the first sub-state information are fused, the calculation amount of fusion is reduced, the fusion efficiency is improved, the efficiency of obtaining the third image feature is improved, the efficiency of video matting is improved, the computing resources of the terminal are saved, and the video matting scheme can be applied in the terminal in a lightweight manner.

[0327] To reduce the calculation amount of fusing and updating the second sub-state information in the fourth image feature and the second sub-state information, in an embodiment of the present application, when the fourth image feature and the second sub-state information are fused and the second sub-state information is updated, the fourth image feature is segmented to obtain a fifth sub-image feature and a sixth sub-image feature; the fifth sub-image feature and the second sub-state information in the second hidden state information are fused and the second sub-state information is updated to obtain a seventh sub-image feature; and the seventh sub-image feature and the sixth sub-image feature are spliced to obtain a feature of the contour mask image of the object.

[0328] The implementation manner of segmenting the fourth image feature can refer to the implementation manner of segmenting the second image described above; the implementation manner of fusing the fifth sub-image feature and the second sub-state information and updating the second sub-state information can refer to the implementation manner of fusing the second sub-image feature and the first sub-state information and updating the first sub-state information described above; and the implementation manner of splicing the seventh sub-image feature and the sixth sub-image feature can refer to the implementation manner of splicing the fourth sub-image feature and the third sub-image feature described above, which will not be described herein again.

[0329] As can be seen from the above, in the scheme provided in the embodiment, the fourth image feature is segmented to obtain the fifth sub-image feature and the sixth sub-image feature, and the data amount of the fifth sub-image feature and the sixth sub-image feature is smaller than the data amount of the fourth image feature. In this way, fusing the fifth sub-image feature and the first sub-state information can reduce the calculation amount of fusion and improve the fusion efficiency, so as to improve the efficiency of obtaining the feature of the object's contour mask image, and further improve the efficiency of video matting. At the same time, the calculation resources of the terminal are also saved, so that the video matting scheme can be applied in the terminal in a light-weight manner.

[0330] In addition, to solve the problem of large calculation amount of the two fusion processes, in an embodiment of the present application, when the second image feature and the first sub-state information are fused and the first sub-state information is updated, the second image feature can be segmented to obtain the second sub-image feature and the third sub-image feature; the second sub-image feature and the first sub-state information in the first hidden state information are fused and the first sub-state information is updated to obtain the fourth sub-image feature; and the fourth sub-image feature and the third sub-image feature are spliced to obtain the third image feature. When the fourth image feature and the second sub-state information are fused and the second sub-state information is updated, the fourth image feature is segmented to obtain the fifth sub-image feature and the sixth sub-image feature; the fifth sub-image feature and the second sub-state information in the second hidden state information are fused and the second sub-state information is updated to obtain the seventh sub-image feature; and the seventh sub-image feature and the sixth sub-image feature are spliced to obtain the feature of the object's contour mask image. In this way, the calculation amount of the two fusion processes can be reduced, so as to further improve the fusion efficiency, and further improve the video matting efficiency, and realize the light-weight application of the video matting scheme in the terminal.

[0331] The implementation manner of compressing information of the target video frame by combining convolution transformation, linear transformation, batch normalization processing, and nonlinear transformation in the step S301 will be described below.

[0332] In an embodiment of the present application, referring to FIG. 22, a flowchart of a third video matting method is provided, and the step S301 can be implemented by the following steps S301A-S301E in the embodiment.

[0333] Step S301A: performing convolution transformation on the target video frame in the video to obtain a fifth image feature.

[0334] Specifically, the fifth image feature can be obtained by performing convolution calculation on the target video frame using a pre-set convolution kernel, or the fifth image feature can be obtained by performing convolution transformation on the target video frame using a trained convolution neural network.

[0335] Step S301B: performing linear transformation on the fifth image feature based on the convolution kernel to obtain a sixth image feature.

[0336] The convolution kernel is a pre-set convolution kernel.

[0337] Specifically, the linear transformation on the fifth image feature can be implemented in the form of convolution transformation on the fifth image feature based on the convolution kernel. Since the network processing unit (NPU) in the terminal has strong computing capability for convolution transformation, the linear transformation in the form of convolution transformation can shorten the time consumption of linear transformation, thereby shortening the time consumption of video matting and improving the efficiency of video matting.

[0338] In an embodiment of the present application, the convolution kernel is a 1x1 convolution kernel. Since the 1x1 convolution kernel has a small amount of data itself, linear transformation on the fifth image feature based on the 1x1 convolution kernel can reduce the calculation amount of linear transformation while achieving linear transformation on the fifth image feature, thereby improving the calculation efficiency of linear transformation and the efficiency of video matting. Moreover, the video matting scheme provided in the embodiment is applied to a terminal, and linear transformation on the fifth image feature based on the 1x1 convolution kernel in the terminal does not require a large amount of computing resources of the terminal, thereby facilitating the terminal to implement linear transformation and promoting the lightweight implementation of video matting on the terminal side.

[0339] Step S301C: performing batch normalization processing on the sixth image feature to obtain a seventh image feature.

[0340] Specifically, the batch normalization processing can be performed on the sixth image feature using a batch normalization algorithm or model to obtain the seventh image feature.

[0341] For example, the BatchNorm2d algorithm can be used to perform batch normalization processing on the sixth image feature.

[0342] Step S301D: performing nonlinear transformation on the seventh image feature to obtain an eighth image feature.

[0343] Specifically, the seventh image feature can be subjected to nonlinear transformation by using a nonlinear transformation function, an algorithm, or an activation function, etc., to obtain an eighth image feature.

[0344] For example, the seventh image feature can be subjected to nonlinear transformation by using a GELU activation function or a RELU activation function. In the case of using the RELU activation function to perform nonlinear transformation on the seventh image feature, since the RELU activation function has a good quantization effect on data processing, using the RELU activation function to perform nonlinear transformation on the seventh image feature can improve the transformation effect of nonlinear transformation, thereby improving the accuracy of the eighth image feature.

[0345] Step S301E: performing linear transformation on the eighth image feature based on the convolution kernel to obtain a first feature of the target video frame.

[0346] The implementation manner of linear transformation in this step is the same as that of linear transformation in the above step S301B, which will not be described here.

[0347] In addition, when obtaining the first feature, the processing procedure shown in steps S301A-S301E can be performed once or multiple times. For example, the processing procedure shown in steps S301A-S301E can be performed 4 times, 5 times, or other number of times.

[0348] In the case of performing the processing procedure shown in steps S301A-S301E multiple times, the input of the first processing procedure is the target video frame in the video, the input of the other processing procedures is the feature output by the last processing procedure, the feature output by the last processing procedure is the first feature, and in this case, the scale of the feature output by each processing procedure becomes smaller with the multiple execution of the processing procedure.

[0349] As can be seen from the above, in the scheme provided by the embodiment, when compressing information of the target video frame, the target video frame is subjected to convolution transformation, linear transformation, batch normalization processing, and nonlinear transformation, etc., which can realize more accurate information compression of the target video, thereby improving the accuracy of the first image feature, and further improving the accuracy of video matting based on the first image feature.

[0350] In addition, in the scheme provided by the embodiment, the sixth image feature is subjected to batch normalization processing first, and then the seventh image feature obtained by the processing is subjected to nonlinear transformation, which can prevent the loss of quantization precision of the feature during information compression, thereby improving the quantization precision of information compression, and further improving the accuracy of the first image feature and the accuracy of video matting.

[0351] The scheme provided by the embodiments of the present application is applied to a terminal. Convolution transformation, linear transformation, batch normalization processing, and nonlinear transformation are relatively friendly to the computing capability of the terminal. Therefore, performing convolution transformation, linear transformation, batch normalization processing, and nonlinear transformation in the terminal can facilitate the terminal to compress information, thereby promoting the lightweight implementation of video matting on the terminal side.

[0352] The video matting scheme provided by the embodiments of the present application can also be implemented based on a neural network model. The video matting scheme is described below in combination with the neural network model.

[0353] In an embodiment of the present application, the above steps can be implemented by using a pre-trained video matting model.

[0354] Referring to FIG. 23, a structural schematic diagram of a first video matting model is provided. As can be seen from FIG. 23, the video matting model includes an information compression network, a first image generation network, a second image generation network, a result output network, three sets of contour feature generation networks, and three sets of transparency feature generation networks. Each set of contour feature generation networks corresponds to a scale of a contour mask image and includes a first reconstruction subnetwork and a first fusion subnetwork. Each set of transparency feature generation networks corresponds to a scale of a transparency mask image and includes a second reconstruction subnetwork and a second fusion subnetwork.

[0355] FIG. 23 is a video matting model with three contour feature generation networks. In addition to this, the number of contour feature generation networks included in the video matting model can also be four, five, or other numbers. The number of transparency feature generation networks is the same as that of contour feature generation networks. The present embodiment does not limit this.

[0356] The connection relationship of each network in the video matting model shown in FIG. 23 is described below.

[0357] The three sets of contour feature generation networks in the video matting model are contour feature generation network 1, contour feature generation network 2, and contour feature generation network 3, which correspond to contour mask images with scales that increase in turn. The first reconstruction subnetwork in each set of contour feature generation networks is connected to the first fusion subnetwork. The three sets of transparency feature generation networks are transparency feature generation network 1, transparency feature generation network 2, and transparency feature generation network 3, which correspond to transparency mask images with scales that increase in turn. The second reconstruction subnetwork in each set of transparency feature generation networks is connected to the second fusion subnetwork.

[0358] The video matting model includes a segmentation branch and a matting branch. The segmentation branch includes three groups of contour feature generation networks and a first image generation network. The matting branch includes three groups of transparency feature generation networks and a second image generation network. The first layer network of the video matting model is an information compression network. The information compression network is connected to the two branches, and the two branches are also connected to each other.

[0359] The connection relationship of the networks included in the segmentation branch and the matting branch and the connection relationship between the two branches will be described below.

[0360] First, the connection relationship of the networks included in the segmentation branch is described. The information compression network is connected to the first reconstruction subnetwork 1 included in the contour feature generation network 1. The first fusion subnetwork 1 included in the contour feature generation network 1 is connected to the first reconstruction subnetwork 2 included in the contour feature generation network 2. The first fusion subnetwork 2 included in the contour feature generation network 2 is connected to the first reconstruction subnetwork 3 included in the contour feature generation network 3. The first fusion subnetwork 3 included in the contour feature generation network 3 is connected to the first image generation network.

[0361] Second, the connection relationship of the networks included in the matting branch is described. The information compression network is connected to the second reconstruction subnetwork 1 included in the transparency feature generation network 1. The second fusion subnetwork 1 included in the transparency feature generation network 1 is connected to the second reconstruction subnetwork 2 included in the transparency feature generation network 2. The second fusion subnetwork 2 included in the transparency feature generation network 2 is connected to the second reconstruction subnetwork 3 included in the transparency feature generation network 3. The second fusion subnetwork 3 included in the transparency feature generation network 3 is connected to the second image generation network.

[0362] Finally, the connection relationship between the segmentation branch and the matting branch is described. The first fusion subnetwork 1 included in the contour feature generation network 1 is connected to the second reconstruction subnetwork 1 included in the transparency feature generation network 1. The first fusion subnetwork 2 included in the contour feature generation network 2 is connected to the second reconstruction subnetwork 2 included in the transparency feature generation network 2. The first fusion subnetwork 3 included in the contour feature generation network 3 is connected to the second reconstruction subnetwork 3 included in the transparency feature generation network 3. In addition, the first image generation network and the second image generation network are respectively connected to the result output network.

[0363] The networks and subnetworks in the video matting model will be described below.

[0364] For the information compression network, when compressing the information of the target video frame, the target video frame is input into the information compression network, and the information compression network compresses the information of the target video frame to obtain the first feature output by the information compression network.

[0365] The implementation of the information compression network for information compression on the target video frame can refer to the foregoing content, which will not be described here.

[0366] Referring to FIG. 24, FIG. 24 is a structural schematic diagram of an information compression network. In the information compression network shown in FIG. 24, the network layers from top to bottom are a convolution layer, a linear layer 1, a batch normalization layer, a nonlinear layer, and a linear layer 2.

[0367] The convolution layer is configured to perform convolution transformation on the target video frame to obtain a fifth image feature.

[0368] The linear layer 1 is configured to perform linear transformation on the fifth image feature based on a convolution kernel to obtain a sixth image feature.

[0369] The batch normalization layer is configured to perform batch normalization processing on the sixth image feature to obtain a seventh image feature.

[0370] The nonlinear layer is configured to perform nonlinear transformation on the seventh image feature to obtain an eighth image feature.

[0371] The linear layer 2 is configured to perform linear transformation on the eighth image feature based on a convolution kernel to obtain the first feature.

[0372] The implementation of the convolution layer, the linear layer 1, the batch normalization layer, the nonlinear layer, and the linear layer 2 for data processing can refer to the foregoing content, which will not be described here.

[0373] For the target first reconstruction sub-network in the target contour feature generation network, when performing feature reconstruction based on the third target feature, the third target feature is input into the target first reconstruction sub-network, and the target first reconstruction sub-network performs feature reconstruction based on the third target feature to obtain a fourth image feature with an increased scale output by the target first reconstruction sub-network. The scale of the contour mask image corresponding to the target contour feature generation network is the same as the scale of the fourth image feature.

[0374] The implementation of the target first reconstruction sub-network for performing feature reconstruction based on the third target feature can refer to the foregoing embodiments, which will not be described here.

[0375] In an embodiment of the present application, the foregoing first reconstruction sub-network is implemented based on a QARepVGG network structure.

[0376] Since the quantization calculation accuracy of the QARepVGG network is high, implementing the foregoing first reconstruction sub-network based on the QARepVGG network structure can improve the quantization calculation capability of the first reconstruction sub-network, thereby improving the accuracy of feature reconstruction performed by the first reconstruction sub-network, and further improving the accuracy of video matting.

[0377] In another embodiment of the present application, the first reconstruction sub-network in the specific contour feature generation network is implemented based on a QARepVGG network structure.

[0378] The specific contour feature generation network is a contour feature generation network corresponding to a contour mask image with a size smaller than the first preset size.

[0379] The first preset size can be a preset size.

[0380] When constructing the video matting model, the size of the contour mask image corresponding to each group of contour feature generation networks can be determined, so that the contour feature generation network corresponding to the contour mask image with a size smaller than the first preset size can be determined as the specific contour feature generation network, so that when constructing the specific contour feature generation network, the specific contour feature generation network is constructed based on the QARepVGG network structure. For other contour feature generation networks, other network structures can be used for construction.

[0381] Since the calculation amount of the U-shaped residual block in the specific contour feature generation network constructed based on the QARepVGG network structure increases with the size of the contour mask image corresponding to the network, when constructing each contour feature generation network, the first reconstruction sub-network in the specific contour feature generation network corresponding to the contour mask image with a size smaller than the first preset size can be implemented based on the QARepVGG network structure, which can reduce the calculation amount of each contour feature generation network, improve the efficiency of obtaining the contour mask image of the object, and thus improve the efficiency of video matting, and also enable lightweight deployment of the video matting model in the terminal.

[0382] For the target first fusion sub-network in the target contour feature generation network, when the fourth image feature and the second sub-state information are fused and the second sub-state information is updated, the fourth image feature can be input into the target first fusion sub-network, and the second sub-hidden state information provided by the target first fusion sub-network is the second sub-state information. Thus, the fourth image feature and the second sub-hidden state information provided by the target first fusion sub-network are fused and the second sub-hidden state information provided by the target first fusion sub-network is updated, so as to obtain the feature of the contour mask image of the object output by the target first fusion sub-network.

[0383] The implementation manner of the target first fusion sub-network for fusing the fourth image feature and the second sub-hidden state information and updating the second sub-hidden state information can refer to the foregoing embodiments, which will not be described here.

[0384] In an embodiment of the present application, the first fusion sub-network is a Gated Recurrent Unit (GRU) or a Long Short-Term Memory (LSTM) unit.

[0385] Both the GRU and the LSTM unit have information memory function. If either of the two units is used as the first fusion sub-network, the unit itself can store hidden state information representing the fusion features of the object contour mask image in the segmented video frame, so as to accurately fuse the fourth image feature and the second sub-hidden state information provided by the unit itself, improve the accuracy of the features of the object contour mask image output by the sub-network, and thus improve the accuracy of the video matting.

[0386] For the first image generation network, when obtaining the first contour mask image based on the features of the finally obtained object contour mask image, the features of the finally obtained object contour mask image are input into the first image generation network, and an image is generated by the first image generation network based on the features, so that the image output by the network is obtained as the first contour mask image.

[0387] The implementation of the first image generation network generating an image based on the features can refer to the foregoing embodiments, which will not be described here.

[0388] For the target second reconstruction sub-network in the target transparency feature generation network, when performing feature reconstruction based on the first target feature and the second target feature, the first target feature and the second target feature are input into the target second reconstruction sub-network, and feature reconstruction is performed by the target second reconstruction sub-network based on the first target feature and the second target feature, so that the second image feature with increased scale output by the target second reconstruction sub-network is obtained, wherein the scale of the transparency mask image corresponding to the target transparency feature generation network is the same as the scale of the second image feature.

[0389] The implementation of the target second reconstruction sub-network performing feature reconstruction based on the first target feature and the second target feature can refer to the foregoing content, which will not be described here.

[0390] In an embodiment of the present application, the second reconstruction sub-network is implemented based on the QARepVGG network structure.

[0391] Since the quantization calculation accuracy of the QARepVGG network is high, implementing the second reconstruction sub-network based on the QARepVGG network structure can improve the quantization calculation capability of the second reconstruction sub-network, thereby improving the accuracy of feature reconstruction of the second reconstruction sub-network based on the first target feature and the second target feature, and further improving the accuracy of video matting.

[0392] In another embodiment of the present application, the second reconstruction sub-network in the specific transparency feature generation network is implemented based on a QARepVGG network structure.

[0393] The specific transparency feature generation network is a transparency feature generation network corresponding to a transparency mask image with a scale less than the second preset scale.

[0394] The second preset scale can be a preset scale. The second preset scale can be the same as or different from the first preset scale.

[0395] When constructing the video matting model, the scale of the transparency mask image corresponding to each group of transparency feature generation networks can be determined. In this way, the transparency feature generation network corresponding to the transparency mask image with a scale less than the second preset scale can be determined as the specific transparency feature generation network. Thus, when constructing the specific transparency feature generation network, the specific transparency feature generation network can be constructed based on the QARepVGG network structure. For other transparency feature generation networks, other network structures can be used for construction.

[0396] Since the calculation amount of the U-shaped residual block in the specific transparency feature generation network constructed based on the QARepVGG network structure increases with the increase of the scale of the transparency mask image corresponding to the network, when constructing each transparency feature generation network, the second reconstruction sub-network in the specific transparency feature generation network corresponding to the transparency mask image with a scale less than the second preset scale can be implemented based on the QARepVGG network structure. In this way, the calculation amount of each transparency feature generation network can be reduced, the efficiency of obtaining the fusion result can be improved, the efficiency of video matting can be improved, and the video matting model can be deployed in a terminal in a lightweight manner.

[0397] For the target second fusion sub-network in the target transparency feature generation network, when the second image feature and the first sub-state information are fused and the first sub-state information is updated, the second image feature can be input into the target second fusion sub-network, and the first sub-hidden state information provided by the target second fusion sub-network is the first sub-state information. In this way, the target second fusion sub-network fuses the second image feature and the first sub-hidden state information provided by itself and updates the first sub-hidden state information provided by itself, so as to obtain the third image feature output by the target second fusion sub-network.

[0398] The implementation manner of the target second fusion sub-network fusing the second image feature and the first sub-hidden state information and updating the first sub-hidden state information can be referred to the foregoing embodiments, which will not be described herein.

[0399] In an embodiment of the present application, the second fusion sub-network is a Gated Recurrent Unit (GRU) or a Long Short-Term Memory (LSTM) unit.

[0400] Both the GRU and the LSTM unit have information memory function. If either of the two units is used as the second fusion sub-network, the unit itself can store hidden state information representing the fusion features of the transparency mask image of the object in the segmented video frame, so as to accurately fuse the second image features and the first sub-hidden state information provided by the unit itself, improve the accuracy of the third image features output by the sub-network, and thus improve the accuracy of the video matting.

[0401] For the second image generation network, when the first contour mask image is obtained based on the finally obtained fusion result, the fusion result is input into the second image generation network, and the second image generation network generates an image based on the fusion result, so as to obtain the image output by the network as the target transparency mask image.

[0402] The implementation of the second image generation network for generating an image based on the fusion result can refer to the foregoing embodiments, which will not be described here again.

[0403] For the result output network, when the region matting is performed, the target transparency mask image and the first contour mask image are input into the result output network, and the result output network performs the region matting on the target video frame according to the target transparency mask image and the first contour mask image, so as to obtain the matting result output by the result output network.

[0404] The implementation of the result output network for performing the region matting on the target video frame can refer to the foregoing embodiments, which will not be described here again.

[0405] As can be seen from the above, in the scheme provided in the present embodiment, the video matting is performed by using each network and sub-network included in the video matting model. Since the video matting model is a pre-trained video matting model, the accuracy of the video matting can be improved by using the video matting model. Moreover, the video matting model does not need to interact with other devices, so the video matting model can be deployed in an offline device, which can improve the convenience of video segmentation.

[0406] Since the data amount of the first target feature and the second target feature is usually large, the calculation amount of the feature reconstruction performed by the target second reconstruction sub-network is large, which leads to a low efficiency of the feature reconstruction performed by the second reconstruction sub-network, and thus the efficiency of the video matting performed by using the video matting model is also low.

[0407] In view of this, in one embodiment of the present application, referring to FIG. 25, a structural diagram of a second video matting model is provided. Compared with FIG. 23, in the video matting model shown in FIG. 25, each transparency feature generation network further includes a feature screening subnetwork, the first fusion subnetwork 1 is connected to the second reconstruction subnetwork 1 through the feature screening subnetwork 1, the first fusion subnetwork 2 is connected to the second reconstruction subnetwork 2 through the feature screening subnetwork 2, and the first fusion subnetwork 3 is connected to the second reconstruction subnetwork 3 through the feature screening subnetwork 3.

[0408] For the target feature screening subnetwork in the target transparency feature generation network, before inputting the feature output by the target first fusion subnetwork into the target second reconstruction subnetwork, the feature can be input into the target feature screening subnetwork, the target screening feature characteristic of the edge contour of the object is screened from the feature by the target feature screening subnetwork, and the target screening feature is input into the target second reconstruction subnetwork in the target transparency feature generation network, and the target second reconstruction subnetwork performs feature reconstruction based on the target screening feature.

[0409] Specifically, the implementation manner of the target feature screening subnetwork screening feature and the implementation manner of the target second reconstruction subnetwork performing feature reconstruction based on the target screening feature and the first target feature can be referred to the foregoing embodiments, which will not be described here.

[0410] As can be seen from the above, in the scheme provided by the present embodiment, by adding a feature screening subnetwork in the transparency feature generation network, the calculation amount of the second reconstruction subnetwork in the transparency feature generation network for feature reconstruction can be reduced, and the efficiency of the second reconstruction subnetwork for feature reconstruction can be improved, so that the efficiency of the video matting model for video matting can be improved.

[0411] In one embodiment of the present application, referring to FIG. 26, a flow diagram of a fourth video matting method is provided. In FIG. 26, the video matting model processes video frame 1 and video frame 2 included in a video in sequence. When processing the video frame 1, the video frame 1 is processed by each network and subnetwork in the model to obtain a matting result 1 corresponding to the video frame 1, wherein each first fusion subnetwork and second fusion subnetwork in the model outputs information to other network layers on one hand, and updates the hidden state information contained therein on the other hand, which is used for fusion with the input feature when the model processes the next frame video frame 2. When processing the video frame 2, the video frame 2 is processed by each network and subnetwork in the model to obtain a matting result 2 corresponding to the video frame 2.

[0412] In an embodiment of the present application, when the video matting model contains a large number of sets of contour feature generation networks and transparency feature generation networks, the video matting model has a large amount of calculation in processing video frames. In view of this, the first fusion sub-network in the last one or more sets of contour feature generation networks contained in the video matting model can be removed, and the second fusion sub-network in the last one or more sets of transparency feature generation networks contained in the video matting model can also be removed, so as to reduce the calculation amount of the video matting model in processing video frames, and also enable the video matting model to be deployed in a terminal in a lightweight manner.

[0413] Taking the removal of the first fusion sub-network in the last set of contour feature generation networks as an example, referring to FIG. 27, a third structure diagram of a video matting model is provided. Compared with the video matting model shown in FIG. 23, the video matting model shown in FIG. 27 contains only the first reconstruction sub-network 3 in the last set of contour feature generation networks 3, and the output result of the first reconstruction sub-network 3 is the feature of the finally obtained contour mask image of the object.

[0414] In an embodiment of the present application, referring to FIG. 28, a fourth structure diagram of a video matting model is provided. The video matting model shown in FIG. 26 contains multiple layers of cascaded information compression networks. The output result of each layer of information compression network is a first sub-feature, and one layer of information compression network is connected with the first reconstruction sub-network in one set of contour feature generation networks. The last layer of information compression network is connected with the first set of contour feature generation networks. The scale of the first sub-feature output by the connected information compression network is the same as the scale of the third target feature to be processed by the contour feature generation network. The first sub-feature output by the last layer of information compression network is used as the third target feature to be processed by the first set of contour feature generation networks. The first reconstruction sub-network in the other sets of contour feature generation networks performs feature reconstruction based on the third target feature and the first sub-feature output by the information compression network connected with the network, so as to improve the accuracy of feature reconstruction and improve the accuracy of video matting.

[0415] The training process of the above-mentioned video matting model is described below.

[0416] In an embodiment of the present application, referring to FIG. 29, a flowchart of a first model training method is provided. In the present embodiment, the above-mentioned method includes the following steps S1701-S1706.

[0417] Step S1701: input a first sample video frame in a sample video into an initial model of a video matting model for processing, to obtain a first sample contour mask image of an object in the first sample video frame output by a first image generation network in the initial model.

[0418] The sample video can be any video obtained through a network, a video library or other channels. In addition, after obtaining the video through a network, a video library or other channels, a plurality of videos can be spliced into one video to obtain a spliced video as a sample video.

[0419] The first sample contour mask image has the same size as the first sample video frame, and a pixel value of a pixel point in the first sample contour mask image represents a confidence degree of a pixel point at a same position in the first sample video frame predicted by the model belonging to a region where the object is located.

[0420] The initial model is configured to process a video frame input into the model according to a model parameter of the initial model which is not trained completely. In the process of processing the video frame by the model, an image output by a first image generation network in the model can be obtained.

[0421] Specifically, after the first sample video frame is input into the initial model, the initial model can process the first sample video frame according to the model parameter of the initial model, and obtain an image output by the first image generation network in the model as a first sample contour mask image of the object in the first sample video frame.

[0422] Step S1702: obtaining a first difference between a label mask image corresponding to the first sample video frame and a label mask image corresponding to a second sample video frame.

[0423] The second sample video frame is a video frame in the sample video before the first sample video frame and separated by a preset number of frames.

[0424] The preset number of frames is a preset number of frames, for example, 3 frames, 4 frames or other values of frames.

[0425] The first sample video frame can be a video frame at a preset number of frames or any video frame after a preset number of frames in the sample video.

[0426] The first difference can be calculated by the terminal itself or other devices, and the terminal device can obtain the calculated first difference from the other devices.

[0427] The implementation of the terminal or other devices calculating the first difference is described below.

[0428] The terminal or other device can obtain a label mask image corresponding to each sample video frame in the sample video. The label mask image corresponding to the sample video frame can be understood as the actual mask image of the object in the sample video frame. Thus, after determining the second sample video frame according to the frame number of the first sample video frame and the preset frame number, the label mask image corresponding to the first sample video frame and the label mask image corresponding to the second sample video frame can be obtained from the obtained label mask images corresponding to each sample video frame, so as to calculate the first difference between the two obtained label mask images.

[0429] In calculating the first difference between the two label mask images, in one implementation, the pixel values of the pixel points at the same positions of the two images can be subtracted, and the number of results other than "0" in the operation results of each pixel point can be counted as the first difference, or the proportion of the number of results other than "0" in the total number of pixel points of the label mask image can be counted as the first difference. In another implementation, the similarity of the two images can be calculated, and the operation result obtained by subtracting the calculated similarity from 1 can be taken as the first difference.

[0430] Step S1703: obtaining a second difference between the first sample contour mask image and a second sample contour mask image, wherein the second sample contour mask image is a mask image output by the first image generation network when the initial model processes the second sample video frame.

[0431] Specifically, similar to the aforementioned video matting process, in the model training process, each sample video frame in the sample video can be input into the model frame by frame to obtain the sample contour mask image of the object in each sample video frame output by the first image generation network in the model. After obtaining the first sample contour mask image, the second sample video frame with a preset frame number before the first sample video frame can be determined in the video frame processed by the model, and the second sample contour mask image output by the first image generation network when the model processes the second sample video frame can be obtained, and the second difference between the first sample contour mask image and the second sample contour mask image can be calculated.

[0432] The implementation of calculating the second difference is the same as the implementation of calculating the first difference in the aforementioned step S1702, and thus will not be described herein.

[0433] Step S1704: obtaining a third difference between a first sample transparency mask image and a second sample transparency mask image.

[0434] The first sample transparency mask image is a mask image output by the second image generation network when the initial model processes the first sample video frame.

[0435] The second sample transparency mask image is an image output by the second image generation network in the initial model when the initial model processes the second sample video frame.

[0436] Specifically, after the first sample video frame is input into the initial model, the initial model can process the first sample video frame according to the model parameters configured by itself, and obtain an image output by the second image generation network in the model as the first sample transparency mask image; after the second sample video frame is input into the initial model, the initial model can process the second sample video frame according to the model parameters configured by itself, and obtain an image output by the second image generation network in the model as the second sample transparency mask image. In this way, after the first sample transparency mask image and the second sample transparency mask image are obtained, the third difference between the first sample transparency mask image and the second sample transparency mask image can be calculated.

[0437] The implementation of calculating the third difference is the same as that of calculating the first difference in the foregoing step S1702, and will not be described here again.

[0438] Step S1705: calculating a training loss based on the first difference, the second difference, and the third difference.

[0439] The training loss can be calculated based on the first difference, the second difference, and the third difference by using a loss function, an algorithm, or the like.

[0440] Step S1706: adjusting model parameters of the initial model based on the training loss to obtain a video matting model.

[0441] Specifically, based on the training loss, the model parameters of the initial model can be adjusted by any one of the following three implementation manners.

[0442] In the first implementation manner, for each model parameter in the initial model, a corresponding relationship between the training loss and an adjustment amplitude of the model parameter can be set in advance. After the training loss is calculated, the actual adjustment amplitude of adjusting the model parameter can be calculated according to the corresponding relationship, so that the model parameter is adjusted according to the actual adjustment amplitude.

[0443] In the second implementation manner, the initial model usually needs to be trained using a large amount of sample data. During the training process, the training loss needs to be calculated constantly, and the model parameters of the initial model need to be adjusted constantly based on the training loss. Therefore, after the training loss is calculated, the change difference of the training loss can be determined according to the training loss and the training loss calculated before, and the model parameters of the initial model are adjusted according to the change difference.

[0444] In the third implementation, based on the training loss, the initial model can be adjusted by using a model parameter adjustment algorithm or function.

[0445] As can be seen from the above, in the scheme provided by the embodiment, the first sample video frame and the second sample video frame with the interval of the preset number of frames often have a time domain correlation. Thus, the first difference between the label mask image corresponding to the first sample video frame and the label mask image corresponding to the second sample video frame, the second difference between the first sample contour mask image and the second sample contour mask image, and the third difference between the first sample transparency mask image and the second sample transparency mask image are obtained. The training loss is calculated based on the first difference, the second difference, and the third difference. When the initial model is adjusted based on the training loss, the initial model can learn the time domain correlation between different video frames of the video. Thus, the accuracy of the trained model can be improved. Furthermore, when the model is used for video matting, the accuracy of the video matting can be improved.

[0446] When the second difference is obtained, in addition to the manner mentioned in step S1703, the manner mentioned in step S1703A in the embodiment shown in FIG. 30 can also be used.

[0447] In one embodiment of the present application, referring to FIG. 30, a flowchart of a second model training method is provided.

[0448] In the embodiment, the first sample contour mask image includes a first mask sub-image indicating a region where the object is located in the first sample video frame and a second mask sub-image indicating a region other than the object in the first sample video frame.

[0449] The pixel value of the pixel point in the first mask sub-image represents the confidence that the pixel point at the same position in the first sample video frame predicted by the model belongs to the region where the object is located. The pixel value of the pixel point in the second mask sub-image represents the confidence that the pixel point at the same position in the first sample video frame predicted by the model belongs to the region other than the object.

[0450] As shown in FIG. 31, FIG. 31 is a mask image provided by an embodiment of the present application, which is a first sample mask image.

[0451] The mask image shown in FIG. 31 includes two sub-images, which are a first mask sub-image indicating a region where the object is located in the first sample video frame and a second mask sub-image indicating a region other than the object in the first sample video frame.

[0452] The second sample contour mask image includes a third mask sub-image indicating a region where the object is located in the second sample video frame and a fourth mask sub-image indicating a region other than the object in the second sample video frame.

[0453] The pixel value of the pixel point in the third mask subgraph represents the confidence of the pixel point at the same position in the second sample video frame predicted by the model belonging to the region where the object is located, and the pixel value of the pixel point in the fourth mask subgraph represents the confidence of the pixel point at the same position in the second sample video frame predicted by the model belonging to the region other than the object.

[0454] In this case, the step S1703 can be implemented by the following step S1703A.

[0455] Step S1703A: obtain the difference between the first mask subgraph and the third mask subgraph, and obtain the difference between the second mask subgraph and the fourth mask subgraph, to obtain a second difference containing the obtained differences.

[0456] The implementation of obtaining the difference between the first mask subgraph and the third mask subgraph and the difference between the second mask subgraph and the fourth mask subgraph is the same as the implementation of obtaining the first difference or the second difference, which will not be repeated here.

[0457] After obtaining the difference between the first mask subgraph and the third mask subgraph and the difference between the second mask subgraph and the fourth mask subgraph, the two differences can be accumulated to obtain a second difference containing the two differences, or the average of the two differences can be taken as the second difference, or the larger difference of the two differences can be determined as the second difference, etc.

[0458] As can be seen from the above, in the scheme provided by the embodiment, since the region in the video frame is composed of the region where the object is located and the region other than the object, the greater the difference in the region where the object is located in different video frames, the greater the difference in the region other than the object in different video frames. It can be seen that the difference in the region other than the object can also reflect the difference in the region where the object is located. Therefore, obtaining the second difference according to the difference between the first mask subgraph and the third mask subgraph and the difference between the second mask subgraph and the fourth mask subgraph is a comprehensive calculation of the second difference from two different angles, which can improve the accuracy of the second difference, thereby improving the accuracy of model training and improving the accuracy of video matting using the model.

[0459] The second fusion subnetwork in the above video matting model can ensure that the features of the transparency mask image corresponding to the segmented video frame are considered in the process of matting the target video frame, that is, the temporal continuity between video frames is ensured. However, in this case, if the second image generation network is not hard limited to output a binary image, there may be a semi-transparent area in the image finally output by the second image generation network. The pixel value of a pixel point in the target transparency mask image output by the second image generation network ranges from 0 to 1, and a semi-transparent area may appear in the target transparency mask image output by the second image generation network. It is difficult to determine the cause of the semi-transparent area in the target transparency mask image, and it is also difficult to train the networks and subnetworks in the matting branch of the model.

[0460] In view of this, in one embodiment of the present application, referring to FIG. 32, a structural diagram of a second image generation network is provided. In FIG. 32, the second image generation network includes a hard segmentation subnetwork, a result fusion subnetwork, and an image generation subnetwork.

[0461] The input of the hard segmentation subnetwork is the fusion result output by the second fusion subnetwork in the last set of transparency feature generation networks.

[0462] The hard segmentation subnetwork is used to obtain a second contour mask image of an object in the target video frame based on the fusion result.

[0463] The result fusion subnetwork is used to fuse the second contour mask image and the fusion result to obtain a target fusion feature.

[0464] The image generation subnetwork is used to obtain a target transparency mask image of the edge of the object in the target video frame based on the target fusion feature.

[0465] In the model training process, not only can the initial model be trained based on the first, second, and third differences, but also the mask images corresponding to different sample video frames output by the hard segmentation subnetwork can be obtained, and a fourth difference between the obtained mask images can be calculated, so that the initial model can be trained according to the fourth difference. Since the hard segmentation subnetwork outputs a contour mask image belonging to a binary image, the possibility of a semi-transparent area appearing in the mask image output by the second image generation network can be excluded in the training process, so that the training of the initial model can be accurately and quickly realized.

[0466] As shown in FIG. 33, the left side of FIG. 33 shows the final matting result obtained without using the fourth difference for model training, and the right side of FIG. 33 shows the final matting result obtained by using the fourth difference for model training.

[0467] Next, the electronic device provided by the embodiments of the present application is introduced.

[0468] Electronic equipment can be equipped or portable terminal devices with other operating systems, such as mobile phones, tablet computers, desktop computers, laptop computers, handheld computers, notebook computers, ultra-mobile personal computers (UMPCs), netbooks, as well as cellular phones, personal digital assistants (PDAs), augmented reality (AR) devices, virtual reality (VR) devices, artificial intelligence (AI) devices, wearable devices, in-vehicle devices, smart home devices and / or smart city devices, etc.

[0469] 34 exemplarily shows an electronic device 100 provided in an embodiment of the present application. The electronic device 100 may support AOD, and its screen has pixel-level luminescence capability, supporting lighting up only some pixels of the screen, such as an organic light emitting diode (OLED) screen.

[0470] The electronic device 100 may include a processor 110, a display screen 120, a camera 130, an internal memory 140, a SIM (Subscriber Identification Module) card interface 150, a USB (Universal Serial Bus) interface 160, a charging management module 170, a power management module 171, a battery 172, a sensor module 180, a mobile communication module 190, a wireless communication module 200, an antenna 1, and an antenna 2. The sensor module 180 may include a pressure sensor 180A, a fingerprint sensor 180B, a touch sensor 180C, an ambient light sensor 180D, and the like.

[0471] The structures illustrated in the embodiments of this application do not constitute specific limitations on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0472] The processor 110 can include one or more processing units, for example: the processor 110 can include a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video code, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent components or integrated in one or more processors. In some embodiments, the electronic device 100 can also include one or more processors 110. Among them, the controller can generate operation control signals according to instruction operation codes and timing signals, complete the control of fetching instructions and executing instructions. In other embodiments, the processor 110 can also be provided with a memory for storing instructions and data. Exemplarily, the memory in the processor 110 can be a cache memory. The memory can save instructions or data that the processor 110 has just used or repeatedly uses. If the processor 110 needs to use the instructions or data again, it can be directly called from the memory. In this way, repeated access is avoided, the waiting time of the processor 110 is reduced, and the efficiency of the electronic device 100 in processing data or executing instructions is improved.

[0473] In some embodiments, the processor 110 can include one or more interfaces. The interfaces can include Inter-Integrated Circuit (I2C) interfaces, Inter-Integrated Circuit Sound (I2S) interfaces, Pulse Code Modulation (PCM) interfaces, Universal Asynchronous Receiver / Transmitter (UART) interfaces, Mobile Industry Processor Interface (MIPI) interfaces, General-Purpose Input / Output (GPIO) interfaces, SIM card interfaces, and / or USB interfaces, etc. The USB interface 160 is an interface that conforms to the USB standard specification, and can be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 160 can be used to connect a charger to charge the electronic device 100, and can also be used to transmit data between the electronic device 100 and a peripheral device. The USB interface 160 can also be used to connect a headset to play audio through the headset.

[0474] The interface connection relationship between the modules shown in the embodiments of the present application is used for illustrative description and does not constitute a structural limitation of the electronic device 100. In some other embodiments of the present application, the electronic device 100 can also use different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0475] The wireless communication function of the electronic device 100 can be implemented by the antenna 1, the antenna 2, the mobile communication module 190, the wireless communication module 200, the modem processor, and the baseband processor, etc.

[0476] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example, the antenna 1 can be multiplexed as a diversity antenna of a wireless local area network. In some other embodiments, the antennas can be used in combination with tuning switches.

[0477] The electronic device 100 implements the display function through the GPU, the display screen 120, and the application processor, etc. The GPU is a microprocessor for image processing, which is connected to the display screen 120 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs that execute program instructions to generate or change display information.

[0478] The display screen 120 is configured to display images, videos, etc. The display screen 120 includes a display panel, which can be an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), etc. having a pixel-level light-emitting capability.

[0479] The electronic device 100 can implement a photographing function through an ISP, a camera 130, a video codec, a GPU, a display screen 120, and an application processor, etc. The camera 130 includes a front camera and a rear camera.

[0480] The ISP is configured to process data fed back by the camera 130. For example, when photographing, the shutter is opened, and light is transmitted to the camera photosensitive element through the lens, and the light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing, and converts it into an image visible to the naked eye. The ISP can optimize the noise, brightness, and color of the image through an algorithm, and can also optimize the exposure and color temperature of the shooting scene, etc. In some embodiments, the ISP can be arranged in the camera 130.

[0481] The camera 130 is configured to take photos or videos. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into a standard Red Green Blue (RGB), YUV, etc. format image signal. In some embodiments, the electronic device 100 can include one or N cameras 130, where N is a positive integer greater than 1.

[0482] The digital signal processor is configured to process digital signals, which can include digital image signals and other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is configured to perform Fourier transform on the frequency point energy, etc.

[0483] A video codec is used for compressing or decompressing digital video. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in a variety of encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG 2, MPEG 3, and MPEG 4.

[0484] An NPU is a Neural-Network (NN) computing processor that quickly processes input information by drawing on the structure of a biological neural network, such as the mode of transmission between neurons in the human brain, and can also constantly self-learn. Through the NPU, the electronic device 100 can implement intelligent cognitive applications, such as image recognition, face recognition, voice recognition, text understanding, and the like.

[0485] The internal memory 140 can be used to store one or more computer programs including instructions. The processor 110 can cause the electronic device 100 to perform the video matting method provided in some embodiments of the present application, and various applications and data processing, and the like, by running the above-mentioned instructions stored in the internal memory 140. The internal memory 140 can include a program storage area and a data storage area. The program storage area can store an operating system, and the program storage area can also store one or more applications (such as a gallery, contacts, and the like), and the like. The data storage area can store data created during use of the electronic device 100 (such as photos, contacts, and the like), and the like. In addition, the internal memory 140 can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more disk storage components, flash memory components, Universal Flash Storage (UFS), and the like. In some embodiments, the processor 110 can cause the electronic device 100 to perform the video matting method provided in the embodiments of the present application, and other applications and data processing, by running the instructions stored in the internal memory 140 and / or the instructions stored in the memory disposed in the processor 110.

[0486] The internal memory 140 can be used to store program codes or program instructions for implementing the wallpaper display method and the video matting method provided in the embodiments of the present application. The processor 110 can invoke the program codes or program instructions of the wallpaper display method and the video matting method stored in the internal memory 140 to execute the wallpaper display method and the video matting method of the embodiments of the present application.

[0487] The sensor module 180 can include a pressure sensor 180A, a fingerprint sensor 180B, a touch sensor 180C, an ambient light sensor 180D, and the like.

[0488] The pressure sensor 180A is configured to sense a pressure signal and convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 180A can be disposed on the display 120. The pressure sensor 180A can be of various types, such as a resistive pressure sensor, an inductive pressure sensor, or a capacitive pressure sensor. The capacitive pressure sensor can include at least two parallel plates of conductive material. When a force is applied to the pressure sensor 180A, the capacitance between the electrodes changes, and the electronic device 100 determines the intensity of the pressure based on the change in capacitance. When a touch operation is applied to the display 120, the electronic device 100 detects the touch operation based on the pressure sensor 180A. The electronic device 100 can also calculate the position of the touch based on the detection signal of the pressure sensor 180A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation instructions. For example, when a touch operation with an intensity less than a first pressure threshold is applied to a short message application icon, an instruction to view a short message is executed; when a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the short message application icon, an instruction to create a new short message is executed.

[0489] The fingerprint sensor 180B is configured to acquire a fingerprint. The electronic device 100 can use the acquired fingerprint characteristics to implement functions such as unlocking, accessing an application lock, taking a photograph, and answering an incoming call.

[0490] The touch sensor 180C, also referred to as a touch device. The touch sensor 180C can be disposed on the display 120, and the touch sensor 180C and the display 120 together form a touch screen, also referred to as a touch screen. The touch sensor 180C is configured to detect a touch operation applied to or near the touch sensor 180C. The touch sensor 180C can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display 120. In other embodiments, the touch sensor 180C can also be disposed on the surface of the electronic device 100 and disposed at a different position from the display 120.

[0491] The ambient light sensor 180D is configured to sense the brightness of ambient light. The electronic device 100 can adaptively adjust the brightness of the display 120 based on the sensed brightness of the ambient light. The ambient light sensor 180D can also be used to automatically adjust the white balance when photographing. The ambient light sensor 180D can also transmit information about the environment in which the device is located to the GPU.

[0492] The ambient light sensor 180D is also configured to obtain the brightness, light ratio, color temperature, and the like of the acquisition environment of the image acquired by the camera 130.

[0493] FIG. 35 illustrates one software system architecture employed by the electronic device 100. The software system architecture can employ a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture.

[0494] A layered architecture divides the software system of the terminal into several layers, each of which has a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the software system can be divided into five layers, namely, applications, application framework, system libraries, hardware abstraction layer (HAL), and kernel.

[0495] The application layer can include a series of application packages, and the application layer runs applications by calling application programming interfaces (APIs) provided by the application framework layer. As shown in FIG. 35, the application packages can include browser, gallery, music, and video applications. It can be understood that each of the above-mentioned application ports can be used to receive data.

[0496] The application framework layer provides APIs and programming frameworks for the applications of the application layer. The application framework layer includes some pre-defined functions. As shown in FIG. 35, the application framework layer can include a window manager, a content provider, a view system, a resource manager, a notification manager, and a DHCP (Dynamic Host Configuration Protocol) module.

[0497] The system library can include a plurality of functional modules, such as a surface manager, a three-dimensional graphics processing library, a two-dimensional graphics engine, and a file library.

[0498] The hardware abstraction layer can include a plurality of library modules, such as a display library module and a motor library module. The terminal system can load the corresponding library modules for the device hardware, thereby achieving the purpose of the application framework layer accessing the device hardware.

[0499] The kernel layer is the layer between hardware and software. The kernel layer is used to drive hardware to work. The kernel layer at least includes display driver, audio driver, sensor driver, and motor driver, and the present application embodiment does not limit this. It can be understood that the display driver, the audio driver, the sensor driver, and the motor driver can be regarded as a driving node. Each of the above-mentioned driving nodes includes an interface that can be used to receive data.

[0500] The term "user interface (UI)" in the specification and claims of the present application and the accompanying drawings is a medium interface for interaction and information exchange between an application or an operating system and a user, which realizes conversion between an internal form of information and a form acceptable by the user. The user interface of an application is source code written in a specific computer language such as Java, extensible markup language (XML), etc., and the interface source code is parsed, rendered, and finally presented as content recognizable by the user such as a picture, a text, a button, etc. on a terminal device. A control (also referred to as a widget) is a basic element of a user interface, and typical controls include a toolbar, a menu bar, a text box, a button, a scrollbar, a picture, and a text. The properties and content of a control in an interface are defined by a tag or a node, such as an XML tag. <textview> 、 <imgview> 、 <videoview>The interface is defined by nodes that specify the controls contained in the interface. One node corresponds to one control or property in the interface, and the nodes are parsed and rendered to present the content visible to the user. In addition, many applications, such as hybrid applications, also contain web pages in the interface. A web page, also referred to as a page, can be understood as a special control embedded in the interface of an application. The web page is a source code written in a specific computer language, such as hyper text markup language (HTML), cascading style sheets (CSS), JavaScript (JS), etc. The web page source code can be loaded and displayed by a browser or a web page display component similar to the function of a browser to present content recognizable to the user. The specific content contained in the web page is also defined by tags or nodes in the web page source code, such as HTML defines the content of the web page by tags such as 、 、 <video> 、 <canvas>To define the elements and attributes of a web page.

[0501] A common form of user interface is a graphic user interface (GUI), which refers to a user interface that displays in a graphical manner. It can be an icon, window, control, etc. interface element displayed in the display screen of an electronic device, wherein the control can include an icon, button, menu, tab, text box, dialog box, status bar, navigation bar, Widget, etc. visual interface element.

[0502] The steps in the above method embodiments provided by the present application can be completed by integrated logic circuits of hardware in the processor or instructions in the form of software. The method steps disclosed in combination with the embodiments of the present application can be directly embodied as hardware processor execution completion, or executed by a combination of hardware and software modules in the processor.

[0503] The present application also provides an electronic device, which can include a memory and a processor. The memory can be used to store a computer program, and the processor can be used to call the computer program in the memory to enable the electronic device to execute the method in any one of the above embodiments.

[0504] The present application also provides a chip system, as shown in FIG. 36, which can include a processor 2201 for realizing the functions involved in the method executed by the electronic device in any one of the above embodiments. The chip system can also include a memory for saving program instructions and data, which is located inside or outside the processor. The chip system can also include an input / output interface, which can serve as a bridge connecting the processor and the peripheral input / output devices. Specifically, when the peripheral input device collects input data, such as the touch screen collecting user touch operations, the audio input circuit collecting audio data, the camera collecting image data, etc., the input / output interface can transmit the input data to the processor for processing; after the processor completes the processing of the input data and generates a processing result, the input / output interface can transmit the processing result to the peripheral output device for output, such as transmitting the processing result to the display screen for display, transmitting the processing result to the audio output circuit for voice output, etc.

[0505] The chip system can be composed of a chip, or can include a chip and other discrete devices.

[0506] Optionally, the processor in the chip system can be one or more. The processor can be implemented by hardware or software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc. When implemented by software, the processor can be a general-purpose processor that realizes by reading software codes stored in the memory.

[0507] Optionally, the memory in the chip system can also be one or more. The memory can be integrated with the processor, or can be arranged separately from the processor, and the embodiments of the present application are not limited. Illustratively, the memory can be a non-transient processor, for example, a read-only memory (ROM), which can be integrated on the same chip as the processor, or can be arranged separately on different chips, and the embodiments of the present application do not make specific limitations on the type of memory and the arrangement of the memory and the processor.

[0508] Illustratively, the chip system can be a field programmable gate array (FPGA), can be an application specific integrated chip (ASIC), can also be a system chip (SoC), can also be a central processor unit (CPU), can also be a network processor (NP), can also be a digital signal processing circuit (DSP), can also be a micro controller unit (MCU), can also be a programmable logic device (PLD) or other integrated chip.

[0509] The present application also provides a computer program product, which comprises a computer program (also referred to as code or instructions). When the computer program is executed, the computer executes the method performed by the electronic device in any one of the above embodiments.

[0510] The present application also provides a computer readable storage medium, which stores a computer program (also referred to as code or instructions). When the computer program is executed, the computer executes the method performed by the electronic device in any one of the above embodiments.

[0511] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, all or part of the processes or functions according to the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server, data center, etc. that includes one or more available media sets. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk), etc.

[0512] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by a computer program instructing the relevant hardware, which can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-mentioned method embodiments. The aforementioned storage medium includes ROM or random access memory (RAM), magnetic disk or optical disk, and various media that can store program codes.

[0513] In summary, the above is only an embodiment of the technical scheme of the present application, and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. according to the disclosure of the present application shall be included in the protection scope of the present application.< / canvas> < / video> < / videoview> < / imgview> < / textview>

Claims

1. A wallpaper display method, characterized in that: include: An electronic device in a locked or unlocked state detects a first operation, and in response to the first operation, the electronic device displays a first interface, where the first interface includes an AOD wallpaper; When the electronic device displays the first interface, in response to a second operation, the electronic device displays a second interface, where the second interface includes a lock screen wallpaper; In response to a third operation for unlocking, the electronic device displays a third interface, wherein the third interface includes a desktop wallpaper; Among them, the AOD wallpaper, the lock screen wallpaper and the desktop wallpaper respectively correspond to the first video frame, the second video frame and the third video frame in the first video, the second video frame is after the first video frame, and the third video frame is after the second video frame.

2. The method according to claim 1, wherein The first video frame, the second video frame, and the third video frame are all one frame, the frame interval between the first video frame and the second video frame is greater than the first frame interval, and the frame interval between the second video frame and the third video frame is greater than the second frame interval.

3. The method according to claim 1, wherein The first video frame, the second video frame, and the third video frame are all multi-frames, the start frame in the second video frame is the end frame in the first video frame, and the start frame in the third video frame is the end frame in the second video frame.

4. The method according to any one of claims 1 to 3, wherein The first video is selected from a gallery of videos.

5. The method according to any one of claims 1 to 4, wherein The first video is selected from videos provided by a settings application.

6. The method according to any one of claims 1 to 5, wherein The method further comprises: The electronic device displays a fourth interface, wherein the fourth interface includes a video frame sequence of the second video; The first video is selected from the video frame sequence, wherein the first video is a partial video clip in the second video or the entire second video.

7. The method according to claim 6, wherein There is a selection box on the video frame sequence, and the video segment in the selection box is the first video; the user operation of selecting the video segment is specifically an operation of sliding the selection box along the video frame sequence.

8. The method according to claim 6 or 7, wherein: The fourth interface also includes a first preview, which is used to show the display status of the AOD wallpaper, the lock screen wallpaper and the desktop wallpaper at the first interface, the second interface and the third interface respectively.

9. The method according to claim 8, wherein The first preview is a video, and when played, the first preview presents a dynamic display process of switching from the AOD wallpaper to the lock screen wallpaper and then to the desktop wallpaper.

10. The method according to any one of claims 1 to 9, wherein The method further includes: the electronic device displays a fifth interface, and the main object and background of the wallpaper in the fifth interface are allowed to be selected separately by the user.

11. The method according to claim 10, wherein Before the electronic device displays the fifth interface, the method further includes: the electronic device performing a cutout operation on the main object in the first video to separate the main object from the background.

12. The method according to claim 11, wherein The method further comprises: In response to a user operation of selecting the subject object for subject object editing, the electronic device updates the image of the subject object, and refreshes the AOD wallpaper, the lock screen wallpaper, and the desktop wallpaper based on the updated subject object image.

13. The method according to claim 11 or 12, wherein: Also includes: In response to a user operation of selecting the background of the wallpaper for background editing, the electronic device updates the background image and refreshes the AOD wallpaper, the lock screen wallpaper, and the desktop wallpaper based on the updated background image.

14. An electronic device, characterized in that: The method comprises one or more processors and one or more memories; wherein the one or more memories are coupled to the one or more processors, and the one or more memories are used to store computer program code, wherein the computer program code includes computer instructions, and when the one or more processors execute the computer instructions, the method according to any one of claims 1 to 13 is executed.

15. A chip system, applied to electronic equipment, comprising one or more processors, characterized in that: The processor is configured to call computer instructions so as to execute the method according to any one of claims 1 to 13.

16. A computer-readable storage medium comprising a computer-executable program, characterized in that: When the computer-executable program is run on an electronic device, the method according to any one of claims 1 to 13 is executed.