Video playback methods, devices, electronic devices, and computer-readable media

By adjusting masks based on crossover union values between frames, the method addresses mask flickering issues, improving the user experience in video playback with barrage features.

JP2026525370APending Publication Date: 2026-07-29SHANGHAI BILIBILI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SHANGHAI BILIBILI TECH CO LTD
Filing Date
2024-10-09
Publication Date
2026-07-29

AI Technical Summary

Technical Problem

Existing video playback methods with barrage features suffer from mask flickering due to abrupt changes in mask coverage areas, leading to a poor user experience.

Method used

A video playback method that adjusts masks based on crossover union values between adjacent frames to ensure the adjusted masks satisfy a ratio threshold, displaying bullet lines outside the adjusted mask area during playback.

Benefits of technology

This approach effectively prevents mask flickering and enhances the viewing experience by ensuring smooth and continuous display of video content and bullet patterns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026525370000001_ABST
    Figure 2026525370000001_ABST
Patent Text Reader

Abstract

This application provides a video playback method, apparatus, electronic device, and computer-readable medium. This application obtains a first mask in a first video frame of a video to be played back, where the first mask is used to independently display a first object, and in response to the existence of a second mask in a second video frame that is simultaneously used to independently display the first object, a first cross-over union value between the first mask and the second mask is determined. Hereinafter, in response to the second video frame being the frame preceding the first video frame and the first cross-over union value not meeting the requirements of a first ratio threshold, the first mask is adjusted using at least the second mask so that the adjusted first cross-over union value meets the first ratio threshold. Furthermore, when playing back the video to be played back, the bullet lines in the video to be played back are displayed using an area other than the display area corresponding to the adjusted first mask. This ensures the effectiveness of the mask, avoids mask flickering due to changes in the mask's coverage area, and improves the viewing experience of the video and bullet lines.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-reference to Related Applications

[0001] This application claims the priority of a Chinese patent application with the application number 202311356511.2 and the invention title "Video Playback Method, Apparatus, Electronic Device and Computer-readable Medium", which was filed with the China National Intellectual Property Administration on October 18, 2023, and all of its contents are incorporated herein by reference.

Technical Field

[0002] This application relates to the field of computer vision, and particularly to a video playback method, apparatus, electronic device and computer-readable medium.

Background Art

[0003] With the development of computer technology, users can obtain content online via the Internet. For example, users can view and obtain video content via the Internet. In this situation, in order to further improve the user interaction experience, video sharing and providing websites enable users to realize interaction by sending barrage messages. For example, when viewing a video, users can provide opinions and views on the video content through the barrage sending control and realize communication with other users.

[0004] Under this background, how to improve the user experience of the barrage function is worthy of attention and has become an urgently required issue.

Summary of the Invention

[0005] In several aspects of this application, a video playback method, apparatus, electronic device, and computer-readable storage medium are provided, which, in a scene where a mask is used to determine the display area of ​​a bullet pattern in order to avoid obstruction between video content and bullet patterns, further select whether or not to adjust the mask based on the cross-over union of masks in adjacent frames, thereby ensuring the effectiveness of using the mask and avoiding flickering of the mask due to abrupt changes in the mask's coverage area. This can improve the viewing experience of video and bullet patterns.

[0006] One aspect of this application provides a video playback method comprising: obtaining a first mask in a first video frame of a video to be played back, wherein the first mask is used to independently display a first object; determining a first crossover union value between the first mask and the second mask in response to the presence of a second mask in a second video frame that is simultaneously used to independently display the first object, wherein the second video frame is the frame preceding the first video frame; adjusting the first mask using at least the second mask in response to the first crossover union value not satisfying the requirements of a first ratio threshold, such that the adjusted first crossover union value satisfies the first ratio threshold; and playing back the video to be played back, and displaying a barrage of text in the video to be played back using an area other than the display area corresponding to the adjusted first mask during playback.

[0007] In another aspect of this application, an apparatus for video playback is provided, comprising: an acquisition module configured to acquire a first mask in a first video frame of a video to be played back, wherein the first mask is used to independently display a first object; a first determination module configured to determine a first cross-overunion value between the first mask and the second mask in response to the presence of a second mask in a second video frame that is simultaneously used to independently display the first object, wherein the second video frame is the frame preceding the first video frame; an update module configured to adjust the first mask using at least the second mask in response to the first cross-overunion value not satisfying the requirements of a first ratio threshold, such that the adjusted first cross-overunion value satisfies the first ratio threshold; and a first playback module configured to play back the video to be played back and to display a barrage of text in the video to be played back using an area other than the display area corresponding to the adjusted first mask during playback.

[0008] Another aspect of this application provides an electronic device including at least one processor and a memory communicably connected to the at least one processor, wherein the memory stores computer-readable instructions that can be executed by the at least one processor, and the execution of the computer-readable instructions by the at least one processor enables the at least one processor to perform the video playback method provided above.

[0009] In another aspect of this application, a computer-readable storage medium is provided, which stores computer-readable instructions, and in which the video playback method described above is realized when the computer-readable instructions are executed by a processor.

[0010] In the solution provided by the embodiment of this application, a first mask is obtained in a first video frame of the video to be played back, and in response to the fact that the first mask is used to display a first object independently and a second mask is present in a second video frame that is used simultaneously to display the first object independently, a first crossing overunion value between the first mask and the second mask is determined, and in response to the fact that the second video frame is the frame preceding the first video frame and the first crossing overunion value does not satisfy the requirements of a first ratio threshold, the first mask is adjusted using at least the second mask so that the adjusted first crossing overunion value satisfies the first ratio threshold, and the video to be played back is played back, and the bullet lines in the video to be played back are displayed using an area other than the display area corresponding to the adjusted first mask during playback. This makes it possible to avoid mask flickering due to changes in the mask's coverage area and improve the viewing experience of the video and bullet lines. [Brief explanation of the drawing]

[0011] To more clearly explain the technical solutions in the embodiments of this application, the drawings necessary for describing the embodiments or the prior art will be briefly described below. Naturally, the drawings in the following description are only some embodiments of this application, and those skilled in the art can obtain other drawings based on these without requiring any creative effort.

[0012] Further features, purposes, and advantages of this application will become clearer by referring to the detailed description provided for non-limiting embodiments with reference to the following drawings. [Figure 1] This is a schematic diagram of the process of a video playback method provided in one embodiment of this application. [Figure 2] This is a schematic diagram of the process of adding a mask to an object provided in one embodiment of this application. [Figure 3] This is a schematic diagram illustrating an example of adjusting the first mask provided in one embodiment of this application. [Figure 4]This is a schematic diagram of a single-frame optimization process for a first video frame provided in one embodiment of this application. [Figure 5] This is a schematic diagram of the structure of a video playback device provided in one embodiment of this application. [Figure 6] This is a schematic diagram of the structure of an electronic device applicable to realizing the solution in the embodiment of this application. The same or similar reference numerals in the drawing represent the same or similar components. [Modes for carrying out the invention]

[0013] To further clarify the purpose, technical solutions, and advantages of the embodiments of this application, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the drawings of the embodiments. Naturally, the embodiments described are only some, not all, embodiments of this application. All other embodiments obtained by a person skilled in the art based on the embodiments of this application without requiring any creative effort are all within the scope of protection of this application.

[0014] In a typical configuration of this application, each terminal and service network device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0015] Memory can take the form of non-persistent storage devices in computer-readable media, such as random-access memory (RAM) and / or non-volatile memory, for example, read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.

[0016] Computer-readable media include persistent and non-persistent, removable and non-removable media, and can store information by any means or technique. Information may be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random-access memory (SRAM), dynamic random-access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital purpose discs (DVDs) or other optical storage, magnetic cassette tapes, magnetic tape / disk storage or other magnetic storage devices, or any other non-transmission media, which can be used to store information accessible by computing devices.

[0017] As explained above, improving the user experience with bullet hell features is a noteworthy and urgent issue.

[0018] Some solutions involve adding masks to objects in the video (e.g., character objects) and using the masks to distinguish between objects and bullet patterns within a single video frame, thereby avoiding occlusion. For example, a mask can be added to objects within a video frame. The position and size of the mask can be determined based on the display area of ​​the object within the video frame, and then the mask can be added to the video frame accordingly. After the mask is added, when the video is displayed in that video frame, only objects within the mask may be displayed. In other words, if bullet patterns in the video move to the area corresponding to the mask, the bullet pattern content will not be displayed to the user within the area corresponding to the mask (e.g., the bullet patterns are occluded within the mask). This makes it possible to display, for example, "objects" in the "top layer" using masks, and thus the problem of bullet patterns occluding "objects" is solved. This achieves the goal of avoiding occlusion of video content by bullet patterns, maintains the availability of the bullet pattern function, and ensures a smooth viewing experience of the video content.

[0019] However, with this method, because video is a continuous process, if there is a significant change in the mask of an adjacent frame, the bullet hell will "suddenly appear" and "suddenly disappear." This situation is particularly noticeable in situations where there are misrecognitions or misdivisions of the mask's range and position. As a result, when users actually watch the video, the bullet hell may frequently "suddenly appear" and "suddenly disappear" (situations such as "bullet hell flickering" or "bullet hell flashing"), which can significantly impact the user's viewing experience.

[0020] In contrast, the embodiment of the present application provides a video playback method which acquires a first mask in a first video frame of a video to be played back, and in response to the fact that the first mask is used to independently display a first object and a second mask is present in a second video frame that is used simultaneously to independently display the first object, a first crossing overunion value between the first mask and the second mask is determined, and in response to the fact that the second video frame is the frame preceding the first video frame and the first crossing overunion value does not satisfy the requirements of a first ratio threshold, the first mask is adjusted using at least the second mask so that the adjusted first crossing overunion value satisfies the first ratio threshold, and the video to be played back is played back and the bullet lines in the video to be played back are displayed using an area other than the display area corresponding to the adjusted first mask during playback. This ensures the effectiveness of using the mask, avoids mask flickering due to changes in the mask's coverage area, and improves the viewing experience of the video and bullet lines.

[0021] In practice, the entity executing this method may be a user device, or a device in which a user device and a network device are integrated via a network, or an application program running on such a device. User devices include, but are not limited to, various terminal devices such as computers, mobile phones, tablet computers, smartwatches, and smart bands. Network devices include, but are not limited to, implementations such as network hosts, a single network server, multiple network server sets, or a collection of computers based on cloud computing, and can be used to implement some processing functions when setting alarms. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing, where cloud computing is a type of distributed computing and consists of a virtual computer made up of loosely coupled computers.

[0022] FIG. 1 shows a process 100 of a video playback method provided by an embodiment of the present application. The process 100 includes at least the following processing steps S101 to S108.

[0023] In step S101, a first mask in the first video frame of the video to be played is obtained.

[0024] In the embodiment of the present application, after obtaining the video to be played, for any frame in the video to be played (described as the first video frame for ease of understanding), the first mask included in the first video frame is read. As described above, the mask is used to independently display an object. Accordingly, for ease of explanation, the object independently displayed by the first mask may be referred to as the first object. In some embodiments, the video to be played may include a plurality of objects. For example, when the video to be played is an animation video, the object may be an animation character appearing in the animation video. It should be understood that the mask corresponds to the object (e.g., an animation character) one-to-one. For example, in the first video frame, the area indicated by the first mask A is used to display only the animation character A, and the area indicated by the first mask B is used to display only the animation character B.

[0025] In some embodiments, the mask (e.g., the first mask) included in the video frame (e.g., the first video frame) in the video to be played may be determined based on the recognition result of the object included in the video frame. For example, when it is determined that the video includes the animation characters A and B, the first mask A can be determined corresponding to the area for displaying the animation character A in the video frame, and the first mask B can be determined corresponding to the area for displaying the animation character B in the video frame.

[0026] It should be understood that in some videos to be played, if the first video frame does not contain a mask (e.g., the first mask), the video to be played may be played directly in existing ways. For example, when playing the first video frame, bullet patterns may be displayed based on an effect such as floating above an anime character. Such methods will not be discussed in detail here.

[0027] In some embodiments, the process of adding a mask to an object in the video to be played back can be seen in Figure 2. Figure 2 shows a masking process 200 provided in an embodiment of the present application, which includes at least the following processing steps S201 to S205.

[0028] In step S201, the video to be played is retrieved.

[0029] In step S202, a set of shortcut points in the video to be played back is determined based on the inter-frame image change rate algorithm.

[0030] Specifically, after acquiring the video to be played based on step S201 above, a set of shortcut points within the video to be played is determined based on an inter-frame image change rate algorithm. Shortcut points can be used, for example, to indicate points in the video to be played where the scene changes. For example, before time A in the video to be played, the content of the video to be played is that anime character A (hereinafter abbreviated as A) and anime character B (hereinafter abbreviated as B) are talking in a bedroom, and after time A, the content of the video to be played is that A, B, and anime character C (hereinafter abbreviated as C) are playing basketball in a gymnasium. Accordingly, time A can be determined as one of the shortcut points within the video to be played. In some embodiments, shortcut points in a video can be detected using an inter-frame image change rate algorithm, for example, implemented by the PySceneDetect algorithm library. In such a method, the location of a shortcut can be determined by analyzing the differences in images between adjacent video frames. This allows for efficient and accurate determination of shortcut point locations, for example, in animated videos, by using the PySceneDetect algorithm library to perform shortcuts based on the rate of change between video frames.

[0031] In step S203, at least one video frame sequence is segmented from the video to be played based on the shortcut point.

[0032] Specifically, the video shortcuts can be completed based on a set of shortcut points determined in step S202 above. For example, a video frame sequence can be obtained based on a set of video frames contained between two adjacent shortcut points. It should be understood that embodiments of this application can segment one or more video frame sequences from the video to be played back as needed. For example, if a mask is to be added only to the first scene in the video content, the video frame sequence for the first scene can be segmented based only on the start frame of the video to be played back and the first shortcut point determined based on the playback sequence. In some embodiments, it is usually possible to choose to use a set of shortcut points sequentially or all at once to complete the segmentation of the target video to be played back and obtain multiple consecutive video frame sequences.

[0033] In step S204, determine a set of objects that are included in at least one video frame sequence.

[0034] Specifically, the system recognizes a set of objects contained within at least one video frame sequence. In some embodiments, a segmentation model can be used to determine a set of objects contained within at least one video frame sequence. The segmentation model may be, for example, a Mask Recycle Convolutional Neural Network (abbreviated as Mask R-CNN). Typically, video frames within a video frame sequence can be processed individually; for example, Mask R-CNN can be used to segment objects (e.g., anime characters) within an anime video. In some embodiments, Mask R-CNN can be combined with other algorithms; for example, Mask R-CNN can be combined with the PointRend algorithm to directly segment objects and their bounding boxes, and the bounding box of the anime character can be used to determine the corresponding mask. This allows the segmentation model to automatically and accurately determine the image region where an object is located within a video frame.

[0035] In some embodiments, the segmentation model is obtained by using an entity segmentation model trained on non-animated data as a pre-training model, training it with animated data as input to the pre-training model and animated images containing mask annotations of virtual characters as outputs to the pre-training model. Specifically, for data training, an entity segmentation model trained on non-animated data (e.g., real-world physical scene images, video frames) is used as a pre-training model, and then the pre-training model is adjusted and trained using animated data as sample input data and animated images containing annotations of the location regions of virtual characters (e.g., animated characters) as sample output data. This allows the pre-training model to have the ability to segment animated characters in an animated scene. For example, a trained segmentation model can recognize animated characters within video frames of an animated video. In some embodiments, masks can also be directly indicated by annotations, allowing the segmentation model to directly output masks corresponding to animated characters. This allows us to obtain an entity segmentation model (or segmentation model) that has the ability to segment anime characters by adjusting an entity segmentation model trained on non-animated data, thereby reducing the difficulty of training the model.

[0036] In step S205, a mask is added for each object based on the tracking of each object in a set of objects.

[0037] Specifically, after determining a set of objects based on step S204 above, the position of each object within each image frame can be determined for each object in the set using an object tracking method. For example, objects within a video frame sequence can be tracked using the DeepSort algorithm, a multi-target tracking algorithm, and the tracking objective is achieved by performing inter-frame matching on the objects within the video frame sequence. In some embodiments, to facilitate the representation of objects, a corresponding unique identifier (Identity Document, abbreviated as ID) can be added to the tracking trajectory of each object for identification. Taking one object (e.g., A) as an example, by tracking the position of A within a video frame sequence, the image region corresponding to A within each video frame in the video frame sequence can be determined. Furthermore, a mask can be added for A based on the image region corresponding to A within each video frame. For example, in one video frame, the area covered by the mask coincides with the image region where A is located. This allows for the efficient and accurate determination of the corresponding mask for each object within a video frame using an object tracking method. It's important to understand that multiple objects can exist within a single video frame, and accordingly, multiple masks can exist simultaneously within that video frame, each using a different mask to display the corresponding object.

[0038] This allows for a reduction in the number of video frames that need to be processed each time, for example, by segmenting the video to be played into at least one video frame sequence based on a scene shortcut point, and then segmenting the video to be played. Furthermore, such a method can reduce the difficulty of tracking objects by avoiding situations such as loss or untracking of objects due to scene transitions, for example, by associating video content.

[0039] In some embodiments, the use of masks can be further restricted by setting constraints, balancing the needs for displaying video content and bullet patterns. For example, for some objects that appear relatively infrequently or for short periods, their masks can be removed to free up space and satisfy the bullet pattern display needs. In some embodiments, for an object in a set of objects within a video frame sequence (for ease of explanation, the object currently being processed will be referred to as the target object), the number of video frames in the video frame sequence that contain the target mask of the target object can be obtained. Furthermore, if the number of such video frames is below a predetermined threshold, the space for displaying bullet patterns can be expanded by choosing to remove the target mask added to the target object. In some embodiments, the threshold may be determined based on the frequency, proportion, and fulfillment of the needs. For example, if an object appears less than 50%, it is determined to be a low-frequency object, and in this case, the threshold may be determined based on the total number of video frames in the video frame sequence and 50%. For example, if the total number of video frames is 10, the threshold may be 5 frames. In this case, if the number of frames in which the target object appears is 5 or less, more bullet patterns can be displayed by freeing up the space occupied by the target object mask. This expands the display space for the bullet patterns and prevents the display space from being too small, which would negatively impact the bullet pattern experience. Similarly, a corresponding numerical threshold can be determined based on the conversion of the duration corresponding to the shortest appearance time to video frames, but this will not be explained again here.

[0040] In step S102, in response to the fact that a second mask for independently displaying the first object exists simultaneously within the second video frame, a first crossover overunion value between the first mask and the second mask is determined.

[0041] In the embodiments of this application, after obtaining a first mask, it is possible to detect whether a second mask exists for independently displaying a first object in a second video frame that is before the first video frame and adjacent to the first video frame (i.e., the second video frame is the frame preceding the first video frame).

[0042] If present, the intersection overunion value between the first and second masks (referred to as the first intersection overunion value for ease of understanding) is determined in response. The intersection overunion value may also be called the intersection overunion (IoU), and it can be determined based on the ratio of the intersection area to the union area of ​​the first and second masks. Generally, a numerical value of the intersection overunion value closer to 0 indicates a larger difference between the first and second masks (e.g., at least one of the area difference and / or position difference). Conversely, a numerical value of the intersection overunion value closer to 1 indicates a smaller difference between the first and second masks.

[0043] In some embodiments, after extracting the first and second video frames respectively, the first and second video frames can be aligned on the same plane, that is, the screen centers of the first and second video frames can be set to the same point on the same plane. Furthermore, based on the aligned first and second video frames, the intersection area and the combined area of ​​the first and second masks can be obtained to obtain the intersection overunion. It should be understood that in actual operation, the first and second video frames belong to the same video to be played, so in normal cases their sizes match, and this prevents sudden distortion when playing the video. Accordingly, alignment can be achieved by adding the content of the first video frame to the second video frame, or by adding the content of the second video frame to the first video frame.

[0044] For easier understanding, Figure 3 can also be referenced. Figure 3 shows a schematic diagram of Example 300 of adjusting the first mask provided in the embodiments of this application.

[0045] In Example 300, a first video frame 310 and a second video frame 320 are included, where the second video frame 320 is the frame preceding the first video frame 310 in the video being played. The first video frame 310 contains a first mask 311 for displaying the first object 330, and the second video frame 320 contains a second mask 321 for displaying the first object 330. For ease of understanding, the sizes of the first video frame 310 and the second video frame 320 are shown as identical.

[0046] Alignment can be achieved by adding the second mask 321 in the second video frame 320 to the first video frame 310. For example, in Example 300, by adding the second mask 321 to the first video frame 310, the first video frame 310' can be obtained, and alignment between the first video frame 310 and the second video frame 320 can be achieved.

[0047] In the first video frame 310', the intersection region of the first mask 311 and the second mask 321 (e.g., the area of ​​the first mask 311) and the combined region of the first mask 311 and the second mask 321 (i.e., the area of ​​the second mask 321) are determined. Furthermore, a first cross-overunion value can be determined based on the ratio of the intersection region to the combined region.

[0048] In step S103, it is determined whether the first cross-overunion value satisfies the requirements of the first ratio threshold.

[0049] In the embodiments of this application, after determining the first crossover union value in step S103, it is possible to determine whether the first crossover union value satisfies the requirements of the first ratio threshold based on a comparison with the first ratio threshold. In some embodiments, the first ratio threshold may be determined based on the allowable deviation (i.e., difference) between the first mask and the second mask, for example, the first ratio threshold may be 0.3, 0.5, 0.7, etc.

[0050] Furthermore, if it is determined that the first cross-overunion value does not meet the requirements of the first ratio threshold, step S104 can be performed.

[0051] In step S104, the first mask is adjusted using at least the second mask so that the adjusted first cross-overunion value satisfies the first ratio threshold. In embodiments of this application, if the first cross-overunion value does not meet the requirements of the first ratio threshold, the first cross-overunion value can be improved by adjusting the first mask using the second mask. For example, the first cross-overunion value can be improved by adjusting the position of the first mask by moving the first mask to increase the overlap area between the first and second masks. In some embodiments, the first cross-overunion value can also be improved by scaling the first or second mask by a 1:1 ratio. In some embodiments, the object can be further divided based on the type of the first object. For example, if the first object is an anime character, the area occupied by the first and / or second mask can be reduced by dividing it according to the face area, body area, etc., and then deleting part of the division result. This improves the first cross-overunion value by reducing the combined area.

[0052] For illustrative purposes, you can continue to refer to Figure 3. If it is determined that the first crossover overunion value of the first mask 311 and the second mask 321 does not meet the requirements of the first ratio threshold, the first ratio threshold can be improved by adjusting the first mask 311 using the second mask 321.

[0053] For example, the region indicated by the first mask 311 can be directly adjusted to the region indicated by the second mask 321. For example, the second mask can be scaled down by a 1:1 ratio based on a value indicated by the first ratio threshold. For example, by scaling down the second mask 312 by a 1:1 ratio, the minimum area region 341 that satisfies the indication of the first ratio threshold can be obtained, and the first mask 311 can be adjusted according to this minimum area 341. It should be understood that when determining the adjustment target of the first mask 311 by scaling down the second mask 312 by a 1:1 ratio, the display position of the adjusted first mask 311 can be determined by fixing the center of the second mask 312 and then choosing to align the center of the minimum area region 341 with the center of the second mask 312, i.e., the adjusted first mask 311 and second mask 312 are actually two concentric regions.

[0054] In some embodiments, if a third mask for independently displaying a first object exists simultaneously in the frame following the first video frame (for ease of understanding, the frame following the first video frame will be referred to as the third video frame), the second cross-overunion value can be determined based on the second and third masks. If the second cross-overunion value satisfies the requirements of the first ratio threshold, the second and third masks can be combined, and the first mask can be adjusted using the result of combining the second and third masks. That is, the area indicated by the first mask is adjusted to the area indicated by the result of combining the second and third masks. This allows the first mask to be adjusted using the second and third masks if the difference between the masks added to the first object in the frame before and after the first video frame (i.e., the second and third masks) satisfies the requirements. This ensures that the areas corresponding to the masks in each video frame are approximately the same, avoiding bullet barrage flickering due to abrupt changes in the mask area.

[0055] In some embodiments, a similar method can be used to sequentially acquire the fourth mask in the fourth video frame after the third video frame, the fifth mask in the fifth video frame after the fourth video frame, and so on. Furthermore, based on whether the cross-over union value determined by the second and third masks, the fourth and fifth masks, etc., satisfies the first ratio threshold, it is possible to decide whether or not to adjust the first mask using the result of combining the second and third masks. This allows for the determination of whether the difference between the first and second masks is due to a situation such as a recognition error, depending on the mask changes for the same object in multiple consecutive video frames. This enables more accurate recognition of issues such as "bullet hell flickering" caused by inaccurate recognition, and allows for more precise and high-quality adjustment of the first mask.

[0056] Furthermore, after adjusting the first mask so that the adjusted first crossover union value satisfies the first ratio threshold, step S105 can be performed again.

[0057] In step S105, the target video is played, and the bullet lines within the target video are displayed using the area other than the display area corresponding to the first mask that has been adjusted during playback.

[0058] In the embodiments of this application, after the adjustment of the first mask is completed, the video to be played is played (for example, the video to be played is played to the user at the user's request), and during playback, the bullet lines within the video to be played can be displayed using an area other than the display area corresponding to the adjusted first mask (i.e., the area indicated by the first mask).

[0059] In some embodiments, if it is determined in step S103 that the first cross-over union value satisfies the requirements of the first ratio threshold, step S106 can be performed directly. In step S106, the video to be played is played, and during playback, the bullet lines within the video to be played are displayed using areas other than the display area corresponding to the first mask.

[0060] In some embodiments, the range of difference between the first and second masks can be further constrained by setting a second ratio threshold that has a smaller value than the first ratio threshold. This prevents "gaps" in the bullet pattern from occurring within the video frame due to incorrect adjustment of the position of the first mask when, for example, the first object experiences situations such as "teleportation" or "disappearance." In this case, if the first crossover union value is less than the second ratio threshold, the first and second masks can be directly selected to solve the problem of bullet pattern flickering caused by too large a difference between the first and second masks.

[0061] Specifically, for example, if it is determined in step S103 that the first cross-overunion value does not meet the requirements of the first ratio threshold, then step S107 can be performed.

[0062] In step S107, it is determined whether the first crossover overunion value is less than the second ratio threshold. If it is determined in step S107 that the first crossover overunion value is greater than or equal to the second ratio threshold, the execution of step S104 is selected.

[0063] If it is determined in step S107 that the first crossover overunion value is less than the second ratio threshold, then step S108 is executed.

[0064] In step S108, at least the first mask is deleted. This avoids bullet pattern flickering caused by switching from the second mask to the first mask, which has too large a difference from the second mask. In some embodiments, the first and second masks may be deleted simultaneously, which allows the bullet patterns to be displayed more quickly and continuously, further preventing bullet pattern flickering.

[0065] In some embodiments, if a second mask exists, it is possible to directly determine whether the first mask is available based on an area comparison between the first and second masks. For example, after obtaining the first and second masks, the difference between the area of ​​the first mask and the area of ​​the second mask can be determined, and then the area difference rate can be determined based on the ratio of this difference to the area of ​​the second mask. If this area difference rate exceeds a predetermined difference threshold (e.g., 0.4), it can be determined that the first mask has a larger area change compared to the second mask. Accordingly, bullet flickering can be avoided by selecting to delete the first mask or to delete both the first and second masks.

[0066] It's important to understand that even if a mask (for example, at least one of the first and second masks) is removed, the video can still be played. Accordingly, in this case, when playing the video, the position corresponding to the removed mask can be used to display the commentary, which will not be explained in detail here.

[0067] In some embodiments, if the first mask is to be removed, and further, by removing the first mask within a set of video frames within a certain range of the first video frame, the bullet pattern can be moved and displayed continuously within these video frames (the first video frame and the set of video frames within a certain range of the first video frame) without being obstructed. This improves the display effect of the bullet pattern. In some embodiments, based on the positional interval of the first video frame in the video to be played, and, A first decision strategy instructs the system to determine a set of adjustment frames based on each frame from the start frame of the video to be played back to the first video frame, A second decision strategy instructs the system to determine a set of adjustment frames based on each frame from the second video frame to the end frame of the video being played back, A third decision strategy instructs the system to determine a set of adjustment frames based on each frame included in the position interval, and based on one of these decision strategies, a set of adjustment video frames can be determined, and a mask added to the first object within each video frame of the set of adjustment video frames can be removed.

[0068] Specifically, a first position interval can be determined using a predetermined frame interval, starting from the beginning frame of the video to be played. In some embodiments, the frame interval may be determined in accordance with the total number of frames in the video to be played. For example, the frame interval may be determined based on the product of 0.1, 0.2, or 0.3 and the total number of frames. This allows the bullet hell to be displayed more continuously by removing the first mask for the first object in each video frame from the beginning frame to the first video frame, according to the first determination strategy, when the first video frame is located within the first position interval. In some embodiments, when objects are marked using IDs, the masks can be removed all at once by determining the mask corresponding to the ID in each video frame.

[0069] Similarly, the second position interval can be determined by using the end frame of the video to be played as the endpoint and calculating backward based on the number of video frames in the above frame interval to determine the starting point. As a result, if the first video frame is located within the second position interval, the second decision strategy can remove the first mask for the first object in each video frame from the first video frame to the end frame, thereby displaying the bullet patterns more continuously.

[0070] Similarly, any position within the video to be played, other than the first and second position intervals, can be determined as a third position interval. For example, if the first video frame is located within the third position interval, the third decision strategy can be used to remove the first mask for the first object within each video frame located within the third position interval, thereby allowing the bullet patterns to be displayed more continuously.

[0071] In step S106, the target video is played, and the bullet lines within the target video are displayed using the area other than the display area corresponding to the first mask that has been adjusted during playback.

[0072] In some embodiments, if it is determined that a second mask for independently displaying the first object does not simultaneously exist within the second video frame, the video to be played can be played in response without adjusting the first mask. In this case, during playback of the video to be played, when the position within the first video frame is reached, the bullet lines within the video to be played can be displayed using an area other than the display area corresponding to the first mask.

[0073] Furthermore, in some embodiments, the quality of mask usage can be further improved by single-frame optimization of the video frame. For example, after adding a mask based on the process 200 described above, the mask can be optimized (e.g., adjusted or deleted) as follows. This can improve the display effect by using single-frame optimization sequentially or by utilizing at least the previous frame. For ease of understanding, this will be illustrated using the first video frame as an example.

[0074] Figure 4 can be further referenced for this. Figure 4 shows a single-frame optimization process 400 for a first video frame provided in an embodiment of the present application, the process 400 including at least the following processing steps S401 to S404.

[0075] In step S401, it is determined whether or not a fourth mask exists within the first video frame.

[0076] Specifically, it is possible to determine whether there are further masks corresponding to other objects (which may be referred to as fourth masks for ease of explanation) within the first video frame, and the first and fourth masks may correspond to different objects (which may be referred to as second objects for ease of explanation). For example, the first mask corresponds to object A in the first video frame, and the fourth mask corresponds to object B in the first video frame. If there is at least one further fourth mask within the first video frame, step S402 is performed.

[0077] In step S402, the total area of ​​the mask is determined based on the first and fourth masks.

[0078] Specifically, if a fourth mask exists, the total area of ​​the mask can be determined based on the sum of the area indicated by the first mask and the area indicated by the fourth mask. In some embodiments, if multiple fourth masks exist, the determination of the total area of ​​the mask may be based on the sum of the area indicated by the first mask and the areas of all the fourth masks.

[0079] In step S403, it is determined whether the total area of ​​the mask exceeds the area threshold.

[0080] Specifically, it is determined whether the total area of ​​the mask exceeds an area threshold. In some embodiments, the area threshold may be determined based on the minimum area required to display the bullet pattern, thereby ensuring the normal display of the bullet pattern by limiting the total area of ​​the mask by the area threshold. For example, the area threshold may be 65% of the display area of ​​the first video frame. If the total area of ​​the mask exceeds the area threshold, step S404 is performed.

[0081] In Step S404, the first or fourth mask is held.

[0082] Specifically, the decision to retain the first or fourth mask can be made based on a comparison of the areas of the first or fourth mask, the importance of the type of object being addressed (for example, a mask for a face object may have a higher priority), and the mask's position relative to the video frame. For example, an object corresponding to a mask with a larger occupied area is more likely to be the "main object" in the first video frame. In this case, the mask for the "main object" can be retained by choosing to retain the one with the larger corresponding area among the first and fourth masks based on the area comparison. This reduces the overall mask area occupied by retaining either the first or fourth mask.

[0083] In some embodiments, a first comparison result can be generated that indicates which of the first and fourth masks has a larger area, based on an area comparison between the first and fourth masks. Furthermore, based on the first comparison result, only the mask with the larger area between the first and fourth masks is retained. This allows for the retention of the mask that occupies more of the content of the first video frame and is therefore more likely to be the "main character" object, based on the area comparison method.

[0084] In some embodiments, if at least one of the first and second objects includes at least one face object, a second comparison result can be generated that indicates which of the first and fourth masks has a larger area of ​​the associated face region, based on an area comparison of the face regions associated with the first and fourth masks. Furthermore, based on the second comparison result, only the mask with the larger area of ​​the associated face region is retained. This preferentially retains masks that have associated face regions and masks with larger areas of the associated face region, thereby retaining masks that occupy more of the content of the first video frame and are more likely to be the "main character" object.

[0085] In some embodiments, after obtaining a first position where the region center of the first mask is located and a second position where the region center of the fourth mask is located, a third comparison result can be generated that indicates which of the first and fourth masks has a region center closer to the screen center, based on a comparison of a first distance from the first position to the screen center and a second distance from the second position to the screen center. Furthermore, based on the third comparison result, only the mask whose region center is closer to the screen center is retained. This prioritizes retaining masks that occupy more of the content of the first video frame and are more likely to be "main character" objects, based on the position in which the object occupies the screen.

[0086] In some embodiments, after adding a mask to an object within a single frame, it is possible to determine whether the mask needs to be directly adjusted or deleted based on the descriptive information of the newly added mask. For example, if the determined mask occupies more than 80% of the display area of ​​the first video frame, its occupied area can be reduced, for example, by scaling it down to its original size and deleting the mask in areas other than the face. Alternatively, in some embodiments, if the determined mask occupies less than 1% of the display area of ​​the first video frame, it can be chosen to delete the mask. Similarly, if the mask is divisible, for example, if the mask substantially consists of a first submask for the upper body of an animated character and a second submask for the lower body of an animated character, it is possible to determine whether the first and second submasks need to be adjusted or deleted by setting corresponding area thresholds for the first and second submasks. For example, if the first submask occupies more than 80% of the total mask area, the visual effect can be improved by choosing to reduce the area of ​​the first submask or delete it. Similarly, if a second submask occupies more than 90% of the total mask area, you may choose to reduce the area of ​​the second submask in a similar manner or delete the mask. Similarly, if the mask includes a face area, you may choose to further adjust or delete the mask based on constraints corresponding to the area occupied by the face area. For example, if the height of the bbox of the face area is greater than 0.5 or the width of the bbox is greater than 0.4, adjust or delete the mask. Similarly, you may choose to adjust or delete the mask based on the distance from the center position of the mask to the edge of the first video frame. For example, if the center position of the mask is 25% of the way from the top or bottom of the first video frame, or 10% of the way from the left or right of the first video frame, adjust or delete the mask.

[0087] It should be understood that, in order for the adjusted mask (or sub-regions, face regions included in the mask) to satisfy the corresponding constraints, the adjustments in each of the above examples can be made by scaling or displacing at a 1:1 ratio, but are not limited to these methods, and will not be explained again here.

[0088] Next, the video playback method provided in this application obtains a first mask in a first video frame of the video to be played back, and in response to the fact that the first mask is used to display a first object independently and a second mask is present in a second video frame that is used simultaneously to display the first object independently, a first crossing overunion value between the first mask and the second mask is determined, and in response to the fact that the second video frame is the frame before the first video frame and the first crossing overunion value does not satisfy the requirements of a first ratio threshold, the first mask is adjusted using at least the second mask so that the adjusted first crossing overunion value satisfies the first ratio threshold, and when playing back the video to be played back, the bullet lines in the video to be played back are displayed using an area other than the display area corresponding to the adjusted first mask. This ensures the effectiveness of using the mask, avoids mask flickering due to changes in the mask's coverage area, and improves the viewing experience of the video and bullet lines.

[0089] Embodiments of the present application further provide a device for video playback, the structure of which is as shown in the device 500 in Figure 5. The device 500 includes an acquisition module 510 configured to acquire a first mask in a first video frame of a video to be played back, the acquisition module 510 being used to display a first object independently; a first determination module 520 configured to determine a first cross overunion value between the first mask and the second mask in response to the presence of a second mask in a second video frame that is simultaneously used to display the first object independently, the first determination module 520 being the frame preceding the first video frame; an update module 530 configured to adjust the first mask using at least the second mask in response to the first cross overunion value not satisfying the requirements of a first ratio threshold, such that the adjusted first cross overunion value satisfies the first ratio threshold; and a first playback module 540 configured to play back the video and to display a barrage of text in the video to be played back using an area other than the display area corresponding to the adjusted first mask during playback.

[0090] In some embodiments, the apparatus 500 further includes a first deletion module configured to delete at least a first mask in response to a first crossover union value being less than a second ratio threshold, where the numerical value of the second ratio threshold is less than the first ratio threshold.

[0091] In some embodiments, the apparatus 500 further includes a second determination module configured to determine a second cross-overunion value based on the second and third masks in response to the simultaneous presence of a third mask for independently displaying a first object within a third video frame, to combine the second and third masks in response to the second cross-overunion value satisfying the requirements of a first ratio threshold, and to adjust the first mask using at least the second mask, wherein adjusting the first mask using at least the second mask includes adjusting the first mask using the result of combining the second and third masks.

[0092] In some embodiments, the apparatus 500 further includes a masking module configured to acquire a video to be played back, determine a set of shortcut points in the video to be played back based on an interframe image change rate algorithm, segment at least one video frame sequence from the video to be played back based on the shortcut points, determine a set of objects contained in at least one video frame sequence, and add a mask corresponding to each object based on tracking each object of the set of objects, the coverage area of ​​the mask being determined based on the image region where the corresponding object is located.

[0093] In some embodiments, determining a set of objects contained in at least one video frame sequence involves determining a set of objects contained in at least one video frame sequence by processing at least one video frame sequence using a segmentation model.

[0094] In some embodiments, the apparatus 500 further includes a second delete module configured to obtain the number of video frames containing the target mask of a target object within at least one video frame sequence, and to delete the target mask added to the target object in response to the number of video frames being less than or equal to a certain threshold.

[0095] In some embodiments, the video to be played includes animated video, and the segmentation model is obtained by using an entity segmentation model trained on non-animated data as a pre-trained model, with animated data as input to the pre-trained model and animated images including mask annotations of virtual characters as the output of the pre-trained model.

[0096] In some embodiments, the device 500 further includes a third delete module configured to determine a set of video frames to be adjusted based on one of the following decision strategies: a first decision strategy which instructs the device to determine a set of frames to be adjusted based on a position interval of a first video frame in the video to be played, and on each frame from the start frame to the first video frame of the video to be played; a second decision strategy which instructs the device to determine a set of frames to be adjusted based on each frame from the second video frame to the end frame of the video to be played; and a third decision strategy which instructs the device to determine a set of frames to be adjusted based on each frame included in the position interval, and to delete a mask that has been added to a first object in each video frame of the set of video frames to be adjusted.

[0097] In some embodiments, the apparatus 500 further includes a mask holding module configured to determine the total area of ​​the masks based on the first and fourth masks in response to the presence of at least one fourth mask in the first video frame for independently displaying a second object, and to hold the first or fourth mask in response to the total area of ​​the masks exceeding an area threshold.

[0098] In some embodiments, retaining a first mask or a fourth mask includes generating a first comparison result indicating which of the first and fourth masks has a larger area, based on an area comparison between the first and fourth masks, and retaining only the one of the first and fourth masks that has a larger area, based on the first comparison result.

[0099] In some embodiments, at least one of the first object and the second object includes at least a face object, and holding a first mask or a fourth mask includes generating a second comparison result indicating which of the first and fourth masks has a larger area of ​​the relevant face region, based on a comparison of the areas of the face regions associated with the first and fourth masks, and holding only the first and fourth masks with a larger area of ​​the relevant face region, based on the second comparison result.

[0100] In some embodiments, retaining a first mask or a fourth mask includes obtaining a first position where the region center of the first mask is located and a second position where the region center of the fourth mask is located; generating a third comparison result that indicates which of the first and fourth masks has a region center closer to the screen center, based on a comparison of a first distance from the first position to the screen center and a second distance from the second position to the screen center; and retaining only the one of the first and fourth masks whose region center is closer to the screen center, based on the third comparison result.

[0101] In some embodiments, the device 500 further includes a second playback module configured to play a video to be played in response to a first cross-over union value satisfying the requirements of a first ratio threshold, and to display a barrage of text within the video to be played using an area other than the display area corresponding to the first mask during playback.

[0102] Based on the same inventive concept, embodiments of the present application further provide an electronic device, the method corresponding to the electronic device may be the video playback method in the above embodiments, and the problem-solving principle thereof is similar to that method. The electronic device provided in embodiments of the present application includes at least one processor and a memory communicably connected to the at least one processor, wherein the memory stores computer-readable instructions that can be executed by the at least one processor, and when the computer-readable instructions are executed by the at least one processor, the at least one processor can execute the methods and / or technical solutions of the above embodiments of the present application.

[0103] Electronic devices may be user devices, or devices in which user devices and network devices are integrated via a network, or application programs running on such devices. User devices include, but are not limited to, various terminal devices such as computers, mobile phones, tablet computers, smartwatches, and smart bands. Network devices include, but are not limited to, implementations such as network hosts, a single network server, multiple sets of network servers, or a collection of computers based on cloud computing, and may be used to implement some processing functions when setting alarms. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing, where cloud computing is a type of distributed computing and consists of virtual computers made up of loosely coupled computers.

[0104] Figure 6 shows the structure of an electronic device applied to implement the method and / or technical solution in an embodiment of this application, the electronic device 600 including a central processing unit (CPU) 601 capable of performing various appropriate operations and processes based on a program stored in a read-only memory (ROM) 602 or a program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 further stores various programs and data necessary for system operation. The CPU 601, ROM 602, and RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0105] The components of the input unit 606, which includes a keyboard, mouse, touchscreen, microphone, and infrared sensor; the output unit 607, which includes a cathode ray tube (CRT), liquid crystal display (LCD), LED display, OLED display, and speaker; the storage unit 608, which includes one or more computer-readable media such as a hard disk, optical disk, magnetic disk, and semiconductor memory; and the communication unit 609, which includes a network interface card such as a LAN (Local Area Network) card and modem, are all connected to the I / O interface 605. The communication unit 609 performs communication processing via a network such as the Internet.

[0106] In particular, the methods and / or embodiments in the embodiments of this application may be implemented as computer software programs. For example, the embodiments disclosed in this application include a computer program product including a computer program mounted on a computer-readable medium, the computer program including program code for performing the methods shown in the flowchart. When the computer program is executed by a central processing unit (CPU) 601, the above-described functions limited by the methods of this application are performed.

[0107] In another embodiment of the present application, a computer-readable storage medium is provided which stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the methods and / or technical solutions of any one or more embodiments of the present application described above are realized.

[0108] Specifically, this embodiment may employ any combination of one or more computer-readable media. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exclusive list) include electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this specification, the computer-readable storage medium may be any tangible medium containing or storing a program, which can be used in or in combination with a computer-readable instruction execution system, apparatus, or device.

[0109] A computer-readable signal medium may contain data signals propagated within the baseband or as part of a carrier wave, which may contain computer-readable program code. The data signals propagated in this manner may take various forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium may further be any computer-readable medium other than a computer-readable storage medium, which can transmit, propagate, or transmit programs used in or in combination with computer-readable instruction execution systems, apparatus, or devices.

[0110] Program code contained in a computer-readable medium can be transmitted through any suitable medium, including but not limited to wireless, wire, optical cable, RF, or any suitable combination thereof.

[0111] The computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, and the programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as general procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user computer, partially on the user computer, as an independent software package, partially on the user computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user computer by any network, including a local area network (LAN) or wide area network (WAN), or it can be connected to an external computer (for example, via the Internet using an Internet service provider).

[0112] The flowcharts and block diagrams in the drawings illustrate the implementable system architectures, functions, and operations of devices, methods, and computer program products relating to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, program segment, or part of code, which contains one or more executable computer-readable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions represented in a block may be implemented in a different order than those shown in the drawings. For example, two consecutive blocks may be executed substantially simultaneously, or, depending on the function, they may be executed in reverse order. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, may be implemented by a dedicated system based on hardware that performs the specified function or operation, or by a combination of dedicated hardware and computer-readable instructions.

[0113] Those skilled in the art will clearly understand that, in order to simplify and streamline the explanation, the specific operating processes of the systems, apparatus, and units described above can be found by referring to the corresponding processes in the embodiments of the methods described above, and therefore, a detailed explanation is omitted here.

[0114] It should be understood that, in some embodiments provided in this application, the disclosed systems, apparatus and methods can be implemented in other forms. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is merely a division of logical functions and may be divided in other forms in actual implementation; for example, multiple units or page components may be combined or integrated into another system; or some features may be omitted or not performed. On the other hand, the mutual coupling, direct coupling or communication connection shown or described may be an indirect coupling or communication connection via some interface, apparatus or unit, and may be electrical, mechanical or other in nature.

[0115] The units described as separating members may or may not be physically separated, and the members shown as units may or may not be physical units, may be located in one place, or may be distributed among multiple network units. Some or all of these units can be selected according to actual needs to achieve the objective of the solution of this embodiment.

[0116] Furthermore, each functional unit in each embodiment of this application may be integrated into a single processing unit, may exist physically independently, or may be integrated into a single unit of two or more units. The integrated unit may be implemented in hardware form, or in the form of a hardware and software functional unit.

[0117] The integrated unit, implemented in the form of the software function unit described above, may be stored in a single computer-readable storage medium. The software function unit stored in the storage medium contains a plurality of computer-readable instructions for causing a single computer device (which may be a personal computer, server, or network device, etc.) or processor to perform some of the steps of the methods of each embodiment of this application. The storage medium described above includes various media capable of storing program code, such as USB memory, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0118] Finally, it should be noted that the above embodiments are merely for illustrating, and not limiting, the technical solutions of this application. Although this application has been described in detail with reference to the above embodiments, it will be understood by those skilled in the art that the technical solutions described in each of the above embodiments can be modified or some of their technical features can be replaced with equivalent ones, and such modifications or equivalent replacements will not cause the essence of the relevant technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of this application.

[0119] Furthermore, the term "includes" clearly does not exclude other units or steps, and the singular form does not exclude the plural form. Multiple units or devices described in the device claims may be implemented by a single unit or device using software or hardware. Terms such as "first," "second," etc., are used for naming purposes and do not indicate any particular order.

Claims

1. A step of obtaining a first mask within a first video frame of a video to be played back, wherein the first mask is used to independently display a first object; A step of determining a first crossover overunion value of the first mask and the second mask in response to the presence of a second mask in a second video frame that is used simultaneously to display the first object independently, the step of the second video frame being the frame preceding the first video frame, In response to the first cross-overunion value not satisfying the requirements of the first ratio threshold, the steps include adjusting the first mask using at least the second mask so that the adjusted first cross-overunion value satisfies the first ratio threshold, A video playback method comprising the steps of: playing the video to be played, and displaying a barrage of text within the video to be played using an area other than the display area corresponding to the first mask adjusted during playback.

2. The method according to claim 1, further comprising the step of removing at least the first mask in response that the first crossover-union value is less than a second ratio threshold, wherein the numerical value of the second ratio threshold is less than the first ratio threshold.

3. In response to the simultaneous presence of a third mask for independently displaying the first object within a third video frame, the steps include determining a second cross-overunion value based on the second mask and the third mask, The method further includes the step of joining the second mask and the third mask in response to the second cross-overunion value satisfying the requirements of the first ratio threshold, The step of adjusting the first mask using at least the second mask is, The method according to claim 1, comprising adjusting the first mask using the result of combining the second mask and the third mask.

4. The steps include: obtaining the video to be played back, The steps include determining a set of shortcut points in the video to be played back based on an inter-frame image change rate algorithm, The steps include segmenting at least one video frame sequence from the video to be played back based on the aforementioned shortcut point, The steps include determining a set of objects included in at least one video frame sequence, The method according to claim 1, further comprising the steps of adding a mask corresponding to each object based on tracking each object of the set of objects, wherein the coverage area of ​​the mask is determined based on the image area in which the corresponding object is located.

5. The step of determining a set of objects included in at least one video frame sequence is: The method according to claim 4, comprising determining a set of objects contained in the at least one video frame sequence by processing the at least one video frame sequence using a segmentation model.

6. The steps include obtaining the number of video frames in the at least one video frame sequence that contain the target mask of the target object, The method according to claim 4, further comprising the step of removing a target mask added to the target object in response that the number of video frames is below a certain threshold.

7. The method according to claim 5, wherein the video to be played includes an animated video, the segmentation model is obtained by using an entity segmentation model trained on non-animated data as a pre-trained model, using animated data as input to the pre-trained model, and training an animated image including mask annotations of virtual characters as output to the pre-trained model.

8. A step of determining a set of video frames to be adjusted based on one of the following decision strategies: a first decision strategy that instructs to determine a set of adjustment target frames based on the position interval of the first video frame in the video to be played and on each frame from the start frame to the first video frame of the video to be played; a second decision strategy that instructs to determine a set of adjustment target frames based on each frame from the second video frame to the end frame of the video to be played; and a third decision strategy that instructs to determine a set of adjustment target frames based on each frame included in the position interval. The method according to claim 2, further comprising the step of removing a mask that has been added to the first object in each video frame of the set of video frames to be adjusted.

9. A step of determining the total area of ​​a mask based on the first mask and the fourth mask in response to the presence of a fourth mask within the first video frame, wherein the fourth mask is used to independently display a second object. The method according to claim 1, further comprising the step of holding the first mask or the fourth mask in response to the total area of ​​the masks exceeding an area threshold.

10. The step of holding the first mask or the fourth mask is: Based on the area comparison between the first mask and the fourth mask, a first comparison result is generated that indicates the mask with the larger area among the first mask and the fourth mask. The method according to claim 9, comprising, based on the first comparison result, retaining only the one of the first mask and the fourth mask that has a larger area.

11. The step of having at least one of the first object and the second object include at least one face object and holding the first mask or the fourth mask, Based on a comparison of the area of ​​the face regions associated with the first mask and the fourth mask, a second comparison result is generated that indicates the one of the first mask and the fourth mask with the larger area of ​​the associated face region. The method according to claim 9, comprising, based on the second comparison result, retaining only the one of the first mask and the fourth mask that has a larger area of ​​the relevant facial region.

12. The step of holding the first mask or the fourth mask is: To obtain the first position where the center of the region of the first mask is located, and the second position where the center of the region of the fourth mask is located, Based on a comparison of a first distance from the first position to the screen center position and a second distance from the second position to the screen center position, a third comparison result is generated that indicates which of the first mask and the fourth mask has a region center closer to the screen center. The method according to claim 9, comprising, based on the third comparison result, retaining only the one of the first mask and the fourth mask whose region center is closer to the screen center.

13. The method according to claim 1, further comprising the steps of playing the video to be played in response to the first cross-over union value satisfying the requirements of the first ratio threshold, and displaying a barrage of text in the video to be played using an area other than the display area corresponding to the first mask during playback.

14. An acquisition module configured to acquire a first mask within a first video frame of a video to be played back, wherein the first mask is used to independently display a first object, A first determination module configured to determine a first crossover overunion value between the first mask and the second mask in response to the presence of a second mask in a second video frame used simultaneously to independently display the first object, wherein the second video frame is the frame preceding the first video frame, An update module configured to adjust the first mask using at least the second mask so that the adjusted first cross-overunion value satisfies the requirements of the first ratio threshold in response to the first cross-overunion value not satisfying the requirements of the first ratio threshold, A video playback apparatus comprising: a first playback module configured to play the aforementioned video to be played and to display a barrage of text within the video to be played using an area other than the display area corresponding to the first mask adjusted during playback.

15. At least one processor, An electronic device including a memory that is communicably connected to at least one processor, The memory stores computer-readable instructions that can be executed by the at least one processor, and when the computer-readable instructions are executed by the at least one processor, the at least one processor A step of obtaining a first mask within a first video frame of a video to be played back, wherein the first mask is used to independently display a first object; A step of determining a first crossover overunion value of the first mask and the second mask in response to the presence of a second mask in a second video frame that is used simultaneously to display the first object independently, the step of the second video frame being the frame preceding the first video frame, In response to the first cross-overunion value not satisfying the requirements of the first ratio threshold, the steps include adjusting the first mask using at least the second mask so that the adjusted first cross-overunion value satisfies the first ratio threshold, An electronic device capable of performing the steps of: playing the video to be played, and displaying a barrage of text within the video to be played using an area other than the display area corresponding to the first mask that has been adjusted during playback.

16. When the computer-readable instruction is executed by the at least one processor, the at least one processor: The electronic device according to claim 15, wherein, in response to the first crossover union value being less than a second ratio threshold, the step of removing at least the first mask can be further performed, and the numerical value of the second ratio threshold is less than the first ratio threshold.

17. When the computer-readable instruction is executed by the at least one processor, the at least one processor: In response to the simultaneous presence of a third mask for independently displaying the first object within a third video frame, the steps include determining a second cross-overunion value based on the second mask and the third mask, The method further includes the step of joining the second mask and the third mask in response to the second cross-overunion value satisfying the requirements of the first ratio threshold, The step of adjusting the first mask using at least the second mask is, The electronic device according to claim 15, further comprising adjusting the first mask using the result of combining the second mask and the third mask.

18. When the computer-readable instruction is executed by the at least one processor, the at least one processor: The steps include: obtaining the video to be played back, The steps include determining a set of shortcut points in the video to be played back based on an inter-frame image change rate algorithm, The steps include segmenting at least one video frame sequence from the video to be played back based on the aforementioned shortcut point, The steps include determining a set of objects included in at least one video frame sequence, The electronic device according to claim 15, further comprising the step of adding a mask corresponding to each object based on tracking each object of the set of objects, wherein the coverage area of ​​the mask is determined based on the image area in which the corresponding object is located.

19. When the computer-readable instruction is executed by the at least one processor, the at least one processor: A step of determining the total area of ​​a mask based on the first mask and the fourth mask in response to the presence of a fourth mask within the first video frame, wherein the fourth mask is used to independently display a second object. The electronic device according to claim 15, further comprising the step of holding the first mask or the fourth mask in response to the total area of ​​the masks exceeding an area threshold.

20. A computer-readable medium storing computer-readable instructions, wherein when the computer-readable instructions are executed by a processor, A step of obtaining a first mask within a first video frame of a video to be played back, wherein the first mask is used to independently display a first object; A step of determining a first crossover overunion value of the first mask and the second mask in response to the presence of a second mask in a second video frame that is used simultaneously to display the first object independently, the step of the second video frame being the frame preceding the first video frame, In response to the first cross-overunion value not satisfying the requirements of the first ratio threshold, the steps include adjusting the first mask using at least the second mask so that the adjusted first cross-overunion value satisfies the first ratio threshold, A computer-readable medium that enables the steps of: playing the video to be played, and displaying bullet patterns within the video to be played using an area other than the display area corresponding to the first mask adjusted during playback.