A cover generation method, apparatus, device and medium
By extracting key objects and background frame sequences from videos to synthesize video cover frames, the problems of insufficient information and high bandwidth consumption in existing technologies are solved, achieving fast loading and efficient display of cover information.
Patent Information
- Application Number
- CN202111362508.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-17
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2041-11-17
AI Technical Summary
In existing technologies, video covers that use single or multiple keyframes to create GIF animations suffer from insufficient information or high bandwidth consumption, making it difficult to effectively attract users and resulting in long loading times.
The key object frame sequence and key background frame sequence are obtained from the video. The target background frame is determined by processing the key background frame sequence. The objects in the key object frame sequence are then combined with the background in the target background frame to generate the target cover frame of the video.
The generated video cover contains key information, loads quickly, avoids high bandwidth consumption, and improves the display effect of cover information in video scenarios.
Smart Images

Figure CN116137671B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of image processing, and particularly relates to a cover generation method and device, equipment and medium. BACKGROUND
[0002] With the rapid development of Internet technology and intelligent terminals, watching videos has become a part of people's lives, and a video cover can attract clicks to play the video.
[0003] In the related art, a certain key frame in a video is selected by an algorithm as a video cover, and only a frame of picture has a small amount of information, and it is difficult to outline the video content, and the key frame selected by the algorithm has a certain randomness, and the attractiveness of the video cover content is usually poor.
[0004] In addition, in order to make up for the amount of information, multiple key frames are selected by an algorithm to be superimposed into a GIF (Graphics Interchange Format) animation, but the interval time between multiple frames is usually large, the picture has a jumping feeling, and the GIF animation is essentially a short video, and the traffic occupancy is greater than that of a picture, and the loading time is long. SUMMARY
[0005] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a cover generation method, device, equipment and medium.
[0006] The present disclosure provides a cover generation method, which comprises:
[0007] obtaining a key object frame sequence and a key background frame sequence from a video;
[0008] processing the key background frame sequence to determine a target background frame;
[0009] synthesizing each object in the key object frame sequence with the background in the target background frame to generate a target cover frame of the video.
[0010] The present disclosure also provides a cover generation device, which comprises:
[0011] an obtaining module configured to obtain a key object frame sequence and a key background frame sequence from a video;
[0012] a processing and determining module configured to process the key background frame sequence to determine a target background frame;
[0013] a synthesizing and generating module configured to synthesize each object in the key object frame sequence with the background in the target background frame to generate a target cover frame of the video.
[0014] The embodiment of the disclosure further provides an electronic device, comprising: a processor; a memory for storing executable instructions of the processor; the processor is used for reading the executable instructions from the memory and executing the instructions to realize the cover generation method provided by the embodiment of the disclosure.
[0015] The embodiment of the disclosure further provides a computer readable storage medium, the storage medium stores a computer program, the computer program is used for executing the cover generation method provided by the embodiment of the disclosure.
[0016] The embodiment of the disclosure further provides a computer program product, when the instructions in the computer program product are executed by the processor, the cover generation method provided by the embodiment of the disclosure is realized.
[0017] The technical scheme provided by the embodiment of the disclosure has the following advantages compared with the prior art: the cover generation scheme provided by the embodiment of the disclosure obtains a key object frame sequence and a key background frame sequence from a video, processes the key background frame sequence to determine a target background frame, and synthesizes each object in the key object frame sequence with the background in the target background frame to generate a target cover frame of the video. By using the above technical scheme, each object in the key object frame sequence in the video can be segmented and fused into a selected appropriate background to generate a static picture as the cover of the video, so that the video cover has a key information amount, can be quickly loaded, avoids large traffic consumption, and improves the cover information display effect in the video scene. BRIEF DESCRIPTION OF DRAWINGS
[0018] The above and other features, advantages, and aspects of the embodiments of the disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the original and elements are not necessarily drawn according to the scale.
[0019] Figure 1 A flowchart of a cover generation method provided by the embodiment of the disclosure is shown in the figure;
[0020] Figure 2 A flowchart of another cover generation method provided by the embodiment of the disclosure is shown in the figure;
[0021] Figure 3 An example diagram of a cover generation method provided by the embodiment of the disclosure is shown in the figure;
[0022] Figure 4a A flowchart of object segmentation provided by the embodiment of the disclosure is shown in the figure;
[0023] Figure 4b A flowchart of key frame extraction provided by the embodiment of the disclosure is shown in the figure;
[0024] Figure 4c A flowchart of a process selected for providing a background for an embodiment of the present disclosure;
[0025] Figure 4d A flowchart of a process of picture synthesis provided for an embodiment of the present disclosure;
[0026] Figure 5a A schematic diagram of a key frame sequence provided for an embodiment of the present disclosure;
[0027] Figure 5b A schematic diagram of a key object frame sequence provided for an embodiment of the present disclosure;
[0028] Figure 5c A schematic diagram of a key background frame sequence provided for an embodiment of the present disclosure;
[0029] Figure 5d A schematic diagram of a video cover provided for an embodiment of the present disclosure;
[0030] Figure 6 A schematic diagram of a structure of a cover generation apparatus provided for an embodiment of the present disclosure;
[0031] Figure 7 A schematic diagram of a structure of an electronic device provided for an embodiment of the present disclosure. DETAILED DESCRIPTION
[0032] Embodiments of the present disclosure will be described in more detail by referring to the drawings. Although certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein, but rather the embodiments are provided so as to more completely and thoroughly understand the present disclosure. It is understood that the drawings and embodiments of the present disclosure are for exemplary purposes only and are not intended to limit the scope of the present disclosure.
[0033] It is understood that each of the steps recited in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the execution of the steps shown. The scope of the present disclosure is not limited in this respect.
[0034] The term "comprising" and variations thereof as used herein are open-ended, that is, "including but not limited to". The term "based on" is "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related terms are defined in the description that follows.
[0035] It should be noted that the terms "first", "second", and the like mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.
[0036] It should be noted that the terms "one", "multiple" mentioned in the present disclosure are illustrative and not restrictive, and those skilled in the art should understand that unless the context clearly indicates otherwise, "one" or "multiple" should be understood as "one or more".
[0037] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0038] In actual application, the video cover can be understood as a static picture or dynamic picture that highly outlines the video content and attracts users to play the video. However, the information amount of a key frame selected as a video static cover by an algorithm is relatively small, it is difficult to outline the video content and has a certain randomness, the attractiveness of the video cover content is usually poor, in addition, in order to make up for the information amount, multiple key frames are selected by an algorithm to be superimposed into a GIF dynamic picture (a picture format based on video and dynamically changeable), but the interval time between the multiple frames is usually large, the picture has a jumping feeling, and the GIF dynamic picture is essentially a short video, the traffic occupation is greater than that of a picture, and the loading time is long.
[0039] In view of the above problems, the present application provides a cover generation method, which obtains a key object frame sequence and a key background frame sequence from a video, processes the key background frame sequence to determine a target background frame, and synthesizes each object in the key object frame sequence with the background in the target background frame to generate a target cover frame of the video.
[0040] Therefore, each object in the key object frame sequence in the video can be segmented and fused into a selected appropriate background to generate a static picture as a cover of the video, so that the video cover has a key information amount, can be quickly loaded, avoids large traffic consumption, and improves the cover information display effect in the video scene.
[0041] Specifically, Figure 1 A flowchart of a cover generation method provided by the embodiments of the present disclosure is shown in the figure. The method can be executed by a cover generation device, which can be implemented by software and / or hardware, and can be integrated in an electronic device. As shown in the figure, the method comprises the following steps. Figure 1 The method comprises the following steps.
[0042] Step 101, obtaining a key object frame sequence and a key background frame sequence from a video.
[0043] The video can be any video including an object, and the object can be a person, an animal, a plant, an article, etc. The embodiments of the present disclosure do not limit the source of the video. For example, the video can be a sports video shot by a sports field expert, wherein the sports video can be understood as a video with a theme of sports competition or outdoor sports, and usually has a main character. The video can also be a video of an animal chasing a ball.
[0044] In some embodiments, a video segment is obtained from the video, object and background separation is performed on each frame of image in the video segment, a candidate object frame sequence and a corresponding candidate background frame sequence are obtained, a key object frame sequence is determined from the candidate object frame sequence, and a key background frame sequence corresponding to the key object frame sequence is determined from the candidate background frame sequence. The video segment is composed of multiple frames of images obtained from the video, and the time is usually several seconds.
[0045] In another embodiment, multiple frames of images, such as a number of bullet screens greater than a preset first threshold or a number of likes greater than a preset first threshold, are obtained from the video, object and background separation is performed on each frame of image in the multiple frames of images, a key object frame sequence and a key background frame sequence are obtained. The preset first threshold is set according to the need, and the present disclosure does not make specific limitations. It should be noted that the above is only an example, and the embodiments of the present disclosure do not limit the specific way of obtaining the key object frame sequence and the key background frame sequence from the video.
[0046] Step 102, processing the key background frame sequence to determine a target background frame.
[0047] The key background frame sequence is a plurality of image frames including only the background. It can be understood that the finally generated video cover is a static picture, that is, only one frame of image frame. In the embodiments of the present disclosure, only one background frame is needed, so the key background frame sequence needs to be processed to determine the target background frame.
[0048] In some embodiments, a candidate background feature point set in the key background frame sequence is obtained, a target background feature point set is obtained by quantizing the candidate background feature point set, a target center point feature vector is determined according to the target background feature point set, and a target background frame is determined from the key background frame sequence according to the target center point feature vector.
[0049] In another embodiment, a key background frame is directly selected from the key background frame sequence as the target background frame. It should be noted that the above is only an example, and the embodiments of the present disclosure do not limit the specific way of processing the key background frame sequence to determine the target background frame.
[0050] Step 103, synthesizing each object in the key object frame sequence with the background in the target background frame to generate a target cover frame of the video.
[0051] The key object frame sequence refers to a plurality of image frames including only objects, that is, each key object frame includes an object. It can be understood that the finally generated video cover is a static picture, that is, only one image frame. Therefore, each object needs to be synthesized with the background in the target background frame to generate a target cover frame of the video.
[0052] In some embodiments, the transparency corresponding to each object in the key object frame sequence is configured, and the target cover frame of the video is generated by synthesizing each object with the background in the target background frame after image processing of each object according to the transparency corresponding to each object.
[0053] In another embodiment, each object is directly synthesized with the background in the target background frame to generate a target cover frame of the video, or each object is synthesized with the background in the target background frame after being configured with different colors, etc. to generate a target cover frame of the video. In the synthesis of each object with the background in the target background frame, the position of each object in the target cover frame can be set as needed, such as the position of the object in the corresponding background, or arranged in a straight line, etc. It should be noted that the above is only an example, and the specific manner of processing the key background frame sequence to determine the target background frame is not limited in the embodiments of the present disclosure.
[0054] The cover generation scheme provided by the embodiments of the present disclosure obtains a key object frame sequence and a key background frame sequence from a video, processes the key background frame sequence to determine a target background frame, synthesizes each object in the key object frame sequence with the background in the target background frame, and generates a target cover frame of the video. By using the above technical scheme, each object in the key object frame sequence in the video can be segmented and fused into a selected appropriate background to generate a static picture as a cover of the video, so that the video cover has a key information amount while being quickly loaded and avoiding large traffic consumption, thereby improving the cover information display effect in the video scene.
[0055] Considering that, for example, sports videos usually have the most exciting video clips, the information of this part of the clip is the most representative. The length of the video clip is usually between 3-5 seconds, and the target video cover can be obtained by condensing the video clip into a static picture. Therefore, the video clip (the most exciting part of the video, which can reflect the essence of the video content) can be screened from the video for processing to obtain the key object frame sequence and the key background frame sequence. The present embodiment provides an implementation manner for obtaining the key object frame sequence and the key background frame sequence from the video, which can be implemented by referring to the following steps a to d.
[0056] Step a, determining a video clip in the video that meets a preset information screening condition.
[0057] In the present embodiment, obtaining a video clip from a video mainly refers to obtaining the most exciting video clip in the video that can reflect the essence of the video content. In the present embodiment, the video clip can be obtained by pre-setting an information screening condition to screen the video. The preset information screening condition can be set according to the application scenario, for example, if the number of video interactions meets a preset first threshold, it is determined that the preset information screening condition is met, and for another example, if the number of video viewers meets a preset second threshold, it is determined that the preset information screening condition is met.
[0058] In some embodiments, in response to the number of video interactions reaching the preset first threshold, it is determined that the corresponding video frame is a video clip that meets the preset information screening condition. In some other embodiments, in response to the number of video viewers reaching the preset second threshold, it is determined that the corresponding video frame is a video clip that meets the preset information screening condition. Thus, the video clip screened based on the number of video interactions or the number of video viewers is processed for subsequent object and background separation, so that the finally generated target cover frame can better meet the user demand. The preset first threshold and the preset second threshold are set as needed, and the present disclosure does not make specific limitations. It should be noted that the above is only an example, and the present embodiment does not limit the specific manner of determining the video clip in the video that meets the preset information screening condition.
[0059] Step b, performing object and background separation processing on each frame in the video clip to obtain a candidate object frame sequence and a corresponding candidate background frame sequence.
[0060] In the present embodiment, each image frame in the video clip includes an object and a background. In some embodiments, after the object recognition algorithm identifies the object in each frame in the video clip, the object and the background are separated, so that the object frame and the object frame are segmented to generate the candidate object frame sequence and the corresponding candidate background frame sequence.
[0061] Step c, determining a key object frame sequence that meets a preset object screening condition from the candidate object frame sequence.
[0062] The preset object screening condition can be selected and set according to the application scenario, for example, the difference degree between each frame object is greater than a preset threshold value, for example, the difference degree between each frame object shape, position, color, size and action is greater than a preset threshold value, it can be determined that the object frames meet the preset object screening condition, for example, the object actions of the object frames are the same, but the distance between the object and the object in the background (such as the ground and the like) changes greatly, it can also be determined that the object frames meet the preset object screening condition. Thus, the object frames with large difference in object action can be selected for comparison, so that in the case that the same number of object frames exist in the video cover, more object information can be provided to the user.
[0063] In a specific embodiment, the first frame in the candidate object frame sequence is taken as a reference frame, and the cosine similarity between the second frame and the reference frame is calculated in sequence. If the cosine similarity is less than or equal to a preset third threshold value, the second frame is discarded, and the first frame is taken as the reference to calculate the cosine similarity between the other subsequent frames in the candidate object frame sequence and the first frame. If the cosine similarity is greater than the preset third threshold value, the first frame is retained and the second frame is taken as the reference frame to calculate the cosine similarity between the other subsequent frames in the candidate object frame sequence and the second frame. After calculating the candidate object frame sequence frame by frame, the retained object frame is the key object frame sequence that meets the preset object screening condition. The cosine similarity refers to measuring the similarity between two vectors by measuring the cosine value of the included angle. The preset third threshold value is set according to the need, and the present disclosure does not make specific limitations. Thus, by calculating the cosine similarity between the object frames and comparing it with the third threshold value, the key object frame sequence is obtained, and the accuracy of obtaining the key object frame sequence is further improved, so that the target cover frame generated based on the key object frame sequence can better meet the user's demand.
[0064] Step d, determining a key background frame sequence corresponding to the key object frame sequence from the candidate background frame sequence.
[0065] In the embodiments of the present disclosure, after the key object frame sequence is determined, a key background frame sequence corresponding to the key object frame sequence can be determined from the candidate background frame sequence, that is, before the separation processing of the object and the background, the background frame of the same frame image frame as the key object frame in the key object frame sequence is taken as the key background frame.
[0066] Thus, the video segments of interest to the user are filtered based on preset information filtering conditions such as the number of video interactions and the number of video viewers, and the object and background are separated. Further, the key object frame sequence with a large difference is filtered based on preset object filtering conditions such as the position of the object and the difference between actions, so that the display of each object in the video cover can meet the user's needs and show the highlights of the video, further improving the cover information display effect in the video scene.
[0067] It can be understood that the finally generated video cover is a static picture, i.e., only one frame of image, that is, only one background frame is needed, so the key background frame sequence needs to be processed to determine the target background frame. The present disclosure provides an implementation manner for processing the key background frame sequence to determine the target background frame, which can be implemented by referring to the following steps 1 to 4.
[0068] Step 1, obtaining a candidate background feature point set in the key background frame sequence.
[0069] In order to obtain the candidate background feature point set in the key background frame sequence more quickly, in a specific embodiment, the original background feature points of each frame in the key background frame sequence are extracted according to a preset algorithm, the feature intensity of each original background feature point is determined according to the feature vector corresponding to the original background feature point of each frame, and the feature points satisfying the preset feature point filtering condition are obtained as the candidate background feature point set according to the feature intensity of each original background feature point.
[0070] The preset algorithm can be selected and set as needed, and an image feature algorithm can be used, such as the SIFT (Scale-invariant feature transform) algorithm, which can be used to determine whether there are the same objects between two pictures, and can analyze the corresponding relationship between the objects. Even if the two images have rotation, blur, and scale changes, even if different cameras are used, and even if the angles of the images are different, the SIFT algorithm can detect the stable feature points of the two pictures and establish the corresponding relationship between them.
[0071] The feature vector is used to describe a set of numbers representing the strength of the feature, such as the feature vector corresponding to the original background feature point is 128-dimensional to represent the feature strength of the original background feature point. Further, the feature strength of each original background feature point is determined based on the feature vector corresponding to the original background feature point of each frame. The higher the feature strength, the higher the discriminability of the feature point. Conversely, the lower the feature strength, the less obvious the discriminability of the feature point. Therefore, it is necessary to further obtain feature points that meet the preset feature point screening condition as the candidate background feature point set. The preset feature point screening condition can be selected and set according to the application scenario, such as the top 50 feature points in the feature strength ranking as the candidate background feature point set, or the feature points with a feature strength greater than a preset feature strength threshold as the candidate background feature point set, and the like. The present disclosure does not limit the specific screening condition.
[0072] Step 2: Quantizing the candidate background feature point set to obtain a target background feature point set.
[0073] In order to further obtain the target background feature point set, in a specific embodiment, the candidate background feature point set is clustered to obtain K clusters, K center point vectors corresponding to the K clusters are determined, and the K center point vectors are taken as the target background feature point set.
[0074] The quantization process refers to a data processing process of approximating continuous value digital features to a limited number of discrete values. K-order quantization means that there are K discrete values.
[0075] The clustering algorithm, such as the K-means algorithm, is used to cluster the candidate background feature point set to obtain K clusters. The K-means algorithm divides the sample set into K clusters according to the distance between the samples, so that the points in the cluster are closely connected together, and the distance between the clusters is as large as possible. It can be used to realize K-order quantization. Thus, after obtaining the K clusters, K center point vectors corresponding to the K clusters are determined, and the K center point vectors are taken as the target background feature point set. The center point can be understood as a feature point in a multi-dimensional vector space that is closest to each feature point in the feature point set. Therefore, the candidate background feature point set with high discriminability is selected for quantization to obtain the target background feature point set, thereby improving the subsequent cover display effect.
[0076] Step 3: Determining a target center point feature vector based on the target background feature point set.
[0077] Step 4: Determining a target background frame from the key background frame sequence based on the target center point feature vector.
[0078] In one specific embodiment, the target background feature point set is clustered to obtain M clusters, a target class with the most feature points is determined from the M clusters, and a target center point feature vector of the target class is calculated. The Euclidean distance between each frame in the key background frame sequence and the target center point feature vector is calculated and compared, and the key background frame with the shortest distance is taken as the target background frame.
[0079] In the embodiments of the present disclosure, after obtaining the target background feature point set, the target background feature point set is further clustered to obtain M clusters, a class with the most feature points is determined as the target class from the M clusters, and a target center point feature vector of the target class is calculated. The center point can be understood as a feature point in a multi-dimensional vector space that is closest to each feature point in the feature point set, and the target center point feature vector of the center point is calculated.
[0080] In the embodiments of the present disclosure, after obtaining the target center point feature vector of the target class, the Euclidean distance between each frame in the key background frame sequence and the target center point feature vector is calculated and compared, and the key background frame with the shortest distance is taken as the target background frame.
[0081] The Euclidean distance refers to the shortest distance between two points in Euclidean space. In the embodiments of the present disclosure, each point is represented by a feature vector. The Euclidean distance between all frames in the key background frame sequence and the target center point feature vector is calculated, and the respective Euclidean distances are compared. The key background frame with the shortest Euclidean distance is taken as the target background frame.
[0082] Thus, the background frame with higher discrimination and more feature points is obtained as the target background frame, so that the background part in the video cover generated subsequently can contain more key information to meet the user's demand and improve the cover information display effect.
[0083] Figure 2 Another flowchart of a cover generation method provided by the embodiments of the present disclosure is provided. The embodiments further optimize the cover generation method described above.
[0084] As shown in Figure 2 The method comprises the following steps:
[0085] Step 201: determining a video segment in a video that meets a preset information screening condition, performing object and background separation processing on each frame in the video segment, and obtaining a candidate object frame sequence and a corresponding candidate background frame sequence.
[0086] Step 202: determining a key object frame sequence that meets a preset object screening condition from the candidate object frame sequence, and determining a key background frame sequence corresponding to the key object frame sequence from the candidate background frame sequence.
[0087] Step 203, obtaining a candidate background feature point set in the key background frame sequence, and quantifying the candidate background feature point set to obtain a target background feature point set.
[0088] Step 204, determining a target center point feature vector according to the target background feature point set, and determining a target background frame from the key background frame sequence according to the target center point feature vector.
[0089] The specific implementation of the above steps S201 to S204 can refer to the foregoing content, and will not be described here.
[0090] Step 205, configuring the transparency of each object in the key object frame sequence, image processing each object according to the transparency of each object to generate each target object, and synthesizing each target object with the background in the target background frame.
[0091] In the embodiment of the present disclosure, the transparency of each object in the key object frame sequence can be configured through the transparency setting control. For example, the key object frame sequence includes five key object frames, and the transparencies of the five objects are set to 20%, 35%, 50%, 75% and 90% respectively. The above is only an example, and the transparency of each object in the key object frame sequence can be configured according to the application scene as needed. After configuring the transparency of each object in the key object frame sequence, image processing each object according to the transparency of each object to generate each target object, and synthesizing each target object with the background in the target background frame. In this way, by setting the transparency, different objects can be distinguished, and the cover information display effect in the video scene is further improved.
[0092] In order to further improve the synthesis effect, in a specific embodiment, synthesizing each target object with the background in the target background frame can include determining an object position sequence of each object in the corresponding background according to the key object frame sequence and the key background frame sequence, and synthesizing each target object with the background in the target background frame according to the object position sequence. In this way, each target object is synthesized with the background in the target background frame based on the position of the object, and the generated video cover is more intuitive, further improving the cover information display effect in the video scene.
[0093] Through the above Figure 2 The cover generation method makes the video cover have key information, can be quickly loaded, avoids large flow consumption, and improves the cover information display effect in the video scene.
[0094] To help those skilled in the art better understand the above embodiments, a detailed description will be provided below in conjunction with specific scenarios, using videos of sports competitions or outdoor activities, typically featuring one main character, as an example.
[0095] Specifically, for sports videos, the key information in a video clip (in this disclosure, the most exciting segment that embodies the essence of the video content) includes: objects and their actions. The video cover is created by segmenting the key object frames from the video clip, selecting a suitable background image, and then merging the continuous objects into the background image to obtain a condensed image of the video clip as the video cover.
[0096] Specifically, such as Figure 3 As shown, the cover generation method proposed in this disclosure can be executed by an object segmenter, a keyframe extractor, a background selector, and an image compositer, respectively. Figure 4a - Figure 4b The process is complete.
[0097] Specifically, such as Figure 3 The object segmenter shown includes object recognition, foreground / background segmentation, and object location information, such as... Figure 4a As shown, the process includes steps 4a.1 acquiring video segments, 4a.2 extracting all keyframes from the video segments, 4a.3 performing object recognition and object segmentation on all keyframes, and 4a.4 acquiring candidate key object frame sequences, candidate key background frame sequences, and candidate object position sequences.
[0098] For example, a video clip V is decompressed to obtain a keyframe sequence, V = (F1, F2, F3, ..., Fn), where n represents the number of frames in the video and Fn represents the nth frame in the video. Each frame of the keyframe sequence is separated from the background using an object segmentation algorithm, forming a candidate key object frame sequence P = (P1, P2, P3, ..., Pn) and a candidate key background frame sequence B = (B1, B2, B3, ..., Bn). At the same time, a candidate object position sequence L = (L1, L2, L3, ..., Ln) is obtained for the object in the background.
[0099] For example, Figure 5a This is a schematic diagram of a keyframe sequence provided in an embodiment of the present disclosure. The keyframe sequence is obtained by decompressing a video segment as shown below. Figure 5a The first to fifth keyframes shown are for Figure 5a The first to fifth keyframes in the image are used to separate the object and background, respectively, to obtain the following: Figure 5b The first to fifth frames show the candidate key object frame order and as follows: Figure 5c The first to fifth candidate key background frame sequences are shown.
[0100] Specifically, such as Figure 3The key frame extractor shown includes key frame identification, object action determination, frame discarding strategy and frame cache management, as shown in Figure 4b As shown, it includes step 4b.1 retaining the first frame in the candidate key object frame sequence as the reference frame, step 4b.2 starting inter-frame comparison, step 4b.3 determining whether the cosine similarity exceeds the preset third threshold, step 4b.4 if yes, discarding the compared frame, step 4b.5 if no, retaining the compared frame and taking it as the reference frame, step 4b.6 determining whether there are candidate key object frames to be processed, if yes, returning to execute step 4b.2, step 4b.7 if no, obtaining the key object frame sequence, step 4b.8 determining the key background frame sequence and object position sequence corresponding to the key object frame sequence from the candidate background frame sequence.
[0101] For example, for the candidate key object frame sequence P=(P1, P2, P3,..., Pn), the cosine similarity of the subsequent object frame and the previous frame is calculated with the first frame as the reference frame. If the similarity is less than 50% (a preset value, which can be configured as needed), it is considered that the object actions of the two frames are similar, and the object frame is discarded; otherwise, the frame is retained, and the retained frame is taken as the reference frame to determine whether the next frame is retained. In this way, the key object frame sequence P'=(P'1, P'2, P'3,..., P'm) with greater object action distinction is obtained, where m is the number of retained frames, m<=n. At the same time, the corresponding key background frame sequence B'=(B'1, B'2, B'3,..., B'm) and object position sequence L'=(L'1, L'2, L'3,..., L'm) are obtained.
[0102] For example, continuing with Figure 5a - Figure 5c As an example, according to the above cosine similarity calculation and determination, the first, third and fifth frames are retained as the key object frame sequence, the key background frame sequence and the object position sequence.
[0103] Specifically, as shown in Figure 3 The background selector shown includes object position input and background frame selection strategy, as shown in Figure 4c As shown, it includes step 4c.1 extracting SIFT features for the key background frame sequence, step 4c.2 retaining the top preset number of feature points in terms of feature intensity for each frame, step 4c.3 performing K-order quantization on the feature points of all frames, step 4c.4 counting the feature vectors of each frame, step 4c.5 performing K-means clustering, step 4c.6 calculating the maximum cluster center point, and step 4c.7 taking the background frame farthest from the center point as the target background frame.
[0104] For example, for the key background frame sequence B' = (B'1, B'2, B'3, ..., B'm), SIFT features are extracted frame by frame using the SIFT algorithm (e.g., feature points are represented by 128-dimensional vectors) to obtain background feature vectors. The top 50 feature points in intensity of each frame (preset value, which can be configured as needed) are retained. Finally, the feature points of all frames are quantized to order K.
[0105] The specific quantization method is as follows: all retained feature points are subjected to K-means clustering to obtain K clusters K = (Q1, Q2, Q3, ..., Qk), and the centroid vector of each category is S = (S1, S2, S3, ..., Sk). Finally, all feature points contained in each category are quantized into the centroid vector of that category. For example, all feature points of category Qk are quantized into vector Sk, thus obtaining the K-order quantized feature points K' = (Q'1, Q'2, Q'3, ..., Q'k).
[0106] Further, after quantization, the number of feature points of each order in the K-order feature points of each frame is counted to obtain the feature vector N = (N1, N2, N3, ..., Nk) for each frame, where Nk represents the number of K-order feature points in this frame, thus obtaining the feature vectors of all frames. Next, these vectors are clustered using the K-means algorithm to obtain M clusters K” = (Q”1, Q”2, Q”3, ..., Q”k). The number of samples in each category is counted to obtain the cluster with the most samples Q”j, and the feature vector Cj = (C1, C2, C3, ..., Ck) of the center point of Q”j is calculated.
[0107] Finally, the Euclidean distance between all frames in the key background frame sequence B' = (B'1, B'2, B'3, ..., B'm) and the center point feature vector Cj = (C1, C2, C3, ..., Ck) is calculated and compared, and the frame B'i with the closest distance is selected as the target background frame.
[0108] For example, continue with Figure 5a - Figure 5c For example, following the above method, Figure 5c The third background frame in the image is used as the target background frame.
[0109] Specifically, such as Figure 3 The image compositer shown includes object information, object location information, optimal background, and compositing strategy, such as... Figure 4d As shown, the process includes step 4d.1, adaptively adjusting the transparency of each object in the key object frame sequence (from 30% to 100%), step 4d.2, obtaining the object position sequence of each object in the corresponding background, step 4d.3, superimposing the objects onto the target background frame in sequence, and step 4d.4, obtaining the video cover.
[0110] For example, in order to better superimpose, a reasonable transparency needs to be set for each frame. The transparency T ranges from the lowest 30% (which can be configured as needed) to the highest 100%. The embodiment of the disclosure can adopt an equal division algorithm, divide equally according to the number m of the key object frame sequence P'=(P'1, P'2, P'3,..., P'm), and obtain T=(T1, T2, T3,..., Tm), wherein T1=0.3 and Tm=1.0. The key object frame sequence P'=(P'1, P'2, P'3,..., P'm) modifies the transparency frame by frame, and then is sequentially superimposed on the target background frame B'i according to the object position sequence L'=(L'1, L'2, L'3,..., L'm), and finally a video cover condensed from a video clip is obtained.
[0111] For example, continue to take Figure 5a - Figure 5c For example, after the first, third and fifth frames of the object frames are processed for transparency, they are superimposed on the third frame background frame as the target background frame according to the corresponding object positions, and the generated video cover is as shown in Figure 5d
[0112] Thus, the key frames with significant object action changes are extracted from the video clip, the objects are identified, the object background is separated, the object part is segmented, the SIFT features are analyzed, the most suitable target background is selected by using the K-means algorithm, the objects with different transparencies are fused into the target background in time sequence, and the video cover is generated. The video cover has a key information amount, can be quickly loaded, avoids large traffic consumption, and improves the cover information display effect in a video scene
[0113] Figure 6 A structure schematic diagram of a cover generation device provided by the embodiment of the disclosure is shown in FIG. 1. The device can be realized by software and / or hardware, and can be integrated in an electronic device. As shown in Figure 6 The device includes:
[0114] The acquisition module 301 is configured to acquire a key object frame sequence and a key background frame sequence from a video.
[0115] The processing and determination module 302 is configured to process the key background frame sequence to determine a target background frame.
[0116] The synthesis and generation module 303 is configured to synthesize each object in the key object frame sequence with the background in the target background frame to generate a target cover frame of the video.
[0117] Optionally, the acquisition module 301 includes a first determination unit, a separation acquisition unit, a second determination unit and a third determination unit.
[0118] The first determining unit is configured to determine a video segment in the video that meets a preset information screening condition.
[0119] The separation obtaining unit is configured to perform separation processing of objects and backgrounds on each frame in the video segment to obtain a candidate object frame sequence and a corresponding candidate background frame sequence.
[0120] The second determining unit is configured to determine a key object frame sequence that meets a preset object screening condition from the candidate object frame sequence.
[0121] The third determining unit is configured to determine a key background frame sequence corresponding to the key object frame sequence from the candidate background frame sequence.
[0122] Optionally, the first determining unit is specifically configured to:
[0123] in response to a video interaction quantity reaching a preset first threshold, determine that the corresponding video frame is a video segment that meets the preset information screening condition; and / or,
[0124] in response to a video viewer quantity reaching a preset second threshold, determine that the corresponding video frame is a video segment that meets the preset information screening condition.
[0125] Optionally, the key object frame sequence that meets the preset object screening condition includes: object frames in which a difference degree between actions of objects in each frame is greater than a preset threshold value.
[0126] Optionally, the second determining unit is specifically configured to:
[0127] take a first frame in the candidate object frame sequence as a reference frame, and sequentially calculate a cosine similarity between a second frame and the reference frame;
[0128] if the cosine similarity is less than or equal to a preset third threshold value, discard the second frame, continue to take the first frame as the reference, and calculate a cosine similarity between other subsequent frames in the candidate object frame sequence and the first frame;
[0129] if the cosine similarity is greater than the preset third threshold value, retain the first frame and take the second frame as a reference frame, and calculate a cosine similarity between other subsequent frames in the candidate object frame sequence and the second frame;
[0130] after the candidate object frame sequence is calculated frame by frame, the retained object frame is a key object frame sequence that meets the preset object screening condition.
[0131] Optionally, the processing and determining module 302 includes an obtaining unit, a quantifying unit, a clustering calculation unit, and a comparing unit.
[0132] an acquisition unit, configured to acquire a candidate background feature point set in the key background frame sequence;
[0133] a quantization unit, configured to perform quantization processing on the candidate background feature point set to acquire a target background feature point set;
[0134] a fourth determination unit, configured to determine a target center point feature vector according to the target background feature point set;
[0135] a fifth determination unit, configured to determine a target background frame from the key background frame sequence according to the target center point feature vector.
[0136] Optionally, the acquisition unit is specifically configured to:
[0137] extract original background feature points of each frame in the key background frame sequence according to a preset algorithm;
[0138] determine a feature intensity of each original background feature point according to a feature vector corresponding to the original background feature point of each frame;
[0139] acquire feature points satisfying a preset feature point screening condition as the candidate background feature point set according to the feature intensity of each original background feature point.
[0140] Optionally, the quantization unit is specifically configured to:
[0141] perform clustering processing on the candidate background feature point set to acquire K clusters;
[0142] determine K center point vectors corresponding to the K clusters respectively, and take the K center point vectors as the target background feature point set.
[0143] Optionally, the fourth determination unit is specifically configured to:
[0144] perform clustering processing on the target background feature point set to acquire M clusters, determine a target category with the largest number of feature points from the M clusters, and calculate a target center point feature vector of the target category;
[0145] the fifth determination unit is specifically configured to:
[0146] calculate an Euclidean distance between each frame in the key background frame sequence and the target center point feature vector, and take a key background frame with the closest Euclidean distance to the target center point feature vector as a target background frame.
[0147] Optionally, the synthesis generation module 303 includes a configuration unit, a generation unit and a synthesis unit;
[0148] The configuration unit is configured to configure the transparency corresponding to each object in the sequence of key object frames.
[0149] The generation unit is configured to generate each target object by performing image processing on each object according to the transparency corresponding to each object.
[0150] The synthesis unit is configured to perform synthesis processing on the target background frame and each target object.
[0151] Optionally, the synthesis unit is specifically configured to:
[0152] determine a sequence of object positions of each object in the corresponding background according to the sequence of key object frames and the sequence of key background frames;
[0153] perform synthesis processing on the target background frame and each target object according to the sequence of object positions.
[0154] The cover generation apparatus provided in the embodiments of the present disclosure can perform the cover generation method provided in any of the embodiments of the present disclosure, and has the function modules and beneficial effects corresponding to the execution method.
[0155] Figure 7 A structural schematic diagram of an electronic device is provided in the embodiments of the present disclosure. The following specifically refers to Figure 7 which shows a structural schematic diagram of an electronic device 400 suitable for implementing the electronic device in the embodiments of the present disclosure. The electronic device 400 in the embodiments of the present disclosure can include but is not limited to a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a vehicle terminal (for example, a vehicle navigation terminal), and the like, and a fixed terminal such as a digital TV, a desktop computer, and the like. Figure 7 The electronic device shown is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present disclosure.
[0156] As shown in Figure 7 , the electronic device 400 can include a processing device (for example, a central processing unit, a graphics processing unit, and the like) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 are also stored. The processing device 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0157] In general, the following devices can be connected to the I / O interface 405: input devices 406 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 408 including, for example, a magnetic tape, a hard disk, and the like; and communication devices 409. The communication devices 409 can allow the electronic device 400 to communicate wirelessly or wired with other devices to exchange data. Although Figure 7 The electronic device 400 is shown with various devices, but it is understood that all of the shown devices are not required to be implemented or present. More or less devices can alternatively be implemented or present.
[0158] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 409, or installed from the storage devices 408, or installed from the ROM 402. When the computer program is executed by the processing devices 401, the above-mentioned functions defined in the cover generation method of embodiments of the present disclosure are performed.
[0159] It should be noted that the computer-readable medium described above can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave in a propagated data signal, in which the computer-readable program code is contained. Such a propagated data signal can take many forms, including but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium that can send, propagate or transfer the program for use by or in connection with the instruction execution system, apparatus or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to, wire, cable, RF (radio frequency), etc., or any suitable combination of the above.
[0160] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0161] The computer-readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device, and is not assembled into the electronic device.
[0162] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: in a playing process of a video, receive an information display triggering operation of a user; obtain at least two target information associated with the video; display first target information of the at least two target information in an information display area of a playing page of the video, wherein a size of the information display area is smaller than a size of the playing page; and receive a first switching triggering operation of the user, and switch the first target information displayed in the information display area to second target information of the at least two target information.
[0163] Computer program code for carrying out operations of the present disclosure can be written in any one or more of a variety of programming languages or combinations of languages including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0164] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may
[0165] The units described in the embodiments of the present disclosure can be implemented by software, or by hardware, or by a combination of software and hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself.
[0166] The functionality described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0167] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0168] According to one or more embodiments of the present disclosure, the present disclosure provides a cover generation method, comprising:
[0169] obtaining a key object frame sequence and a key background frame sequence from the video;
[0170] processing the key background frame sequence to determine a target background frame;
[0171] synthesizing each object in the key object frame sequence with the background in the target background frame to generate a target cover frame of the video.
[0172] According to one or more embodiments of the present disclosure, the cover generation method provided by the present disclosure comprises:
[0173] determining a video segment in the video that meets a preset information screening condition;
[0174] performing object and background separation processing on each frame in the video segment to obtain a candidate object frame sequence and a corresponding candidate background frame sequence;
[0175] determining a key object frame sequence that meets a preset object screening condition from the candidate object frame sequence;
[0176] determine a key background frame sequence corresponding to the key object frame sequence from the candidate background frame sequence.
[0177] According to one or more embodiments of the present disclosure, in the cover generation method provided by the present disclosure, the determination of the video segment satisfying the preset information screening condition in the video comprises:
[0178] In response to the video interaction quantity reaching a preset first threshold, the corresponding video frame is determined as the video segment satisfying the preset information screening condition; and / or,
[0179] In response to the video viewer quantity reaching a preset second threshold, the corresponding video frame is determined as the video segment satisfying the preset information screening condition.
[0180] According to one or more embodiments of the present disclosure, the key object frame sequence satisfying the preset object screening condition comprises an object frame in which the difference between the actions of each frame is greater than a preset threshold value.
[0181] According to one or more embodiments of the present disclosure, in the cover generation method provided by the present disclosure, the determination of the key object frame sequence satisfying the preset object screening condition from the candidate object frame sequence comprises:
[0182] Taking a first frame in the candidate object frame sequence as a reference frame, the cosine similarity between a second frame following the first frame and the reference frame is calculated in sequence;
[0183] If the cosine similarity is less than or equal to a preset third threshold value, the second frame is discarded, and the cosine similarity between other subsequent frames in the candidate object frame sequence except the first and second frames and the first frame is calculated with the first frame as the reference;
[0184] If the cosine similarity is greater than the preset third threshold value, the first frame is retained and the cosine similarity between other subsequent frames in the candidate object frame sequence except the first and second frames and the second frame is calculated with the second frame as the reference;
[0185] After the frame-by-frame calculation of the candidate object frame sequence, the retained object frame is the key object frame sequence satisfying the preset object screening condition.
[0186] According to one or more embodiments of the present disclosure, in the cover generation method provided by the present disclosure, the processing of the key background frame sequence to determine the target background frame comprises:
[0187] Obtain a candidate background feature point set in the key background frame sequence;
[0188] Quantitatively process the candidate background feature point set to obtain a target background feature point set;
[0189] determine a target center point feature vector according to the target background feature point set;
[0190] determine a target background frame from the key background frame sequence according to the target center point feature vector.
[0191] According to one or more embodiments of the present disclosure, in the cover generation method provided by the present disclosure, the acquiring of the candidate background feature point set in the key background frame sequence comprises:
[0192] extracting original background feature points of each frame in the key background frame sequence according to a preset algorithm;
[0193] determining the feature intensity of each original background feature point according to the feature vector corresponding to the original background feature point of each frame;
[0194] acquiring feature points satisfying a preset feature point screening condition as the candidate background feature point set according to the feature intensity of each original background feature point.
[0195] According to one or more embodiments of the present disclosure, in the cover generation method provided by the present disclosure, the quantization processing of the candidate background feature point set to acquire the target background feature point set comprises:
[0196] performing clustering processing on the candidate background feature point set to acquire K clusters;
[0197] determining K center point vectors corresponding to the K clusters respectively, and taking the K center point vectors as the target background feature point set.
[0198] According to one or more embodiments of the present disclosure, the determining of the target center point feature vector according to the target background feature point set comprises:
[0199] performing clustering processing on the target background feature point set to acquire M clusters, determining a target category with the largest number of feature points from the M clusters, and calculating a target center point feature vector of the target category;
[0200] The determining of the target background frame from the key background frame sequence according to the target center point feature vector comprises:
[0201] calculating the Euclidean distance between each frame in the key background frame sequence and the target center point feature vector, and taking the key background frame with the closest Euclidean distance to the target center point feature vector as the target background frame.
[0202] According to one or more embodiments of the present disclosure, in the cover generation method provided by the present disclosure, the synthesizing the objects in the key object frame sequence with the background in the target background frame comprises:
[0203] configuring the transparency corresponding to each object in the key object frame sequence;
[0204] generating each target object by image processing according to the transparency corresponding to each object;
[0205] synthesizing each target object with the background in the target background frame.
[0206] According to one or more embodiments of the present disclosure, in the cover generation method provided by the present disclosure, the synthesizing the objects in the key object frame sequence with the background in the target background frame comprises:
[0207] determining an object position sequence of each object in the corresponding background according to the key object frame sequence and the key background frame sequence;
[0208] synthesizing each target object with the background in the target background frame according to the object position sequence.
[0209] According to one or more embodiments of the present disclosure, the present disclosure provides a cover generation device, comprising:
[0210] an acquisition module configured to acquire a key object frame sequence and a key background frame sequence from a video;
[0211] a processing and determination module configured to process the key background frame sequence to determine a target background frame;
[0212] a synthesis generation module configured to synthesize each object in the key object frame sequence with the background in the target background frame to generate a target cover frame of the video.
[0213] According to one or more embodiments of the present disclosure, in the cover generation device provided by the present disclosure, the acquisition module comprises a first determination unit, a separation acquisition unit, a second determination unit and a third determination unit;
[0214] The first determination unit is configured to determine a video segment in the video that meets a preset information screening condition.
[0215] The separation acquisition unit is configured to perform object and background separation processing on each frame in the video segment to acquire a candidate object frame sequence and a corresponding candidate background frame sequence.
[0216] The second determination unit is configured to determine a key object frame sequence that meets a preset object screening condition from the candidate object frame sequence.
[0217] a third determining unit, configured to determine a key background frame sequence corresponding to the key object frame sequence from the candidate background frame sequence.
[0218] According to one or more embodiments of the present disclosure, in the cover generation apparatus provided by the present disclosure, the first determining unit is specifically configured to:
[0219] in response to the number of video interactions reaching a preset first threshold, determining that the corresponding video frame is a video segment satisfying the preset information screening condition; and / or,
[0220] in response to the number of video viewers reaching a preset second threshold, determining that the corresponding video frame is a video segment satisfying the preset information screening condition.
[0221] According to one or more embodiments of the present disclosure, the key object frame sequence satisfying the preset object screening condition comprises: object frames with a difference degree between object actions of each frame greater than a preset threshold value.
[0222] According to one or more embodiments of the present disclosure, in the cover generation apparatus provided by the present disclosure, the second determining unit is specifically configured to:
[0223] taking a first frame in the candidate object frame sequence as a reference frame, sequentially calculating a cosine similarity between a second frame and the reference frame;
[0224] if the cosine similarity is less than or equal to a preset third threshold, discarding the second frame, and continuing to take the first frame as the reference to calculate the cosine similarity between other subsequent frames in the candidate object frame sequence except the first frame and the second frame and the first frame;
[0225] if the cosine similarity is greater than the preset third threshold, retaining the first frame and taking the second frame as a reference frame to calculate the cosine similarity between other subsequent frames in the candidate object frame sequence except the first frame and the second frame and the second frame;
[0226] After calculating the candidate object frame sequence frame by frame, the retained object frame is the key object frame sequence satisfying the preset object screening condition.
[0227] According to one or more embodiments of the present disclosure, in the cover generation apparatus provided by the present disclosure, the processing determination module 302 comprises: an acquisition unit, a quantization unit, a clustering calculation unit and a comparison unit.
[0228] The acquisition unit is configured to acquire a candidate background feature point set in the key background frame sequence.
[0229] The quantization unit is configured to perform quantization processing on the candidate background feature point set to obtain a target background feature point set.
[0230] a fourth determining unit, configured to determine a target center point feature vector according to the target background feature point set;
[0231] a fifth determining unit, configured to determine a target background frame from the key background frame sequence according to the target center point feature vector.
[0232] According to one or more embodiments of the present disclosure, the cover generation apparatus provided by the present disclosure comprises:
[0233] extracting original background feature points of each frame in the key background frame sequence according to a preset algorithm;
[0234] determining a feature intensity of each original background feature point according to a feature vector corresponding to the original background feature point of each frame;
[0235] obtaining feature points satisfying a preset feature point screening condition as the candidate background feature point set according to the feature intensity of each original background feature point.
[0236] According to one or more embodiments of the present disclosure, the cover generation apparatus provided by the present disclosure comprises:
[0237] performing clustering processing on the candidate background feature point set to obtain K clusters;
[0238] determining K center point vectors corresponding to the K clusters respectively, and taking the K center point vectors as the target background feature point set.
[0239] According to one or more embodiments of the present disclosure, the fourth determining unit is specifically configured to: perform clustering processing on the target background feature point set to obtain M clusters, determine a target category with the largest number of feature points from the M clusters, and calculate a target center point feature vector of the target category; and the fifth determining unit is specifically configured to: calculate the Euclidean distance between each frame in the key background frame sequence and the target center point feature vector, and take a key background frame with the closest Euclidean distance to the target center point feature vector as a target background frame.
[0240] According to one or more embodiments of the present disclosure, the cover generation apparatus provided by the present disclosure comprises:
[0241] a configuration unit, configured to configure a transparency corresponding to each object in the key object frame sequence;
[0242] a generation unit, configured to perform image processing on the objects according to the transparency corresponding to each object to generate each target object.
[0243] a synthesizing unit, configured to synthesize each target object with a background in the target background frame.
[0244] According to one or more embodiments of the present disclosure, the synthesizing unit in the cover generation apparatus is specifically configured to:
[0245] determine an object position sequence of each object in a corresponding background according to the key object frame sequence and the key background frame sequence;
[0246] synthesize each target object with a background in the target background frame according to the object position sequence.
[0247] According to one or more embodiments of the present disclosure, the present disclosure provides an electronic device, comprising:
[0248] a processor;
[0249] a memory for storing executable instructions of the processor;
[0250] the processor is configured to read the executable instructions from the memory and execute the instructions to implement the cover generation method according to any one of the present disclosure.
[0251] According to one or more embodiments of the present disclosure, the present disclosure provides a computer readable storage medium, which stores a computer program for executing the cover generation method according to any one of the present disclosure.
[0252] According to one or more embodiments of the present disclosure, the present disclosure provides a computer program product, when the instructions in the computer program product are executed by a processor, the cover generation method according to any one of the present disclosure is implemented.
[0253] The above description is only the preferred embodiments of the present disclosure and the explanation of the applied technical principles. Those skilled in the art should understand that the disclosure range of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features disclosed in the present disclosure (but not limited to) having similar functions to form technical solutions.
[0254] Moreover, while operations are depicted in a particular order, this should not be understood as requiring such an order nor infringing on the scope of the disclosure. Certain of the operations described in the discussion are combinable into a single operation, and certain operations can be separated into several operations. In some embodiments, the operations described in the discussion can be performed in an order different than presented in the discussion. In some embodiments, the operations described in the discussion can be performed concurrently. Also, while several specific implementation details are discussed in the discussion, these should not be interpreted as limiting the scope of the disclosure. Rather, certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.
[0255] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A method for generating a cover, characterized in that, Includes the following steps: Extract key object frame sequences and key background frame sequences from the video; The key background frame sequence is processed to determine the target background frame; wherein, the processing of the key background frame sequence to determine the target background frame includes: obtaining a set of candidate background feature points in the key background frame sequence; quantizing the set of candidate background feature points to obtain a set of target background feature points; determining a target center point feature vector based on the target background feature point set; and determining the target background frame from the key background frame sequence based on the target center point feature vector. The objects in the key object frame sequence are combined with the background in the target background frame to generate the target cover frame of the video.
2. The method according to claim 1, characterized in that, The process of obtaining key object frame sequences and key background frame sequences from the video includes: Identify video segments in the video that meet the preset information filtering conditions; Each frame in the video segment is processed to separate the object and the background, and a candidate object frame sequence and a corresponding candidate background frame sequence are obtained. From the candidate object frame sequence, determine the key object frame sequence that meets the preset object screening conditions; Determine the key background frame sequence corresponding to the key object frame sequence from the candidate background frame sequence.
3. The method according to claim 2, characterized in that, The process of determining video segments in the video that meet preset information filtering conditions includes: In response to the number of video interactions reaching a preset first threshold, the corresponding video frame is determined to be a video segment that meets preset information filtering conditions; and / or, In response to the number of video viewers reaching a preset second threshold, the corresponding video frame is determined to be a video segment that meets the preset information filtering conditions.
4. The method according to claim 2, characterized in that, The key object frame sequence that meets the preset object filtering conditions includes: Frames whose differences in object actions between frames exceed a preset threshold.
5. The method according to claim 2, characterized in that, The step of determining the key object frame sequence that meets the preset object screening conditions from the candidate object frame sequence includes: Using the first frame in the candidate object frame sequence as the reference frame, the cosine similarity between the subsequent second frame and the reference frame is calculated sequentially. If the cosine similarity is less than or equal to a preset third threshold, the second frame is discarded, and the first frame is used as a reference to calculate the cosine similarity between the first frame and other subsequent frames in the candidate object frame sequence other than the first and second frames. If the cosine similarity is greater than the preset third threshold, then the first frame is retained and the second frame is used as the reference frame to calculate the cosine similarity between the other subsequent frames in the candidate object frame sequence other than the first and second frames and the second frame. After calculating frame by frame in the candidate object frame sequence, the retained object frames are the key object frame sequences that meet the preset object screening conditions.
6. The method according to claim 1, characterized in that, The step of obtaining the candidate background feature point set in the key background frame sequence includes: The original background feature points of each frame in the key background frame sequence are extracted according to a preset algorithm; The feature intensity of each original background feature point is determined based on the feature vector corresponding to the original background feature point of each frame; Based on the feature intensity of each original background feature point, feature points that meet the preset feature point screening conditions are obtained as the candidate background feature point set.
7. The method according to claim 1, characterized in that, The step of quantizing the candidate background feature point set to obtain the target background feature point set includes: The candidate background feature point set is clustered to obtain K clusters; Determine the K center point vectors corresponding to the K clusters respectively, and use the K center point vectors as the target background feature point set.
8. The method according to claim 1, characterized in that, Determining the target center point feature vector based on the target background feature point set includes: Clustering is performed on the target background feature point set to obtain M clusters. The target category with the most feature points is determined from the M clusters, and the target center point feature vector of the target category is calculated. Determining the target background frame from the key background frame sequence based on the target center point feature vector includes: Calculate the Euclidean distance between each frame in the key background frame sequence and the feature vector of the target center point, and take the key background frame with the closest Euclidean distance to the feature vector of the target center point as the target background frame.
9. The method according to claim 1, characterized in that, The step of compositing each object in the key object frame sequence with the background in the target background frame includes: Configure the transparency of each object in the key object frame sequence; Image processing is performed on each object based on the transparency of each object to generate each target object; The target objects are then composited with the background in the target background frame.
10. The method according to claim 9, characterized in that, The step of compositing the target objects with the background in the target background frame includes: Determine the object position sequence of each object in the corresponding background based on the key object frame sequence and the key background frame sequence; The target objects are composited with the background in the target background frame based on the object position sequence.
11. A cover generating device, characterized in that, include: The acquisition module is used to acquire key object frame sequences and key background frame sequences from the video; A processing and determination module is used to process the key background frame sequence to determine a target background frame; wherein, the processing and determination module is specifically used to: obtain a set of candidate background feature points in the key background frame sequence; perform quantization processing on the set of candidate background feature points to obtain a set of target background feature points; determine a target center point feature vector based on the set of target background feature points; and determine the target background frame from the key background frame sequence based on the target center point feature vector. The compositing module is used to compose each object in the key object frame sequence with the background in the target background frame to generate the target cover frame of the video.
12. An electronic device, characterized in that, The electronic device includes: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the cover generation method according to any one of claims 1-10.
13. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, is used to implement the cover generation method according to any one of claims 1-10.
14. A computer program product, characterized in that, When the instructions in the computer program product are executed by a processor, the cover generation method as described in any one of claims 1-10 is implemented.
Citation Information
Patent Citations
Image summarization system and method
US20190251364A1