Video thumbnail generation method, video thumbnail display method and corresponding devices
By detecting motion vector and structural similarity of video keyframes, determining the target keyframe to generate video thumbnails, solving the problems of low generation efficiency and waste of resources in the prior art, and achieving more efficient and high-quality video thumbnail generation.
Patent Information
- Application Number
- CN202510716893.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art is inefficient and wastes resources when generating video thumbnails, especially in scenarios where video subject changes less, affecting the generation quality and system performance.
By performing motion vector detection and structural similarity detection on keyframes in the video, target keyframes are determined, and video thumbnails are generated based on these keyframes to reduce redundant generation.
Improve the efficiency of video thumbnail generation, reduce the use of computing resources, improve server performance, and improve generation quality.
Smart Images

Figure CN120390105A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a video thumbnail generation method, a video thumbnail display method, and corresponding devices. Background Art
[0002] With the development of computer technology, applications can provide users with more and more functions. For example, in a video application, when a video progress bar is dragged, a video thumbnail is displayed above the video progress bar.
[0003] However, in the related art, when generating video thumbnails, there are problems such as low generation efficiency and waste of resources. Summary of the Invention
[0004] This summary is provided to introduce concepts in a brief form that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0005] In a first aspect, the present disclosure provides a video thumbnail generation method, which is applied to a server. The video thumbnail generation method includes: Extract key frames from the video to obtain multiple key frames; Performing motion vector detection on the multiple key frames to obtain a motion change value used to characterize the degree of change in motion of the video subject between each two key frames, and / or performing structural similarity detection on the multiple key frames to obtain the structural similarity between each two key frames; determining a target key frame among the multiple key frames according to the motion change value and / or the structural similarity; A video thumbnail of the video is generated according to the target key frame.
[0006] In a second aspect, the present disclosure provides a video thumbnail display method, which is applied to a client. The video thumbnail display method includes: Obtaining a mosaic image from a server, wherein the mosaic image includes a video thumbnail, the video thumbnail being generated by the server based on a target key frame in the video, the target key frame being determined by the server based on a motion change value and / or structural similarity between every two key frames in the video, the motion change value being used to represent a degree of motion change of the video subject between every two key frames; In response to a control operation on the playback progress of the video, a target video thumbnail is determined from the mosaic graph for display.
[0007] In a third aspect, the present disclosure provides a video thumbnail generation device, which includes: An extraction module, configured to extract key frames from a video to obtain a plurality of key frames; A detection module, configured to perform motion vector detection on the plurality of key frames to obtain a motion change value for characterizing the motion change degree of the video main body between every two of the key frames, and / or perform structural similarity detection on the plurality of key frames to obtain the structural similarity between every two of the key frames; A determination module, configured to determine target key frames from the plurality of key frames according to the motion change value and / or the structural similarity; A generation module, configured to generate a video thumbnail of the video according to the target key frames.
[0008] In a fourth aspect, the present disclosure provides a video thumbnail display device, which includes: An acquisition module, configured to acquire a spliced image from a server, where the spliced image includes a video thumbnail, the video thumbnail is generated by the server according to target key frames in a video, the target key frames are determined by the server according to the motion change value and / or the structural similarity between every two key frames in the video, and the motion change value is used to characterize the motion change degree of the video main body between every two of the key frames; A display module, configured to determine a target video thumbnail from the spliced image for display in response to a control operation on the playback progress of the video.
[0009] In a fifth aspect, the present disclosure provides a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processing device, the steps of the method described in the first aspect are implemented.
[0010] In a sixth aspect, the present disclosure provides an electronic device, including: A storage device, on which a computer program is stored; A processing device, configured to execute the computer program in the storage device to implement the steps of the method described in the first aspect.
[0011] In a seventh aspect, the present disclosure provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.
[0012] Through the above technical solutions, motion vector detection can be performed on multiple key frames extracted from a video to obtain a motion change value for characterizing the motion change degree of the video subject between every two key frames, and / or structural similarity detection can be performed on the multiple key frames to obtain the structural similarity between every two key frames. Moreover, the target key frame can be determined from the multiple key frames according to the motion change value and / or the structural similarity, and a video thumbnail of the video can be generated based on the target key frame. Thus, when generating a video thumbnail, only the key frames in the video need to be extracted and processed, thereby reducing the amount of data processing. On the one hand, the generation efficiency of the video thumbnail can be improved, and on the other hand, the occupation of computing resources can be reduced, enhancing the server performance. In addition, since the motion change value and / or the structural similarity between key frames can reflect the content change degree between key frames, when determining the target key frame according to the motion change value and / or the structural similarity between key frames and generating the video thumbnail of the video based on the target key frame, the generation of redundant thumbnails can be reduced, improving the generation quality of the video thumbnail.
[0013] Other features and advantages of the present disclosure will be described in detail in the following specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In combination with the accompanying drawings and with reference to the following specific implementation, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the original elements and elements are not necessarily drawn to scale. In the drawings: Figure 1 is a flowchart of a method for generating a video thumbnail according to an exemplary embodiment of the present disclosure; Figure 2 is a schematic diagram of generating a mosaic according to an exemplary embodiment of the present disclosure; Figure 3 is a schematic diagram of dividing the blank area in a canvas into child nodes of a parent node according to an exemplary embodiment of the present disclosure; Figure 4 is another schematic diagram of dividing the blank area in a canvas into child nodes of a parent node according to an exemplary embodiment of the present disclosure; Figure 5 is a schematic flowchart of generating a differential package according to an exemplary embodiment of the present disclosure; Figure 6 is a flowchart of a method for displaying a video thumbnail according to an exemplary embodiment of the present disclosure; Figure 7 is a schematic diagram of intercepting and displaying a target video thumbnail according to an exemplary embodiment of the present disclosure; Figure 8 It is a schematic diagram of video thumbnail generation and display shown according to an exemplary embodiment of the present disclosure; Figure 9 It is a block diagram of the structure of a video thumbnail generation device shown according to an exemplary embodiment of the present disclosure; Figure 10 It is a block diagram of the structure of a video thumbnail display device shown according to an exemplary embodiment of the present disclosure; Figure 11 It is a schematic diagram of the structure of an electronic device shown according to an exemplary embodiment of the present disclosure. Specific embodiments
[0015] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0016] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0017] The term "including" and its variants used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0018] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions executed by these devices, modules or units or their interdependent relationships.
[0019] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0020] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0021] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to users in an appropriate manner and the authorization of users should be obtained in accordance with relevant laws and regulations.
[0022] For example, when responding to an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the technical solutions of the present disclosure according to the prompt message.
[0023] As an optional but non-limiting implementation manner, when responding to an active request from a user, the manner of sending a prompt message to the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0024] It is understandable that the above process of notifying and obtaining user authorization is only illustrative and does not limit the implementation manners of the present disclosure, and other manners that meet relevant laws and regulations can also be applied to the implementation manners of the present disclosure.
[0025] Meanwhile, it is understandable that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws and regulations and related regulations.
[0026] As mentioned in the background art, when generating video thumbnails in related technologies, there are problems such as low generation efficiency and resource waste.
[0027] Specifically, when generating video thumbnails in related technologies, video frames are generally extracted from a video regularly or quantitatively, and then video thumbnails are generated based on the extracted video frames. For example, one video frame is extracted from the video every 5 video frames or every 5 seconds, and then video thumbnails are generated based on all the extracted video frames.
[0028] When generating video thumbnails based on this method, since a large number of video frames need to be extracted from the video and corresponding processing is required, on the one hand, the generation efficiency of the thumbnails is reduced, and on the other hand, a large amount of computing resources need to be consumed, affecting the system performance. In addition, when facing a video scene with a small change in the video subject, a large number of redundant thumbnails will also be generated, thus affecting the generation quality of the video thumbnails.
[0029] In view of this, the present disclosure provides a method for generating video thumbnails, a method for displaying video thumbnails, and a corresponding device to solve the above technical problems.
[0030] The following further explains the embodiments of the present disclosure in conjunction with the accompanying drawings.
[0031] Figure 1 is a flowchart of a method for generating video thumbnails shown according to an exemplary embodiment of the present disclosure, which is applied to a server. Referring to Figure 1 , the method may include the following steps: S101: Extract key frames from the video to obtain multiple key frames.
[0032] The video in this embodiment may be an advertising video, a film and television video, a game video, or a travel video. Of course, it may also be other videos, and the embodiments of the present disclosure do not impose any restrictions on this.
[0033] S102: Detect motion vectors for multiple key frames to obtain motion change values for characterizing the motion change degree of the video subject between every two key frames, and / or detect the structural similarity for multiple key frames to obtain the structural similarity between every two key frames.
[0034] Exemplarily, for each extracted key frame, according to the playback order of the key frame in the video, motion vector detection and / or structural similarity detection may be performed on the current key frame and each key frame located after the current key frame, so as to obtain the motion change value and / or structural similarity between every two key frames.
[0035] For example, the key frames may include key frame 1, key frame 2, key frame 3, and key frame 4, and the playback order of the key frames in the video is: key frame 1, key frame 2, key frame 3, key frame 4. Thus, when performing motion vector detection and / or structural similarity detection on multiple key frames, the motion change value and / or structural similarity between key frame 1 and key frame 2, the motion change value and / or structural similarity between key frame 1 and key frame 3, and the motion change value and / or structural similarity between key frame 1 and key frame 2 may be detected first; then, the motion change value and / or structural similarity between key frame 2 and key frame 3, and the motion change value and / or structural similarity between key frame 2 and key frame 4 may be detected; finally, the motion change value and / or structural similarity between key frame 3 and key frame 4 may be detected, so as to obtain the motion change value and / or structural similarity between every two key frames.
[0036] S103: Determine target key frames for generating video thumbnails from multiple key frames according to the motion change value and / or structural similarity.
[0037] It should be understood that the greater the motion change value between key frames, the greater the degree of motion change of the video subject between key frames, and the greater the content difference between key frames; the smaller the motion change value, the smaller the degree of motion change of the video subject between key frames, and the smaller the content difference between key frames. Therefore, in order to reduce the repeated use of those key frames with small content differences for thumbnail generation, a motion threshold can be preset. Thus, when determining the target key frames for generating video thumbnails from multiple key frames, the key frames with a motion change value greater than or equal to the preset motion threshold among the multiple key frames can be determined as the target key frames, thereby reducing the generation of redundant thumbnails and improving the quality of video thumbnail generation.
[0038] In addition, it should be understood that the smaller the structural similarity between key frames, the less similar the image structures between key frames, and the greater the content difference between key frames; the greater the structural similarity between key frames, the more similar the image structures between key frames, and the smaller the content difference between key frames. Therefore, in order to reduce the repeated use of those key frames with small content differences for thumbnail generation, a similarity threshold can be preset. Thus, when determining the target key frames for generating video thumbnails from multiple key frames, the key frames with a structural similarity less than or equal to the preset similarity threshold among the multiple key frames can be determined as the target key frames, thereby reducing the generation of redundant thumbnails and improving the quality of video thumbnail generation.
[0039] That is to say, in a possible way, determining the target key frames from multiple key frames according to the motion change value and / or the structural similarity may include: Among multiple key frames, determining the key frames with a motion change value greater than or equal to the preset motion threshold as the target key frames; and / or, among multiple key frames, determining the key frames with a structural similarity less than or equal to the preset similarity threshold as the target key frames.
[0040] In this embodiment, the preset motion threshold and / or the preset similarity threshold can be a fixed value or a dynamically changing value, and the embodiments of the present disclosure do not impose any restrictions on this. When the preset motion threshold and / or the preset similarity threshold is a dynamically changing value, the preset motion threshold and / or the preset similarity threshold can change dynamically according to the degree of content change of the video. For example, when the degree of content change of the video is small, a smaller preset motion threshold or a larger similarity threshold can be selected; when the degree of content change of the video is large, a larger preset motion threshold or a lower similarity threshold can be selected.
[0041] It should be understood that for different types of videos, the degree of change in video content is different. Therefore, in order to generate better video thumbnails for different types of videos, among possible methods, different types of videos correspond to different preset motion thresholds, and / or different types of videos correspond to different preset similarity thresholds.
[0042] Exemplarily, for lecture videos, since the degree of change in video content is not large, in order to extract different key frames from lecture videos for generating video thumbnails, a larger similarity threshold can be selected, and / or a smaller motion threshold can be selected. For example, the similarity threshold can be set to 0.95 and the motion threshold can be set to 6%.
[0043] Exemplarily, for sports videos, since the degree of change in video content is large, in order to extract different key frames from sports videos for generating video thumbnails, a smaller similarity threshold can be selected, and / or a larger motion threshold can be selected. For example, the similarity threshold can be set to 0.85 and the motion threshold can be set to 15%.
[0044] It should be understood that after the server generates a video thumbnail based on the target key frames, it is used for display on the client. Therefore, in order to reduce the situation where the number of video thumbnails is small due to the small number of target key frames, which affects the user's browsing experience. Among possible methods, when the motion change value corresponding to a key frame is less than the preset motion threshold, the previous key frame can be determined as the target key frame; and / or when the structural similarity corresponding to a key frame is greater than or equal to the preset similarity threshold, the previous key frame can be determined as the target key frame, so as to increase the number of target key frames, and then enrich the display content of the video thumbnail and improve the user experience.
[0045] S104: Generate a video thumbnail of the video according to the target key frames.
[0046] Exemplarily, the video thumbnail size can be preset in advance, and then based on the video thumbnail size, each target key frame is scaled to obtain the video thumbnail of the video.
[0047] It should be understood that since the video content included in different key frames is different, the importance levels between different key frames are also different. Therefore, in order to further improve the generation efficiency of video thumbnails and reduce the waste of server resources, among possible methods, generating a video thumbnail of the video according to the target key frames may include: Perform content analysis on the target key frames to obtain the content analysis results, and determine the first type of key frames and the second type of key frames in the target key frames according to the content analysis results, where the importance level of the first type of key frames is higher than that of the second type of key frames; generate the first video thumbnail of the video according to the first type of key frames; generate the second video thumbnail of the video according to the second type of key frames, where the resolution of the first video thumbnail is greater than that of the second video thumbnail.
[0048] In this embodiment, the importance level of the key frames can be determined according to the area size of the main content included in the key frames, and / or the number of video content features, and of course, it can also be determined according to other methods, and the embodiments of the present disclosure do not impose any restrictions on this.
[0049] Exemplarily, continuing to refer to the above example, if it is determined that key frame 1, key frame 2, and key frame 4 are the first type of key frames and key frame 3 is the second type of key frames according to the number of video content features included in the key frames, then when generating the video thumbnail, lossless compression can be used for key frame 1, key frame 2, and key frame 4 to obtain a video thumbnail with a higher resolution; lossy compression can be used for key frame 3 to obtain a video thumbnail with a lower resolution.
[0050] Through the above technical solutions, motion vector detection can be performed on multiple key frames extracted from the video to obtain a motion change value for characterizing the motion change degree of the video main body between every two key frames, and / or structural similarity detection can be performed on multiple key frames to obtain the structural similarity between every two key frames, and the target key frames can be determined among the multiple key frames according to the motion change value and / or the structural similarity, and the video thumbnail of the video can be generated according to the target key frames. Thus, when generating the video thumbnail, only the key frames in the video need to be extracted and processed, thereby reducing the data processing volume. On the one hand, the generation efficiency of the video thumbnail can be improved, and on the other hand, the occupation of computing resources can be reduced, and the server performance can be improved. In addition, since the motion change value and / or the structural similarity between the key frames can reflect the content change degree between the key frames, thus, when determining the target key frames according to the motion change value and / or the structural similarity between the key frames and generating the video thumbnail of the video based on the target key frames, the generation of redundant thumbnails can be reduced, and the generation quality of the video thumbnail can be improved.
[0051] In a possible way, there can be multiple video thumbnails, and correspondingly, the video thumbnail generation method can further include: Insert any one of the multiple video thumbnails into the initially blank canvas, and use the area of the canvas where the thumbnail to be processed is inserted as the initial parent node, and loop to execute the following process: Take any one of the video thumbnails that are not inserted into the canvas among multiple video thumbnails as the thumbnail to be processed. Divide the blank area in the canvas into child nodes of the parent node, insert the video thumbnail to be processed into the child nodes, and use the child nodes as the new parent node until all the multiple video thumbnails are inserted into the canvas, obtaining a spliced image corresponding to the multiple video thumbnails.
[0052] In this embodiment, dividing the blank area in the canvas into child nodes of the parent node can be to divide the blank area into child nodes of the parent node in the vertical direction or in the horizontal direction. The embodiments of the present disclosure do not impose any restrictions on this.
[0053] Exemplarily, as Figure 2 shown, multiple key frames can include Key Frame 1, Key Frame 2, Key Frame 3, and Key Frame 4. After inserting Key Frame 1 into the upper left corner of the canvas, Key Frame 2 can be taken as the thumbnail to be processed, the canvas area where Key Frame 1 is located can be taken as the parent node, and the blank area can be divided into child nodes of the parent node in the vertical direction, namely Child Node 1 and Child Node 2; then, Key Frame 2 can be randomly inserted into any one of Child Node 1 and Child Node 2. For example, after inserting Key Frame 2 into Child Node 1, Key Frame 3 can be taken as the thumbnail to be processed, the canvas area where Key Frame 2 is located can be taken as the new parent node, and the blank area can continue to be divided into child nodes of the parent node in the vertical direction, namely Child Node 3 and Child Node 4; then, Key Frame 3 can be randomly inserted into any one of Child Node 3 and Child Node 4. For example, after inserting Key Frame 3 into Child Node 4, Key Frame 4 can be taken as the thumbnail to be processed, the canvas area where Key Frame 3 is located can be taken as the new parent node, and the blank area can continue to be divided into child nodes of the parent node in the vertical direction, namely Child Node 5 and Child Node 6; then, Key Frame 4 can be randomly inserted into any one of Child Node 5 and Child Node 6. For example, Key Frame 4 can be inserted into Child Node 5, thereby obtaining a spliced image corresponding to the multiple video thumbnails.
[0054] As mentioned above, according to the importance degree of the key frames, different compression methods can be adopted for the key frames to be compressed, thereby obtaining thumbnails of different sizes. If the thumbnails are spliced into the canvas based on the conventional splicing method, since it is difficult to neatly arrange the thumbnails of different sizes on the canvas, there is a problem of low utilization rate of the canvas space. In this embodiment, however, the canvas is segmented based on the binary tree space segmentation algorithm, whereby the canvas can be recursively divided into smaller parts according to the thumbnails of different sizes, so that it can better adapt to the thumbnails of different sizes, and further improve the utilization rate of the canvas space.
[0055] In a possible way, dividing the blank area in the canvas into child nodes of the parent node may include: According to the size information of the blank area in the canvas and the size information of the thumbnail to be processed, divide the blank area in the canvas into child nodes of the parent node so that the canvas area corresponding to the child nodes can accommodate the thumbnail to be processed.
[0056] Exemplarily, as Figure 3 shown, after inserting key frame 5 into the upper left corner of the canvas, key frame 5 can be used as the thumbnail to be processed, and the blank area can be randomly divided into child node 1 and child node 2 in the vertical direction or the horizontal direction. For example, the blank area can be divided into child node 1 and child node 2 in the vertical direction; then, the lengths and widths of key frame 6, child node 1, and child node 2 can be determined. If the length of key frame 6 is greater than the lengths of child node 1 and child node 2, the blank area can be divided into child node 1 and child node 2 in the horizontal direction, as Figure 3 shown in (a), to avoid the situation where neither child node 1 nor child node 2 can accommodate key frame 6 when key frame 6 is inserted into child node 1 or child node 2 after dividing the blank area into child node 1 and child node 2 in the vertical direction, as Figure 3 shown in (b).
[0057] Exemplarily, as Figure 4 shown, after inserting key frame 7 into the upper left corner of the canvas, key frame 7 can be used as the thumbnail to be processed, and the blank area can be randomly divided into child node 1 and child node 2 in the vertical direction or the horizontal direction. For example, the blank area can be divided into child node 1 and child node 2 in the horizontal direction; then, the lengths and widths of key frame 8, child node 1, and child node 2 can be determined. If the width of key frame 8 is greater than the widths of child node 1 and child node 2, the blank area can be divided into child node 1 and child node 2 in the vertical direction, as Figure 4 shown in (a), to avoid the situation where neither child node 1 nor child node 2 can accommodate key frame 8 when key frame 8 is inserted into child node 1 or child node 2 after dividing the blank area into child node 1 and child node 2 in the horizontal direction, as Figure 4 shown in (b).
[0058] In possible ways, the video thumbnail generation method may further include: When the first video thumbnail in the splicing map meets the preset conditions, trace back to the parent node upward from the node corresponding to the first video thumbnail to obtain the target parent node corresponding to the first video thumbnail; when there is an idle node among the child nodes of the target parent node, merge the node corresponding to the first video thumbnail with the idle node, where the idle node is a node that has not been inserted with a video thumbnail.
[0059] In this embodiment, the preset condition may be that the corresponding video segment in the video is deleted, modified, the motion change value and / or the structural similarity exceeds the threshold, or the cache in the opinion client expires, etc. Of course, it may also be others, and the embodiments of the present disclosure do not impose any restrictions on this.
[0060] In the above manner, when the first video thumbnail in the spliced image meets the preset condition, the parent node can be traced back upward from the node corresponding to the first video thumbnail to obtain the target parent node corresponding to the first video thumbnail. And when there is an idle node among the child nodes of the target parent node, the node corresponding to the first video thumbnail is merged with the idle node. Thus, the unused blank area can be reduced by dynamically adjusting the nodes in the canvas, further improving the space utilization rate of the canvas.
[0061] In a possible way, the video thumbnail generation method further includes: Performing change detection on the spliced image; when it is detected that there is a change in the second video thumbnail in the spliced image, determining the differential block information that has changed in the second video thumbnail, and generating a differential package according to the differential block information, where the differential package is used for the client to update the locally stored spliced image.
[0062] It should be understood that after the spliced image is generated, due to reasons such as size adjustment, the video thumbnails in the spliced image may be adjusted, thereby obtaining different versions of the spliced image. And after the server generates the spliced image, it sends the spliced image to the client so that when the client controls the playback progress of the video, the target video thumbnail can be determined from the spliced image for display. Thus, in order to reduce the data transmission volume, after obtaining the new version of the spliced image, change detection can be performed on the spliced image between the new version and the initial version, and when it is detected that there is a change in the second video thumbnail in the spliced image, the differential block information that has changed in the second video thumbnail can be determined, and a differential package can be generated according to the differential block information. Thus, the server can only transmit the differential package to the client so that the client can obtain the latest spliced image according to the differential package and the locally stored initial version of the spliced image. Compared with the server transmitting the new version of the spliced image to the client, the data transmission volume between the server and the client can be reduced, and the data transmission efficiency can be improved.
[0063] Exemplarily, such as Figure 5As shown, if there are two versions of the spliced images, denoted as spliced image V1.0 and spliced image V1.1, and spliced image V1.0 is the spliced image of the initial version, then for each video thumbnail in each version of the spliced image, the video thumbnail can be divided into 16×16 pixel blocks, and the hash value of each pixel block can be calculated; then for each pixel block, compare its hash values in spliced image V1.0 and spliced image V1.1; if the difference in hash values exceeds a preset threshold, for example, exceeds 5%, then mark this block as a "dirty block"; after obtaining all the "dirty blocks", the server generates a difference package according to the image information corresponding to the "dirty blocks" and the position coordinates in the spliced image and other information, and can transmit the difference package to the server so that the client can obtain the latest spliced image based on the difference package and the locally stored spliced image V1.0; if the difference in hash values does not exceed the preset threshold, no processing is performed.
[0064] Based on the same concept, an embodiment of the present disclosure further provides a method for displaying video thumbnails, which is applied to a client, as Figure 6 shown, the method for displaying video thumbnails may include: S601: Obtain a spliced image from the server, where the spliced image includes video thumbnails, and the video thumbnails are generated by the server according to the target key frames in the video, and the target key frames are determined by the server according to the motion change value and / or structural similarity between every two key frames in the video, and the motion change value is used to characterize the degree of motion change of the video subject between every two key frames; S602: In response to a control operation on the playback progress of the video, determine a target video thumbnail from the spliced image for display.
[0065] The client in this embodiment may be a video playback platform in a terminal device, or a video application program, and of course, it may also be others, and the embodiments of the present disclosure do not make any restrictions on this.
[0066] Exemplarily, when the client is a video application program, during the process of playing a video through the video application program, the progress bar of the video can be dragged, and according to the position of the progress bar dragged to, determine the target video thumbnail corresponding to the position of the progress bar from the spliced image, and can be displayed above the progress bar.
[0067] Through the above technical solution, a spliced image including video thumbnails can be obtained from the server, and during the process of controlling the video playback progress, a target video thumbnail can be determined from the spliced image for display. Since the video thumbnails in the spliced image are generated by the server according to the target key frames in the video, and the target key frames are determined by the server according to the motion change value and / or structural similarity between every two key frames in the video, when the server generates video thumbnails, it can extract and process only the key frames in the video, thereby reducing the number of generated video thumbnails. Furthermore, when the client obtains the spliced image including video thumbnails from the server, the communication resources between the client and the server can be reduced, and the processing performance of the client can be improved. In addition, since the motion change value and / or structural similarity between key frames can reflect the content change degree between key frames, when the server determines the target key frames according to the motion change value and / or structural similarity between key frames and generates video thumbnails of the video based on the target key frames, the generation of redundant thumbnails can be reduced, and the generation quality of video thumbnails can be improved. Furthermore, when the client obtains the spliced image including video thumbnails from the server, the communication resources between the client and the server can be further reduced, the processing performance of the client can be further improved, and the display quality of video thumbnails can be improved.
[0068] In a possible manner, the video thumbnails in the spliced image have a start display time and an end display time, and the video thumbnails in the spliced image are stored in sequence according to the start display time. Correspondingly, in response to a control operation on the playback progress of the target video, determining a target video thumbnail from the spliced image for display may include: In response to an operation of dragging the playback progress of the video to the first moment, determine a third video thumbnail in the spliced image, where the time stamp of the target key frame corresponding to the third video thumbnail is the closest to the first moment; according to the first moment, the start display time and the end display time of the third video thumbnail, determine the time ratio of the first moment relative to the start display time of the third video thumbnail; according to the time ratio and the relevant information of the third video thumbnail, intercept the target video thumbnail for display from the third video thumbnail and the fourth video thumbnail, where the fourth video thumbnail is the next video thumbnail stored after the third video thumbnail, and the relevant information includes the size information and coordinate position information of the third video thumbnail.
[0069] It should be understood that when dragging the playback progress of a video in the related art, the video thumbnail closest to the current playback time is found based on binary search or other methods, and then the video thumbnail is displayed. There is a situation where the video frame thumbnail picture jumps, thus affecting the user experience. In this embodiment, through the above method, linear interpolation can be performed in the mosaic picture based on the time ratio, so as to achieve smooth transition between video thumbnails during the dragging process, thereby improving the user experience.
[0070] In a possible way, the video thumbnails in the mosaic picture are stored sequentially in the horizontal direction according to the start display time. Correspondingly, according to the time ratio and the display information of the third video thumbnail in the mosaic picture, intercepting the target video thumbnail for display in the third video thumbnail and the fourth video thumbnail may include: Multiply the time ratio by the width of the third video thumbnail to obtain the coordinate offset; add the abscissa of the third video thumbnail to the coordinate offset to obtain the target abscissa; in the third video thumbnail and the fourth video thumbnail, intercept the target video thumbnail with the same width as the third video thumbnail starting from the target abscissa for display.
[0071] Exemplarily, as Figure 7 shown, if the key frame at the 5th second corresponds to video thumbnail A, and the display time of video thumbnail A is from the 5th second to the 7.5th second, and the key frame at the 10th second corresponds to video thumbnail B, and the display time of video thumbnail B is from the 7.5th second to the 10th second.
[0072] When displaying video thumbnails according to the related technical solution, if the playback progress of the video is dragged to the 6th second, because binary search is used in the related art, the key frame at the 5th second will be used as the closest key frame, so video thumbnail A will be displayed. When the playback progress of the video is dragged to the 7th second, because it is binary search, the key frame at the 5th second will still be used as the closest key frame, and video thumbnail A will also be displayed, thus smooth transition cannot be achieved.
[0073] In this embodiment, if the playback progress of the video is dragged to the 6th second, since the closest key frame is the key frame at the 5th second, thus, the key frame at the 5th second can be used as the third video thumbnail, and according to the first moment, the start display time and the end display time of the key frame at the 5th second, the time ratio of the first moment relative to the start display time of the third video thumbnail can be determined, that is: the first moment, the start display time of the key frame at the 5th second, and the end display time of the key frame at the 5th second can be substituted into the following formula to obtain the time ratio:
[0074] Wherein, represents the time ratio, represents the first moment, represents the start display moment of the third video thumbnail, represents the end display moment of the third video thumbnail.
[0075] After obtaining the time ratio, the time ratio can be substituted into the following formula to obtain the target abscissa:
[0076] wherein, represents the target abscissa, represents the abscissa of video thumbnail A, represents the width of video thumbnail A.
[0077] After obtaining the target abscissa it is possible to start from the target abscissa and intercept a new video thumbnail from the spliced image, which is equivalent to displaying the video thumbnail corresponding to the video frame after the key frame at the 5th second is shifted backward by 1 second. And so on, when dragging to the 7th second, it is equivalent to displaying the video thumbnail obtained by shifting the video thumbnail corresponding to the key frame at the 5th second backward by 2 seconds. Thus, smooth transition can be achieved during the dragging process.
[0078] Wherein, the origin of coordinates in the spliced image can be the upper left vertex of the spliced image, the positive direction of the x-axis can be the horizontal direction from the upper left vertex to the upper right vertex, and the positive direction of the y-axis can be the vertical direction from the upper left vertex to the lower right vertex.
[0079] In a possible way, in response to a control operation on the playback progress of the video, determining a target video thumbnail from the spliced image for display may include: In response to an operation of dragging the playback progress of the video to the second moment, loading and displaying the video thumbnail corresponding to the second moment from the spliced image; Correspondingly, the video thumbnail display method may further include: When loading the video thumbnail corresponding to the second moment, preloading a preset number of video thumbnails to be displayed after the second moment from the spliced image.
[0080] In this embodiment, the preset number can be determined according to the actual situation, and the present disclosure embodiment does not make any limitation thereto. Exemplarily, the preset number can be 3.
[0081] It should be understood that different video thumbnails have different sizes. If the corresponding video is loaded from the stitched image to make a thumbnail only during the dragging process, there will be a problem of lag in loading the video thumbnail. If all the video thumbnails are loaded when the video starts to be displayed, there will be a situation where the video viewer switches to other videos after watching a part of the video content, resulting in waste of traffic resources. In this embodiment, when loading the video thumbnail corresponding to the second moment, a preset number of video thumbnails to be displayed after the second moment can be pre-loaded from the stitched image, thereby reducing the frequency of lag in loading the video thumbnail and reducing the problem of waste of traffic resources.
[0082] In a possible way, the video thumbnail display method may further include: In response to a dragging operation that drags the playback progress of the video to the third moment, determine the dragging speed corresponding to the dragging operation; in the case where the dragging speed is greater than or equal to a preset dragging speed threshold, sample the video to obtain a video frame at the third moment, and render and display the video thumbnail at the third moment according to the video frame at the third moment.
[0083] In this embodiment, the preset dragging speed threshold can be determined according to the actual situation, and the embodiments of the present disclosure do not impose any restrictions on this. Exemplarily, the dragging speed threshold can be 500 pixels per second.
[0084] It should be understood that it takes a certain amount of time to load the video thumbnail from the stitched image. When the dragging speed is too fast, there will be a problem that the video thumbnail is not loaded in time, resulting in no video thumbnail being displayed during the dragging process, thus affecting the user experience. Through this method, when the dragging speed is too fast, the video can be sampled to obtain a video frame at the third moment, and the video thumbnail at the third moment can be rendered and displayed according to the video frame at the third moment, so that the video thumbnail can be displayed in time, thereby improving the user experience.
[0085] In a possible way, the video thumbnail display method may further include: Verify the stitched image to obtain a verification result indicating whether the stitched image is available; Correspondingly, in response to a control operation on the playback progress of the video, determining a target video thumbnail to be displayed from the stitched image may include: In the case where the verification result indicates that the stitched image is available, in response to a control operation on the playback progress of the video, determine a target video thumbnail to be displayed from the stitched image.
[0086] In this embodiment, verifying the stitched image may be to perform a CRC32 check on the stitched image, or to perform a hash check on the stitched image. Of course, the stitched image can also be verified by other means, and the embodiments of the present disclosure do not impose any restrictions on this.
[0087] It should be understood that since the spliced image is obtained from the server, there may be problems such as damage to the spliced image, loss of some video thumbnails in the spliced image, or update of the spliced image by the server during the transmission of the spliced image from the server to the client. Therefore, in this embodiment, after receiving the spliced image, the spliced image is verified to determine whether the received spliced image is accurate, and when the received spliced image is accurate, the target video thumbnail is determined from the spliced image for display according to the control operation of the video playback progress, so that the video thumbnail of normal snow can be displayed on the client.
[0088] If the verification result indicates that the spliced image is unavailable, it means that there are problems such as damage to the spliced image, loss of some video thumbnails in the spliced image, or update of the spliced image by the server during the transmission of the spliced image from the server to the client. Therefore, in order to be able to display the video thumbnail of normal snow on the client, the client can regenerate the acquisition request for the spliced image obtained from the server.
[0089] In a possible way, in order to reduce the data transmission volume and the resource occupation of the client, when the verification result indicates that the spliced image is unavailable, the unavailable block in the spliced image can be determined, and then the acquisition request for the unavailable block in the spliced image is sent to the server, and when the unavailable block returned by the server is received, the spliced image is updated, so as to reduce the repeated generation of the spliced image and the consumption of resources.
[0090] To facilitate understanding of the video thumbnail generation method and video thumbnail display method provided by the embodiments of the present disclosure, the following is described in conjunction with the server and the client: Exemplarily, as Figure 8 shown, after receiving the video, the server can first extract multiple key frames from the video, and can analyze the multiple key frames to obtain the motion change value and structural similarity between every two key frames; then, according to the motion change value and structural similarity, the degree of video content change between every two key frames can be determined. If the motion change value is greater than or equal to the preset motion threshold, or the structural similarity is less than or equal to the preset similarity threshold, it indicates that the degree of video content change between every two key frames is relatively significant, and it can be used as the target key frame and a new video thumbnail is generated; if the motion change value is less than the preset motion threshold, or the key frame with a structural similarity greater than the preset similarity threshold, it indicates that the degree of video content change between every two key frames is relatively small, and it can be not used as the target key frame, and the video thumbnail generated by the previous key frame can be reused.
[0091] After generating video thumbnails for a video, all the generated video thumbnails can be sequentially stitched onto a blank canvas in the playback order of the corresponding key frames in the video, thereby obtaining a stitched image corresponding to the video thumbnails, and recording the position information of each video thumbnail in the stitched image to obtain position metadata.
[0092] After that, when the client needs to display video thumbnails, or when playing a target video on the client, a request for obtaining video thumbnails can be generated and transmitted to the server, so that the server can transmit the stitched image containing the video thumbnails to the client. After receiving the stitched image, if the client detects a control operation on the video playback progress, it can load the corresponding video thumbnail from the stitched image and display it according to the control operation.
[0093] In this embodiment, since the video thumbnails in the stitched image are generated by the server according to the target key frames in the video, and the target key frames are determined by the server according to the motion change value and / or structural similarity between every two key frames in the video, it enables the server to only extract and process the key frames in the video when generating video thumbnails, thereby reducing the number of generated video thumbnails. Furthermore, when the client obtains the stitched image including video thumbnails from the server, it can reduce the communication resources with the server and improve the processing performance of the client. In addition, since the motion change value and / or structural similarity between key frames can reflect the content change degree between key frames, when the server determines the target key frames according to the motion change value and / or structural similarity between key frames and generates video thumbnails of the video based on the target key frames, it can reduce the generation of redundant thumbnails and improve the generation quality of video thumbnails. Furthermore, when the client obtains the stitched image including video thumbnails from the server, it can further reduce the communication resources with the server, further improve the processing performance of the client, and improve the display quality of video thumbnails.
[0094] In addition, to verify the feasibility of this solution, in this embodiment, video thumbnails are generated for the same video in different ways, and the comparison results shown in Table 1 are obtained: Table 1: Comparison results of generating video thumbnails for the same video in different ways
[0095] According to Table 1, when generating video thumbnails based on the video thumbnail generation method in this solution, the volume of the stitched image can be reduced, thereby reducing memory occupancy and transmission overhead and improving the system response speed; the generation time of the stitched image can be reduced, thereby improving the generation efficiency and reducing the user waiting time; the drag response delay can be reduced, thereby improving the response speed of the drag progress bar and enhancing the user interaction experience.
[0096] Based on the same concept, embodiments of the present disclosure further provide a video thumbnail generation device. As Figure 9 shown, the video thumbnail generation device 900 may include: An extraction module 901, configured to extract key frames from a video to obtain a plurality of key frames; A detection module 902, configured to perform motion vector detection on the plurality of key frames to obtain a motion change value for characterizing the motion change degree of the video main body between every two key frames, and / or perform structural similarity detection on the plurality of key frames to obtain the structural similarity between every two key frames; A determination module 903, configured to determine target key frames from the plurality of key frames according to the motion change value and / or the structural similarity; A generation module 904, configured to generate a video thumbnail of the video according to the target key frames.
[0097] Through the above video thumbnail generation device 900, motion vector detection can be performed on a plurality of key frames extracted from the video to obtain a motion change value for characterizing the motion change degree of the video main body between every two key frames, and / or structural similarity detection can be performed on the plurality of key frames to obtain the structural similarity between every two key frames, and target key frames can be determined from the plurality of key frames according to the motion change value and / or the structural similarity, and a video thumbnail of the video can be generated according to the target key frames. Thus, when generating a video thumbnail, only the key frames in the video need to be extracted and processed, thereby reducing the data processing amount. On the one hand, the generation efficiency of the video thumbnail can be improved, and on the other hand, the occupation of computing resources can be reduced, and the server performance can be improved. In addition, since the motion change value and / or the structural similarity between key frames can reflect the content change degree between key frames, thus, when determining target key frames according to the motion change value and / or the structural similarity between key frames and generating a video thumbnail of the video based on the target key frames, the generation of redundant thumbnails can be reduced, and the generation quality of the video thumbnail can be improved.
[0098] In a possible manner, the determination module 903 may be configured to determine, from the plurality of key frames, the key frames with a motion change value greater than or equal to a preset motion threshold as target key frames; and / or determine, from the plurality of key frames, the key frames with a structural similarity less than or equal to a preset similarity threshold as target key frames.
[0099] In a possible manner, different types of videos correspond to different preset motion thresholds, and / or different types of videos correspond to different preset similarity thresholds.
[0100] In a possible manner, the generation module 904 may include: An analysis unit for performing content analysis on a target key frame to obtain a content analysis result, and determining a first type of key frame and a second type of key frame in the target key frame according to the content analysis result, where the importance level of the first type of key frame is higher than that of the second type of key frame; A first generation unit for generating a first video thumbnail of the video according to the first type of key frame; A second generation unit for generating a second video thumbnail of the video according to the second type of key frame, where the resolution of the first video thumbnail is greater than that of the second video thumbnail.
[0101] In a possible manner, there are multiple video thumbnails. Correspondingly, the video thumbnail generation device 900 may further include: A first processing module for inserting any one of the multiple video thumbnails into an initially blank canvas, and taking the area of the canvas where the thumbnail to be processed is inserted as the initial parent node, and cyclically executing the following process: Taking any one of the multiple video thumbnails that has not been inserted into the canvas as the thumbnail to be processed, dividing the blank area in the canvas into child nodes of the parent node, inserting the video thumbnail to be processed into the child nodes, and taking the child nodes as the new parent node until all the multiple video thumbnails are inserted into the canvas to obtain a spliced image corresponding to the multiple video thumbnails.
[0102] In a possible manner, the first processing module may be used to divide the blank area in the canvas into child nodes of the parent node according to the size information of the blank area in the canvas and the size information of the thumbnail to be processed, so that the canvas area corresponding to the child nodes can accommodate the thumbnail to be processed.
[0103] In a possible manner, the first processing module may include: A first processing unit for, when the first video thumbnail in the spliced image meets a preset condition, backtracking up the parent node from the node corresponding to the first video thumbnail to obtain the target parent node corresponding to the first video thumbnail; A second processing unit for, when there is an idle node among the child nodes of the target parent node, merging the node corresponding to the first video thumbnail with the idle node, where the idle node is a node into which no video thumbnail has been inserted.
[0104] In a possible manner, the video thumbnail generation device 900 may further include: A detection module for performing change detection on the spliced image; A second processing module, configured to determine differential block information that has changed in a second video thumbnail in the stitched image when it is detected that the second video thumbnail in the stitched image has changed, and generate a differential packet according to the differential block information, where the differential packet is used for a client to update the stitched image stored locally.
[0105] Based on the same concept, an embodiment of the present disclosure further provides a video thumbnail display device, as Figure 10 shown, the video thumbnail display device 100 may include: An acquisition module, configured to acquire a stitched image from a server, where the stitched image includes a video thumbnail, the video thumbnail is generated by the server according to target key frames in a video, and the target key frames are determined by the server according to a motion change value and / or a structural similarity between every two key frames in the video, and the motion change value is used to characterize the degree of motion change of the video main body between every two key frames; A display module, configured to determine a target video thumbnail from the stitched image for display in response to a control operation on the playback progress of the video.
[0106] In a possible manner, the video thumbnail in the stitched image has a start display time and an end display time, and the video thumbnails in the stitched image are stored in sequence according to the start display time. Correspondingly, the display module may include: A first determination unit, configured to determine a third video thumbnail in the stitched image in response to an operation of dragging the playback progress of the video to a first time, where the time stamp of the target key frame corresponding to the third video thumbnail is the closest to the first time; A second determination unit, configured to determine a time ratio of the first time relative to the start display time of the third video thumbnail according to the first time, the start display time, and the end display time of the third video thumbnail; A first display unit, configured to intercept a target video thumbnail from the third video thumbnail and a fourth video thumbnail for display according to the time ratio and relevant information of the third video thumbnail, where the fourth video thumbnail is the next video thumbnail stored after the third video thumbnail, and the relevant information includes the size information and coordinate position information of the third video thumbnail.
[0107] In a possible manner, the video thumbnails in the stitched image are stored in sequence along the horizontal direction according to the start display time. The first display unit may include: A first processing subunit, configured to multiply the time ratio by the width of the third video thumbnail to obtain a coordinate offset; A first processing subunit, configured to add the abscissa of the third video thumbnail to the coordinate offset to obtain a target abscissa; A display subunit, configured to, from the target abscissa, intercept a target video thumbnail with a width equal to that of the third video thumbnail from among the third video thumbnail and the fourth video thumbnail for display.
[0108] In a possible manner, the display module may be configured to, in response to an operation of dragging the playback progress of the video to the second moment, load the video thumbnail corresponding to the second moment from the spliced image for display; Correspondingly, the video thumbnail display device may further include: A loading module, configured to, when loading the video thumbnail corresponding to the second moment, preload a preset number of video thumbnails to be displayed after the second moment from the spliced image.
[0109] In a possible manner, the video thumbnail display device may further include: A dragging module, configured to, in response to a dragging operation of dragging the playback progress of the video to the third moment, determine the dragging speed corresponding to the dragging operation; A processing module, configured to, when the dragging speed is greater than or equal to a preset dragging speed threshold, sample the video to obtain a video frame at the third moment, and render and display the video thumbnail at the third moment according to the video frame at the third moment.
[0110] In a possible manner, the video thumbnail display device may further include: A verification module, configured to verify the spliced image to obtain a verification result indicating whether the spliced image is available; Correspondingly, the display module may be configured to, when the verification result indicates that the spliced image is available, in response to a control operation of the playback progress of the video, determine a target video thumbnail from the spliced image for display.
[0111] Based on the same concept, an embodiment of the present disclosure further provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processing device, the steps of any of the above video thumbnail generation methods or video thumbnail display methods are implemented.
[0112] Based on the same concept, an embodiment of the present disclosure further provides an electronic device, which may include: A storage device, on which a computer program is stored; A processing device, configured to execute the computer program in the storage device to implement the steps of any of the above video thumbnail generation methods or video thumbnail display methods.
[0113] Based on the same concept, an embodiment of the present disclosure further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any of the above video thumbnail generation methods or video thumbnail display methods are implemented.
[0114] Refer to the following Figure 11 , which shows a schematic structural diagram of an electronic device 1100 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 11 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0115] As Figure 11 shown, the electronic device 1100 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 1101, which may perform various appropriate actions and processes according to the programs stored in the read-only memory (ROM) 1102 or the programs loaded from the storage device 1108 into the random access memory (RAM) 1103. In the RAM 1103, various programs and data required for the operation of the electronic device 1100 are also stored. The processing device 1101, the ROM 1102, and the RAM 1103 are connected to each other through a bus 1104. The input / output (I / O) interface 1105 is also connected to the bus 1104.
[0116] Generally, the following devices may be connected to the I / O interface 1105: an input device 1106 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1107 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1108 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1109. The communication device 1109 may allow the electronic device 1100 to communicate with other devices wirelessly or wirelesly to exchange data. Although Figure 11 the electronic device 1100 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.
[0117] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowchart can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication device 1109, or installed from the storage device 1108, or installed from the ROM 1102. When the computer program is executed by the processing device 1101, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0118] It should be noted that the computer-readable medium described above in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0119] In some embodiments, any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol) can be used for communication, and it can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed network.
[0120] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.
[0121] The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to: extract key frames from the video to obtain a plurality of key frames; perform motion vector detection on the plurality of key frames to obtain a motion change value for characterizing the degree of motion change of the video subject between every two key frames, and / or perform structural similarity detection on the plurality of key frames to obtain the structural similarity between every two key frames; determine target key frames from the plurality of key frames according to the motion change value and / or the structural similarity; generate a video thumbnail of the video according to the target key frames.
[0122] Alternatively, the above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to: obtain a spliced image from a server, where the spliced image includes a video thumbnail, the video thumbnail is generated by the server according to the target key frames in the video, the target key frames are determined by the server according to the motion change value and / or the structural similarity between every two key frames in the video, and the motion change value is used to characterize the degree of motion change of the video subject between every two key frames; in response to a control operation on the playback progress of the video, determine a target video thumbnail from the spliced image for display.
[0123] Computer program code for performing the operations of this disclosure may be written in one or more programming languages or combinations thereof. The foregoing programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network including a local area network (LAN) or a wide area network (WAN), or alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0124] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0125] The modules described in the embodiments of the present disclosure may be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the module itself.
[0126] The functions described above herein may be performed, at least in part, by one or more hardware logic components. By way of example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and the like.
[0127] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0128] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.
[0129] Furthermore, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0130] Although the subject matter has been described in language specific to structural features and / or methodological act logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated herein.
Claims
1. A method for generating video thumbnails, characterized in that, Applied to the server side, the video thumbnail generation method includes: Extract key frames from the video to obtain a plurality of key frames; Perform motion vector detection on the plurality of key frames to obtain a motion change value for characterizing the motion change degree of the video main body between every two of the key frames, and / or, perform structural similarity detection on the plurality of key frames to obtain the structural similarity between every two of the key frames; Determine target key frames from the plurality of key frames according to the motion change value and / or the structural similarity; Generate a video thumbnail of the video according to the target key frames.
2. The video thumbnail generation method according to claim 1, wherein The determining target key frames from the plurality of key frames according to the motion change value and / or the structural similarity includes: Among the plurality of key frames, determine the key frames with the motion change value greater than or equal to a preset motion threshold as target key frames; and / or, Among the plurality of key frames, determine the key frames with the structural similarity less than or equal to a preset similarity threshold as target key frames.
3. The method for generating a video thumbnail according to claim 2, wherein Different types of videos correspond to different preset motion thresholds, and / or, different types of videos correspond to different preset similarity thresholds.
4. The video thumbnail generation method according to any one of claims 1 to 3, characterized in that, The generating a video thumbnail of the video according to the target key frames includes: Perform content analysis on the target key frames to obtain a content analysis result, and according to the content analysis result, determine a first type of key frame and a second type of key frame among the target key frames, wherein the importance degree of the first type of key frame is higher than that of the second type of key frame; Generate a first video thumbnail of the video according to the first type of key frame; Generate a second video thumbnail of the video according to the second type of key frame, wherein the resolution of the first video thumbnail is greater than that of the second video thumbnail.
5. The video thumbnail generation method according to any one of claims 1-3, characterized in that, There are a plurality of the video thumbnails, and the video thumbnail generation method further includes: Insert any one of the plurality of video thumbnails into an initially blank canvas, and use the area of the canvas where the video thumbnail is inserted as the initial parent node, and loop to execute the following process: Use any one of the plurality of video thumbnails that has not been inserted into the canvas as the thumbnail to be processed, divide the blank area in the canvas into child nodes of the parent node, insert the thumbnail to be processed into the child nodes, and use the child nodes as the new parent node until all the plurality of video thumbnails are inserted into the canvas to obtain a spliced image corresponding to the plurality of video thumbnails.
6. The video thumbnail generation method according to claim 5, wherein The dividing the blank area in the canvas into child nodes of the parent node includes: According to the size information of the blank area in the canvas and the size information of the thumbnail to be processed, divide the blank area in the canvas into child nodes of the parent node so that the canvas area corresponding to the child nodes can accommodate the thumbnail to be processed.
7. The method for generating a video thumbnail according to claim 5, wherein The video thumbnail generation method further includes: When the first video thumbnail in the stitched graph meets the preset conditions, trace back to the parent node upward from the node corresponding to the first video thumbnail to obtain the target parent node corresponding to the first video thumbnail; When there is an idle node among the child nodes of the target parent node, merge the node corresponding to the first video thumbnail with the idle node, where the idle node is a node into which no video thumbnail has been inserted.
8. The video thumbnail generation method according to claim 5, wherein The video thumbnail generation method further includes: Performing change detection on the stitched graph; When it is detected that there is a change in the second video thumbnail in the stitched graph, determining the differential block information that has changed in the second video thumbnail, and generating a differential package according to the differential block information, where the differential package is used for the client to update the locally stored stitched graph.
9. A method for displaying video thumbnails, characterized in that, Applied to the client, the video thumbnail display method includes: Obtaining a stitched graph from the server, where the stitched graph includes video thumbnails, the video thumbnails are generated by the server according to the target key frames in the video, the target key frames are determined by the server according to the motion change value and / or structural similarity between every two key frames in the video, and the motion change value is used to characterize the degree of motion change of the video main body between every two key frames; In response to a control operation on the playback progress of the video, determining a target video thumbnail from the stitched graph for display.
10. The video thumbnail display method according to claim 9, characterized in that, The video thumbnails in the stitched graph have a start display time and an end display time, and the video thumbnails in the stitched graph are stored in sequence according to the start display time. The determining a target video thumbnail from the stitched graph for display in response to a control operation on the playback progress of the target video includes: In response to an operation of dragging the playback progress of the video to the first time, determining a third video thumbnail in the stitched graph, where the time stamp of the target key frame corresponding to the third video thumbnail is the closest to the first time; Determining the time ratio of the first time relative to the start display time of the third video thumbnail according to the first time, the start display time and the end display time of the third video thumbnail; Intercepting a target video thumbnail for display from the third video thumbnail and the fourth video thumbnail according to the time ratio and the relevant information of the third video thumbnail, where the fourth video thumbnail is the next video thumbnail stored after the third video thumbnail, and the relevant information includes the size information and coordinate position information of the third video thumbnail.
11. The video thumbnail display method according to claim 10, characterized in that The video thumbnails in the stitched graph are stored in sequence along the horizontal direction according to the start display time. The intercepting a target video thumbnail for display from the third video thumbnail and the fourth video thumbnail according to the time ratio and the display information of the third video thumbnail in the stitched graph includes: Multiplying the time ratio by the width of the third video thumbnail to obtain a coordinate offset; Adding the abscissa of the third video thumbnail to the coordinate offset to obtain the target abscissa; In the third video thumbnail and the fourth video thumbnail, a target video thumbnail with a width equal to the width of the third video thumbnail is intercepted starting from the target abscissa for display.
12. The video thumbnail display method according to claim 9, wherein The determining a target video thumbnail from the spliced image for display in response to a control operation on the playback progress of the video includes: In response to an operation of dragging the playback progress of the video to the second moment, loading the video thumbnail corresponding to the second moment from the spliced image for display; The video thumbnail display method further includes: When loading the video thumbnail corresponding to the second moment, preloading a preset number of video thumbnails to be displayed after the second moment from the spliced image.
13. The video thumbnail display method according to claim 9, wherein The video thumbnail display method further includes: In response to a drag operation of dragging the playback progress of the video to the third moment, determining the drag speed corresponding to the drag operation; In a case where the drag speed is greater than or equal to a preset drag speed threshold, sampling the video to obtain a video frame at the third moment, and rendering the video thumbnail at the third moment for display according to the video frame at the third moment.
14. The video thumbnail display method according to claim 13, wherein The video thumbnail display method further includes: Verifying the spliced image to obtain a verification result for indicating whether the spliced image is available; The determining a target video thumbnail from the spliced image for display in response to a control operation on the playback progress of the video includes: In a case where the verification result indicates that the spliced image is available, determining a target video thumbnail from the spliced image for display in response to a control operation on the playback progress of the video.
15. A video thumbnail generation device, characterized in that, The video thumbnail generation device includes: An extraction module, configured to extract key frames from a video to obtain a plurality of key frames; A detection module, configured to perform motion vector detection on the plurality of key frames to obtain a motion change value for characterizing the motion change degree of the video subject between every two of the key frames, and / or perform structural similarity detection on the plurality of key frames to obtain the structural similarity between every two of the key frames; A determination module, configured to determine target key frames from the plurality of key frames according to the motion change value and / or the structural similarity; A generation module, configured to generate a video thumbnail of the video according to the target key frames.
16. A video thumbnail display device, characterized in that The video thumbnail display device includes: An acquisition module, configured to acquire a spliced image from a server, where the spliced image includes video thumbnails, the video thumbnails are generated by the server according to target key frames in the video, the target key frames are determined by the server according to the motion change value and / or the structural similarity between every two key frames in the video, and the motion change value is used to characterize the motion change degree of the video subject between every two of the key frames; A display module, configured to determine a target video thumbnail from the spliced image for display in response to a control operation on the playback progress of the video.
17. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processing device, the steps of the method according to any one of claims 1-14 are implemented.
18. An electronic device, characterized in that, including: A storage device, on which a computer program is stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-14.
19. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1-14 are implemented.