Method and apparatus for video playback at variable speed
By separating video into image and audio streams, calculating similarity, discarding image frames, and superimposing audio frames, the problem of discontinuous visuals and unnatural audio during video playback at double speed is solved, providing a natural and smooth user experience and reducing resource consumption.
Patent Information
- Application Number
- CN202311412211.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-27
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2043-10-27
AI Technical Summary
Existing video speed-up technology results in discontinuous video footage, unnatural audio, poor user experience, and high resource consumption.
The video is separated into image streams and audio streams. Similar frames are discarded by calculating the similarity of image frames, and similar frames are superimposed in the audio stream. The two streams are then merged into a video at double speed.
It achieves a natural and smooth visual and auditory experience, reduces resource consumption, and is suitable for devices with poor performance.
Smart Images

Figure CN117478964B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of multimedia, in particular to the field of video processing, and specifically to a method and device for video playback at a speed multiple. BACKGROUND
[0002] With the wide application of Internet technology in the field of multimedia, more and more users choose to watch online movies, teaching courses or live interaction, etc. Since the users have different levels of interest in different video segments during the video watching process, video playback at a speed multiple can facilitate the users to concentrate on the interesting segments. For example, when a teacher records a course, the content of the course is rich and the speed is relatively slow in order to accommodate most students. However, for some students who have a good foundation, they want to improve the learning efficiency by increasing the playback speed when learning through the video. Therefore, the playback at a speed multiple is beneficial to improve the learning efficiency of the students. Therefore, all kinds of players on the current market have the function of playback at a speed multiple to meet the various needs of users such as fast browsing or slow appreciation.
[0003] In the prior art, the playback at a speed multiple of a video is achieved by removing a certain proportion of image frames and audio frames (e.g. 1 / 2 of the data is removed for 2 times speed) while keeping the frame rate of the video unchanged, so that the video is played at a certain speed multiple. Randomly removing image frames of the video is easy to remove key information, resulting in an incoherent and unnatural picture and a poor viewing experience.
[0004] The algorithm used in the processing of the audio is easy to change the interval between frames (i.e. change the overlap between frames), which can make the user obviously perceive the fast-forward and end of the audio, and easily introduce noise, resulting in a poor speed multiple effect. SUMMARY
[0005] The present disclosure provides a method, device, equipment, storage medium and computer program product for video playback at a speed multiple.
[0006] According to a first aspect of the present disclosure, a method for video playback at a speed multiple is provided, comprising: separating a video into an image stream and an audio stream; removing similar image frames in each unit time in the image stream by a predetermined proportion to obtain a new image stream; superimposing similar audio frames in a set of audio frames obtained by framing the audio stream by a predetermined proportion to obtain a new audio stream; and combining the new image stream and the new audio stream together to form a video after speed multiple.
[0007] According to a second aspect of the present disclosure, a device for video playback at a speed is provided, comprising: a separation unit configured to separate a video into an image stream and an audio stream; a discarding unit configured to discard similar image frames in a predetermined proportion in each unit time in the image stream to obtain a new image stream; a superimposing unit configured to superimpose similar audio frames in a set of audio frames obtained by framing the audio stream in a predetermined proportion to obtain a new audio stream; and a merging unit configured to merge the new image stream and the new audio stream together to form a video after speed-up.
[0008] According to a third aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of the first aspect.
[0009] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to perform the method of any one of the first aspect.
[0010] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method of any one of the first aspect.
[0011] The method and device for video playback at a speed provided by the embodiments of the present disclosure have particularly good robustness in the case of editing, splicing, rotating, implanting advertisements, bullet screens, adding logos, etc., so that the picture is very natural and smooth. By searching for the most similar audio frame for superimposition, the voice is natural, smooth, and has little noise. Thus, the video after speed-up can have natural and smooth user experience in both vision and hearing. The algorithm is very efficient and does not consume too many resources, and can have good effect on devices with poor performance.
[0012] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0013] The accompanying drawings are used to better understand the present scheme and do not limit the present disclosure. Among them:
[0014] Figure 1 is an exemplary system architecture diagram to which an embodiment of the present disclosure can be applied;
[0015] Figure 2 is a flowchart of one embodiment of a method of video fast play according to the present disclosure;
[0016] Figures 3a-3d is a schematic diagram of one application scenario of a method of video fast play according to the present disclosure;
[0017] Figure 4 is a flowchart of yet another embodiment of a method of video fast play according to the present disclosure;
[0018] Figure 5 is a structural schematic diagram of one embodiment of an apparatus of video fast play according to the present disclosure;
[0019] Figure 6 is a structural schematic diagram of a computer system of an electronic device suitable for implementing embodiments of the present disclosure. DETAILED DESCRIPTION
[0020] Exemplary embodiments of the present disclosure are described herein with reference to the accompanying drawings, which are incorporated in this specification, wherein various details of the embodiments of the present disclosure are set forth in order to provide an overall understanding of the present disclosure. It should be understood that the various details of the embodiments of the present disclosure can be combined or eliminated in various manners, without departing from the scope and spirit of the present disclosure. Also, for the sake of brevity and clarity, descriptions of well-known functions and structures can be omitted in the following description.
[0021] Figure 1 An exemplary system architecture 100 to which embodiments of the method of video fast play or the apparatus of video fast play of the present disclosure can be applied is shown.
[0022] As shown in Figure 1 , the system architecture 100 can include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is a medium to provide communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 can include various connection types, such as wired, wireless communication links or fiber optic cables, etc.
[0023] A user can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the terminal devices 101, 102, 103, such as a video player, a web browser application, a shopping application, a search application, an instant messaging tool, a mailbox client, a social platform software, etc.
[0024] The terminal device 101, 102, 103 can be hardware or software. When the terminal device 101, 102, 103 is hardware, it can be various electronic devices with a display screen and supporting video playing, including but not limited to a smart phone, a tablet computer, an electronic book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer, a desktop computer, and the like. When the terminal device 101, 102, 103 is software, it can be installed in the above-listed electronic devices. It can be implemented as multiple software or software modules (for example, to provide distributed services) or as a single software or software module. No specific limitation is made herein.
[0025] The server 105 can be a server providing various services, for example, a background video server supporting a video displayed on the terminal device 101, 102, 103. The background video server can analyze and process received data such as a fast-forward playing request, and feed back the processing result (for example, a video after fast-forward playing) to the terminal device.
[0026] It should be noted that the server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers or as a single server. When the server is software, it can be implemented as multiple software or software modules (for example, multiple software or software modules to provide distributed services) or as a single software or software module. No specific limitation is made herein. The server can also be a server of a distributed system or a server combined with a blockchain. The server can also be a cloud server or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0027] It should be noted that the method for video fast-forward playing provided by the embodiments of the present disclosure can be executed by the terminal device 101, 102, 103 or the server 105. Correspondingly, the device for video fast-forward playing can be arranged in the terminal device 101, 102, 103 or the server 105. No specific limitation is made herein.
[0028] It should be understood that Figure 1 The number of terminal devices, networks, and servers in
[0029] With reference back to Figure 2, shows a flow 200 of one embodiment of the method of video speed play according to the present disclosure. The method of video speed play comprises the following steps:
[0030] Step 201, separating a video into an image stream and an audio stream.
[0031] In the present embodiment, the video player installed in the terminal device as shown in the figure can separate the loaded video into an image stream and an audio stream. Then the image stream and the audio stream are processed respectively, which can be processed in parallel, not limited to the order of steps listed below. Figure 1
[0032] Step 202, discarding similar image frames in the image stream by a predetermined proportion in each unit time to obtain a new image stream. In the present embodiment, features can be extracted from each image frame, such as trace transformation, image hash value, sift feature vector, etc. By comparing the similarity of image frames in a unit time (such as 1-2 seconds), if the similarity of adjacent image frames is within a threshold, they are considered as similar frames, then when frames are discarded, part of the frames can be selected for loss, so that the key frames will not be lost and the picture will be smooth. For example, the unit time is 1 second, and there are 30 image frames in total. Starting from the first image frame, the similarity between each image frame and the adjacent image frame is calculated. If the similarity is greater than a predetermined threshold, they are grouped together and the similarity with the next image frame is calculated. Multiple image frames with a similarity higher than the predetermined threshold are grouped into a group. Each group of image frames is named an image frame subset, and the number of image frames in each image frame subset is uncertain. In the present embodiment, the predetermined proportion is positively correlated with the speed value. If the speed value is 2, half of the image frames are discarded. Half of the image frames in each group of image frame subsets are discarded. They can be discarded randomly or at fixed intervals. The number of image frames retained in each group of image frame subsets is inversely proportional to the speed value.
[0033] Step 203, superimposing similar audio frames in the audio frame set obtained by framing the audio stream by a predetermined proportion to obtain a new audio stream.
[0034] In the present embodiment, the audio can be framed every 20-50 ms (the previous frame and the next frame are superimposed at 1 / 2 position). In the present embodiment, the most similar signal frame can be found by using the correlation peak algorithm or AMDF (average magnitude difference function method). Then according to the specified speed value, the two audio frames are windowed and superimposed, for example, 50% of the waveform needs to be superimposed for 2x speed play.
[0035] Step 204, merging the new image stream and the new audio stream together to form a video after speed play.
[0036] In this embodiment, the processed audio stream and the image stream are merged together to obtain a speeded-up video, contrary to the operation of step 201.
[0037] The method provided by the above embodiments of the present disclosure has particularly good robustness in the case of clipping, splicing, rotating, implanting advertisements, bullet screens, adding logos, etc., so that the picture can be very natural and smooth. By searching for the most similar audio frame for superposition, the voice is natural, smooth, and has little noise. Thus, the speeded-up video can have natural and smooth user experience in terms of vision and hearing. The algorithm is highly efficient and does not consume too many resources, and can have good effect on devices with poor performance.
[0038] Figure 3a is a schematic diagram of an application scenario of the method for speeded-up video playback according to the present embodiment. In the application scenario of Figure 3a The data stream of the original video is separated into an image stream and an audio stream. Image frames are extracted from the image stream, similar image frames are grouped by a picture similarity algorithm, and after discarding a part of the image frames from each group of similar image frames, a new image stream is obtained. Audio frames are extracted from the audio stream, and similar audios are found by an audio similarity algorithm for superposition to obtain a new audio stream. Finally, the new image stream and the new audio stream are merged to obtain a speeded-up video.
[0039] In some optional implementations of the present embodiment, the discarding of similar image frames in each unit time in the image stream by a predetermined proportion to obtain a new image stream comprises: similarity calculation is performed on the image frames in each unit time in the image stream, and image frames with a similarity higher than a predetermined threshold are grouped into a group to obtain at least one subset of image frames; and according to a specified speed-up value, a corresponding proportion of image frames are selected from each subset of image frames for discarding to obtain a new image stream. Similarity calculation can be performed on the image frames in each unit time in the image stream by a picture similarity algorithm. The predetermined proportion is positively correlated with the speed-up value. Similar image frames can be discarded at equal intervals.
[0040] In some optional implementations of the present embodiment, the superposition of similar audio frames in the set of audio frames obtained by framing the audio stream by a predetermined proportion to obtain a new audio stream comprises: the audio stream is framed to obtain a set of audio frames; for each audio frame, the next most similar signal frame of the audio frame is searched from the set of audio frames, and the audio frame and the signal frame are superposed by a corresponding proportion according to a specified speed-up value to obtain a new audio stream.
[0041] In some optional implementations of the embodiment, the similarity calculation on the image frames in each unit time in the image stream includes: converting the image frames in the image stream into grayscale images; calculating the hash value of each grayscale image; and calculating the distance between the hash values of adjacent grayscale images as the similarity between the adjacent grayscale images for the grayscale images in each unit time. First, each image frame is converted into a luminance color space picture, i.e., from RGB to YUV, to obtain a grayscale image. Then, the hash value of each grayscale image is calculated by using an average hash, a perceptual hash, a difference hash, or the like. Finally, the similarity between adjacent grayscale images is calculated by using a distance calculation formula (e.g., a Hamming distance, an Euclidean distance, or the like). The similarity is calculated by using the hash value, which can improve the calculation speed and accuracy. The requirement for the hardware device is low, and the method can be widely applied to various terminal devices. The corresponding hash algorithm can be selected according to the performance of the hardware device, for example, a simple average hash algorithm or a PDQ algorithm can be selected for a terminal device with low CPU or GPU performance, and a complex difference hash algorithm can be selected for a terminal device with high performance.
[0042] In some optional implementations of the embodiment, before the image frames in the image stream are converted into grayscale images, the method further includes: reducing the image frames in the image stream to a predetermined size. For a million-pixel input picture, the picture size needs to be reduced. For example, the picture size is adjusted to 512x512. The reduced size can also be specified according to the hardware performance of the execution subject, for example, if the CPU or GPU performance is low and the processing time of the image is too long, the picture is reduced to a smaller size, for example, 256x256 or 128x128, or the like. The lower the hardware performance, the smaller the adjusted image size, so as to avoid that the similarity calculation time is too long and the user experience is affected.
[0043] In some optional implementations of the embodiment, the calculation of the hash value of each grayscale image includes: for each grayscale image, calculating the average value of the grayscale values of all pixel points in the grayscale image, recording the pixel points with the grayscale values greater than the average value in the grayscale image as 1, and recording the remaining pixel points as 0 to obtain the hash value of the grayscale image. The average hash algorithm is simple and easy to implement, and is particularly suitable for terminal devices with low performance. The hash value of the grayscale image can be quickly calculated, so as to improve the efficiency of frame loss, reduce the video playing delay, and improve the user experience.
[0044] In some optional implementations of the embodiment, the calculating the hash value of each grayscale image comprises: for each grayscale image, performing two-dimensional discrete cosine transform and retaining a low-frequency matrix of a predetermined size, recording elements greater than 0 in the low-frequency matrix as 1 and recording the rest of the elements as 0, to obtain the hash value of the grayscale image. Two-dimensional discrete cosine transform (DCT) concentrates the low-frequency area of a signal in the upper left corner to facilitate the acquisition of low-frequency information, and the energy of a natural signal is mostly concentrated in the low-frequency area (i.e., the area with a gentle change in attributes, such as an area with a gentle change in pixel value in an image, which often contains main color information, while an area with a sharp change often is a line boundary and a color alternating area), so DCT is often used in the lossy compression of data in signal and image processing. On a 64x64 sampled picture, DCT is performed, and 1-16 time slots are retained in the X and Y directions. A corresponding DCT matrix size can be selected according to the size of the grayscale image. If a 32*32 DCT transform is selected, a 32*32 DCT matrix is obtained, but we only need to retain the 8*8 matrix in the upper left corner, which presents the lowest frequency in the picture. The result cannot tell the low frequency of authenticity, but can roughly tell the relative proportion relative to the 0 frequency. As long as the overall structure of the picture remains unchanged, the hash result value remains unchanged. The influence of the adjustment of gamma correction or a color histogram can be avoided. The method has a faster calculation speed than the median value comparison method.
[0045] In some optional implementations of the embodiment, the calculating the hash value of each grayscale image comprises: for each grayscale image, performing two-dimensional discrete cosine transform and retaining a low-frequency matrix of a predetermined size, calculating the median value of all elements in the low-frequency matrix, recording elements greater than the median value in the low-frequency matrix as 1 and recording the rest of the elements as 0, to obtain the hash value of the grayscale image. This method needs to calculate the median value of all elements in the low-frequency matrix, belongs to the median value comparison method, can obtain a more accurate similarity result, and retains more details. The influence of the adjustment of gamma correction or a color histogram can be avoided.
[0046] In some optional implementations of the embodiment, before the calculating the hash value of each grayscale image, the method further comprises: downsampling each grayscale image to generate a 64x64 sampled picture. In this way, the picture can be reduced, the amount of similarity calculation can be reduced, the video synthesis speed can be improved, the waiting time of the user can be reduced, and the user experience can be improved.
[0047] In some optional implementations of the embodiment, the downsampling each grayscale image comprises: downsampling each grayscale image by using a tent convolution filter to generate a 64x64 sampled picture. The tent convolution filter has no value jump, and does not need to separate the discrete and continuous cases, thereby improving the sampling efficiency.
[0048] Preferably, as Figure 3b The similar pictures are calculated by the following steps:
[0049] 1. Get each image frame.
[0050] 2. Downsize the image frame to a maximum of 512x512 (if it is a megapixel input picture)
[0051] 3. Convert the image frame to a luminance color space picture, i.e. from RGB to YUV
[0052] 4. Convolve filter with tent convolution filter.
[0053] 5. Downsample to generate a 64x64 sample picture
[0054] 6. Perform a two-dimensional discrete cosine transform (DCT) on the 64x64 sample picture and retain 1-16 bins in the X and Y directions; from the 16x16 DCT output, calculate the median in the transform space.
[0055] 7. For each 16x16 bin of the output hash, if the corresponding element of the transform space is greater than the median, emit a 1, otherwise emit a 0; the final hash value is 256 bits.
[0056] 8. Read the binary values from the lower right corner to the upper left corner.
[0057] 9. Convert the binary values to hexadecimal values to obtain the hash value.
[0058] 10. Calculate the Hamming distance of the hash of the adjacent pictures, and consider them as similar pictures within a certain threshold (usually 0-32).
[0059] Further reference is made to Figure 4 which shows a flow 400 of yet another embodiment of a method of video fast-forwarding. The flow 400 of the method of video fast-forwarding comprises the following steps:
[0060] Step 401 separates a video into an image stream and an audio stream.
[0061] Step 402 discards similar image frames in the image stream by a predetermined ratio per unit time to obtain a new image stream.
[0062] Step 403 frames the audio stream to obtain a set of audio frames.
[0063] Steps 401-403 are basically the same as steps 201-203, and thus will not be described again.
[0064] Step 404: For each audio frame, window the audio frame to obtain the first frame. Select the second frame from the audio frame set within the first range whose phase parameters are aligned with the first frame. Find the third frame from the audio frame set within the second range that is most similar to the second frame. Window the third frame to obtain the signal frame. Then, according to the specified speed value, superimpose the audio frame and the signal frame in the corresponding proportion to obtain a new audio stream.
[0065] In this embodiment, as Figure 3c As shown, the audio similarity detection - WSOLA waveform similarity overlay algorithm:
[0066] a) Extract the first frame x′ from the original audio. m Then, the frame is windowed and output to the y signal to obtain y. m .
[0067] b) In relation to x′ m Find the distance from Hs to x′ m The second frame with the same phase parameters
[0068] c) In relation to x′ m A frame x was found at a distance of Ha. m+1 In x m+1 Centered on the center, extending Δ on both sides max area Find with The most similar third frame x′ m+1 .
[0069] d) For the third frame x′ m+1 After windowing, the output is fed to the y signal to obtain y. m+1 . y m With y m+1 Superposition is performed at Hs.
[0070] Speed multiplier = Ha / Hs, where Hs is fixed and usually half the frame length. The speed multiplier is used to control whether the output signal is fast or slow. When the speed multiplier is less than 1, the output signal becomes slower; otherwise, it becomes faster.
[0071] The effect after superposition is as follows Figure 3d As shown.
[0072] Step 405: Merge the new image stream and the new audio stream together to form a video at double speed.
[0073] Step 405 is basically the same as step 204, so it will not be described again.
[0074] from Figure 4 It can be seen from this that, with Figure 2Compared with the corresponding embodiment, the flow 400 of the method for video speed playback in the embodiment embodies the step of using the WSOLA (Waveform Similarity Overlap Add) algorithm to merge the audio frames. Thus, the scheme described in the embodiment can achieve the technical effect of variable speed and constant pitch of audio. Moreover, the algorithm has high effect and does not consume too many resources, and can have good effect on devices with poor performance.
[0075] Further referring to Figure 5 , as an implementation of the method shown in the above figures, the disclosure provides an embodiment of a device for video speed playback, which corresponds to the method embodiment shown in Figure 2 . The device can be specifically applied to various electronic devices.
[0076] As shown in Figure 5 , the device 500 for video speed playback in the embodiment includes a separation unit 501, a discard unit 502, an overlap unit 503, and a merging unit 504. The separation unit 501 is configured to separate a video into an image stream and an audio stream; the discard unit 502 is configured to discard similar image frames in each unit time in the image stream by a predetermined proportion to obtain a new image stream; the overlap unit 503 is configured to overlap similar audio frames in an audio frame set obtained by framing the audio stream by a predetermined proportion to obtain a new audio stream; and the merging unit 504 is configured to merge the new image stream and the new audio stream together to form a video after speed-up.
[0077] In the embodiment, the specific processing of the separation unit 501, the discard unit 502, the overlap unit 503, and the merging unit 504 of the device 500 for video speed playback can refer to steps 201-204 in the corresponding embodiment. Figure 2
[0078] In some optional implementation manners of the embodiment, the discard unit 502 is further configured to: perform similarity calculation on the image frames in each unit time in the image stream, and divide the image frames with a similarity higher than a predetermined threshold into a group to obtain at least one image frame subset; and select image frames of a corresponding proportion from each image frame subset according to a specified speed-up value to obtain a new image stream.
[0079] In some optional implementation manners of the embodiment, the overlap unit 503 is further configured to: frame the audio stream to obtain an audio frame set; for each audio frame, find the next most similar signal frame of the audio frame from the audio frame set, and overlap the audio frame and the signal frame by a corresponding proportion according to a specified speed-up value to obtain a new audio stream.
[0080] In some optional implementation of the embodiment, the discarding unit 502 is further configured to: convert the image frames in the image stream into grayscale images; calculate hash values of each grayscale image; and calculate distances between hash values of adjacent grayscale images as similarities between adjacent grayscale images for each grayscale image in a unit time.
[0081] In some optional implementation of the embodiment, the apparatus 500 further comprises a scaling unit (not shown in the figure) configured to: downsize the image frames in the image stream to a predetermined size before converting the image frames in the image stream into grayscale images.
[0082] In some optional implementation of the embodiment, the discarding unit 502 is further configured to: for each grayscale image, calculate an average value of grayscale values of all pixel points in the grayscale image, record pixel points with grayscale values greater than the average value in the grayscale image as 1 and record the rest of the pixel points as 0 to obtain a hash value of the grayscale image.
[0083] In some optional implementation of the embodiment, the discarding unit 502 is further configured to: for each grayscale image, perform two-dimensional discrete cosine transform and retain a low-frequency matrix of a predetermined size, record elements greater than 0 in the low-frequency matrix as 1 and record the rest of the elements as 0 to obtain a hash value of the grayscale image.
[0084] In some optional implementation of the embodiment, the discarding unit 502 is further configured to: for each grayscale image, perform two-dimensional discrete cosine transform and retain a low-frequency matrix of a predetermined size, calculate a median value of all elements in the low-frequency matrix, record elements greater than the median value in the low-frequency matrix as 1 and record the rest of the elements as 0 to obtain a hash value of the grayscale image.
[0085] In some optional implementation of the embodiment, the apparatus 500 further comprises a downsampling unit (not shown in the figure) configured to: downsample each grayscale image to generate a 64x64 sample picture before the calculating the hash value of each grayscale image.
[0086] In some optional implementation of the embodiment, the downsampling unit is further configured to: downsample each grayscale image by using a tent convolution filter to generate a 64x64 sample picture.
[0087] In some optional implementation of the embodiment, the superimposing unit 503 is further configured to: perform windowing processing on the audio frame to obtain a first frame; select a second frame with a phase parameter aligned with the first frame in a first range from the audio frame set; find a third frame most similar to the second frame in a second range from the audio frame set; and perform windowing processing on the third frame to obtain a signal frame.
[0088] In the technical solutions of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information comply with relevant laws and regulations and do not violate public order and good customs.
[0089] According to embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.
[0090] An electronic device comprises at least one processor and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described in flow 200 or 400.
[0091] A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable the computer to perform the method described in flow 200 or 400.
[0092] A computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the method described in flow 200 or 400.
[0093] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.
[0094] As shown in Figure 6 The device 600 includes a computing unit 601 that can perform various appropriate actions and processes according to computer programs stored in a read-only memory (ROM) 602 or loaded into a random access memory (RAM) 603 from a storage unit 608. Various programs and data required for the operation of the device 600 can also be stored in the RAM 603. The computing unit 601, the ROM 602 and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0095] A number of components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices through computer networks, such as the Internet, and / or various telecommunication networks.
[0096] The computing unit 601 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 601 performs various methods and processes described above, such as the method of video fast play. For example, in some embodiments, the method of video fast play can be implemented as a computer software program, which is tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded onto the RAM 603 and executed by the computing unit 601, one or more steps of the method of video fast play described above can be performed. Alternatively, in other embodiments, the computing unit 601 can be configured to perform the method of video fast play by any other appropriate means, such as by means of firmware.
[0097] The various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0098] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a standalone software package, or entirely on a remote machine or server.
[0099] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0100] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0101] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0102] The computer system can include clients and servers. This relationship can be. The servers are typically remote from the clients with the interactions between them occurring over a communication network. The relationship between client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The servers can be cloud servers, servers of a distributed system, or servers incorporating blockchain.
[0103] It should be understood that the steps shown in the various forms above can be reordered, added to, or deleted from. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technology disclosed in the present disclosure are achieved, which is not limited herein.
[0104] The specific implementation described above does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements within the spirit and principles of the present disclosure should be included in the protection scope of the present disclosure.
Claims
1. A method for video playback at a speed, comprising: separating a video into an image stream and an audio stream; extracting features from each image frame by any one of trace transformation, image hash value, and sift feature vector, performing similarity calculation on image frames in each unit time in the image stream, and grouping image frames with similarity higher than a predetermined threshold into a group to obtain at least one image frame subset; selecting image frames in a corresponding proportion from each image frame subset according to a specified speed value to be discarded to obtain a new image stream, wherein the predetermined proportion is positively correlated with the speed value, and the discarding comprises random discarding or discarding at a fixed interval; superimposing similar audio frames in an audio frame set obtained by framing the audio stream in a predetermined proportion to obtain a new audio stream, comprising: framing the audio stream to obtain an audio frame set; for each audio frame, finding a next most similar signal frame from the audio frame set, and superimposing the audio frame and the signal frame in a corresponding proportion according to a specified speed value to obtain a new audio stream; merging the new image stream and the new audio stream together to form a video after speed-up.
2. The method of claim 1, wherein, The similarity calculation on image frames in each unit time in the image stream comprises: converting image frames in the image stream into grayscale images; calculating a hash value of each grayscale image; for grayscale images in each unit time, calculating a distance between hash values of adjacent grayscale images as similarity between the adjacent grayscale images.
3. The method of claim 2, wherein, Before converting image frames in the image stream into grayscale images, the method further comprises: reducing image frames in the image stream to a predetermined size.
4. The method of claim 2, wherein, The calculation of the hash value of each grayscale image comprises: for each grayscale image, calculating an average value of grayscale values of all pixel points in the grayscale image, recording pixel points with a grayscale value greater than the average value in the grayscale image as 1, and recording the remaining pixel points as 0 to obtain a hash value of the grayscale image.
5. The method of claim 2, wherein, The calculation of the hash value of each grayscale image comprises: for each grayscale image, performing two-dimensional discrete cosine transformation and retaining a low-frequency matrix of a predetermined size, recording elements greater than 0 in the low-frequency matrix as 1, and recording the remaining elements as 0 to obtain a hash value of the grayscale image.
6. The method of claim 2, wherein, The calculation of the hash value of each grayscale image comprises: for each grayscale image, performing two-dimensional discrete cosine transformation and retaining a low-frequency matrix of a predetermined size, calculating a middle value of all elements in the low-frequency matrix, recording elements greater than the middle value in the low-frequency matrix as 1, and recording the remaining elements as 0 to obtain a hash value of the grayscale image.
7. The method of any one of claims 2-6, wherein, Before the calculation of the hash value of each grayscale image, the method further comprises: down-sampling each grayscale image to generate a 64x64 sample picture.
8. The method of claim 7, wherein, The down-sampling of each grayscale image comprises: down-sampling each grayscale image by using a tent convolution filter to generate a 64x64 sample picture. 9.The method of claim 1, wherein the finding of the next most similar signal frame from the audio frame set comprises: performing windowing processing on the audio frame to obtain a first frame; selecting a second frame with a phase parameter aligned with the first frame within a first range from the audio frame set; finding a third frame most similar to the second frame in a second range from the set of audio frames; windowing the third frame to obtain a signal frame.
10. An apparatus for video playback at a speed, comprising: a separation unit configured to separate a video into an image stream and an audio stream; a discarding unit configured to extract features from each image frame by any one of trace transformation, image hash value, sift feature vector, to calculate the similarity of image frames in each unit time in the image stream, and to group image frames with similarity higher than a predetermined threshold to obtain at least one subset of image frames, and to discard a corresponding proportion of image frames from each subset of image frames according to a specified speed value to obtain a new image stream, wherein the predetermined proportion is positively correlated with the speed value, and the discarding comprises random discarding or discarding at fixed intervals; a superimposition unit configured to superimpose similar audio frames in a set of audio frames obtained by framing the audio stream at a predetermined proportion to obtain a new audio stream, comprising: framing the audio stream to obtain a set of audio frames; for each audio frame, finding a next most similar signal frame from the set of audio frames, and superimposing the audio frame and the signal frame at a corresponding proportion according to a specified speed value to obtain a new audio stream; a merging unit configured to merge the new image stream and the new audio stream together to form a video after speed-up.
11. The apparatus of claim 10, wherein, The discarding unit is further configured to: convert image frames in the image stream into grayscale images; calculate a hash value of each grayscale image; for grayscale images in each unit time, calculate the distance between hash values of adjacent grayscale images as the similarity between adjacent grayscale images.
12. The apparatus of claim 11, wherein, The apparatus further comprises a scaling unit configured to: before converting image frames in the image stream into grayscale images, scale down image frames in the image stream to a predetermined size.
13. The apparatus of claim 11, wherein, The discarding unit is further configured to: for each grayscale image, calculate the average of the grayscale values of all pixel points in the grayscale image, record pixel points with grayscale values greater than the average as 1, and record the remaining pixel points as 0 to obtain the hash value of the grayscale image.
14. The apparatus of claim 11, wherein, The discarding unit is further configured to: for each grayscale image, perform two-dimensional discrete cosine transformation and retain a low-frequency matrix of a predetermined size, record elements greater than 0 in the low-frequency matrix as 1, and record the remaining elements as 0 to obtain the hash value of the grayscale image.
15. The apparatus of claim 11, wherein, The discarding unit is further configured to: for each grayscale image, perform two-dimensional discrete cosine transformation and retain a low-frequency matrix of a predetermined size, calculate the median of all elements in the low-frequency matrix, record elements greater than the median in the low-frequency matrix as 1, and record the remaining elements as 0 to obtain the hash value of the grayscale image.
16. The apparatus of any one of claims 11-15, wherein, The apparatus further comprises a downsampling unit configured to: before calculating the hash value of each grayscale image, downsample each grayscale image to generate a 64x64 sample picture.
17. The apparatus of claim 16, wherein, The downsampling unit is further configured to: downsample each grayscale image using a tent convolution filter to generate a 64x64 sample picture.
18. The apparatus of claim 10, the superposition unit is further configured to: window the audio frame to obtain a first frame; select a second frame from the set of audio frames that has a phase parameter aligned with the first frame within a first range; find a third frame from the set of audio frames that is most similar to the second frame within a second range; window the third frame to obtain a signal frame.
19. An electronic device, comprising: at least one processor; and a memory communicatively connected with the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-9. the computer instructions are for causing the computer to perform the method of any one of claims 1-9.
20. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, 21. A computer program product, comprising a computer program which, when executed by a processor, implements the method of any one of claims 1-9.
Citation Information
Patent Citations
Automatic inductive triggering method and a system thereof based on a hash algorithm
CN109344676A
Audio and video multi-speed playing method and device
CN114339443A
Video information extraction method and device, terminal equipment and computer medium
CN116405745A
Video display method and device therefor
JP1997139913A
Openvg based multi-layer algorithm to determine the position of the nested part
KR1020130037910A