A video frame extraction method, device, equipment and storage medium

By slicing and processing video frames to be extracted in parallel, the problem of video frame extraction time in the prior art is solved, and more efficient resource utilization and faster frame extraction process are achieved.

CN115499662BActive Publication Date: 2025-06-10BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211154134.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-21
Publication Date
2025-06-10
Estimated Expiration
2042-09-21

AI Technical Summary

Technical Problem

The prior art has performance bottlenecks in the process of video frame extraction, which cannot fully utilize equipment resources, resulting in a long time to extract frames.

Method used

By obtaining the GOP length of the video and the frame extraction interval, the video to be processed is sliced ​​to obtain multiple frame extraction fragments, and each fragment is extracted separately according to the frame extraction position, and the system is performed using multiple threads to fully utilize the device resources.

Benefits of technology

It reduces the time-consuming video frame extraction, improves resource utilization, and avoids performance bottlenecks caused by single-thread decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115499662B_ABST
    Figure CN115499662B_ABST
Patent Text Reader

Abstract

The present disclosure provides a video frame extraction method, apparatus, device, and storage medium, which relate to the field of artificial intelligence technology, and particularly relate to fields such as image recognition and video analysis. The specific implementation solution is as follows: obtain the frame extraction interval and the length of the Group of Pictures (GOP) of the video to be processed; determine the frame positions to be extracted from the video to be processed according to the frame extraction interval; slice the video to be processed based on the GOP length to obtain a plurality of frame extraction sub-pieces; and extract frames from each of the frame extraction sub-pieces respectively according to the frame positions to be extracted from the video to be processed. The present disclosure can reduce the time consumption of frame extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and particularly to fields such as image recognition and video analysis. Background Art

[0002] A video refers to a combination of a series of image frames. Since the original data volume of a video is usually large, the original video is generally encoded and compressed first, and then other operations such as storage are performed. During the video frame extraction process, the encoded video needs to be decoded first and then the decoded image frames are extracted. Summary of the Invention

[0003] The present disclosure provides a video frame extraction method, apparatus, device, and storage medium.

[0004] According to a first aspect of the present disclosure, there is provided a video frame extraction method, including:

[0005] Obtaining a frame extraction interval and the length of a group of pictures (GOP) of the video to be processed;

[0006] Determining the positions of frames to be extracted from the video to be processed according to the frame extraction interval;

[0007] Slicing the video to be processed based on the GOP length to obtain a plurality of frame extraction slices;

[0008] Extracting frames from each of the frame extraction slices respectively according to the positions of frames to be extracted from the video to be processed.

[0009] According to a second aspect of the present disclosure, there is provided a video frame extraction apparatus, including:

[0010] A first obtaining module, configured to obtain a frame extraction interval and the length of a group of pictures (GOP) of the video to be processed;

[0011] A determining module, configured to determine the positions of frames to be extracted from the video to be processed according to the frame extraction interval;

[0012] A slicing module, configured to slice the video to be processed based on the GOP length to obtain a plurality of frame extraction slices;

[0013] A frame extraction module, configured to extract frames from each of the frame extraction slices respectively according to the positions of frames to be extracted from the video to be processed.

[0014] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0015] At least one processor; and

[0016] A memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the first aspect.

[0018] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the first aspect.

[0019] According to a fifth aspect of the present disclosure, there is provided a computer program product including a computer program, and the computer program implements the method described in the first aspect when executed by a processor.

[0020] The present disclosure can reduce the frame extraction time.

[0021] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0023] Figure 1 is a flowchart of the video frame extraction method provided by an embodiment of the present disclosure;

[0024] Figure 2 is a schematic diagram of slicing a video to be processed based on the GOP length in an embodiment of the present disclosure;

[0025] Figure 3 is another schematic diagram of slicing a video to be processed based on the GOP length in an embodiment of the present disclosure;

[0026] Figure 4 is yet another schematic diagram of slicing a video to be processed based on the GOP length in an embodiment of the present disclosure;

[0027] Figure 5 is a schematic diagram of frame extraction for a frame extraction slice to be processed in an embodiment of the present disclosure;

[0028] Figure 6 is a schematic structural diagram of a video frame extraction device provided by an embodiment of the present disclosure;

[0029] Figure 7 is a schematic structural diagram of a video frame extraction device provided by an embodiment of the present disclosure;

[0030] Figure 8 is a block diagram of an electronic device for implementing the video frame extraction method of an embodiment of the present disclosure. Detailed Implementation Manner

[0031] The exemplary embodiments of the present disclosure will be described below with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0032] In the related art, usually all the image frames in a video are decoded in chronological order, and then the required image frames are extracted. In this way, even if the electronic device performing video frame extraction has idle resources, only one image frame can be decoded at the same time, which will cause a performance bottleneck, unable to fully utilize the device resources, resulting in a long frame extraction time.

[0033] The video frame extraction provided by the embodiments of the present disclosure can be applied to any video frame extraction scenario. For example, it can be applied in a media cloud scenario.

[0034] The video frame extraction method provided by the embodiments of the present disclosure can be applied to an electronic device. Specifically, the electronic device can be a server, a terminal, etc.

[0035] The embodiments of the present disclosure provide a video frame extraction method, including:

[0036] Obtain the frame extraction interval and the length of the Group of Pictures (GOP) of the video to be processed;

[0037] Determine the frame positions to be extracted from the video to be processed according to the frame extraction interval;

[0038] Slice the video to be processed based on the GOP length to obtain multiple frame extraction slices;

[0039] Extract frames from each frame extraction slice respectively according to the frame positions to be extracted from the video to be processed.

[0040] In the embodiments of the present disclosure, by slicing the video to be processed based on the GOP length, frames can be extracted from the multiple frame extraction slices obtained by slicing according to the frame positions to be extracted from the video to be processed, which can avoid being limited to decoding only one image frame at the same time. In this way, the performance bottleneck can be reduced, the device resources can be fully utilized, and thus the frame extraction time can be reduced. In addition, the resource utilization rate can also be improved.

[0041] Figure 1 is the flowchart of the video frame extraction method provided by the embodiments of the present disclosure. Referring to Figure 1 , the video frame extraction method provided by the embodiments of the present disclosure may include:

[0042] S101, Obtain the frame extraction interval and the length of the Group of Pictures (GOP) of the video to be processed.

[0043] The frame extraction interval can be determined according to actual requirements.

[0044] A Group of Pictures (GOP) includes a set of consecutive image frames, that is, a set of consecutive pictures. The GOP length is the length of all the image frames in the GOP, such as 3 seconds, 4 seconds, and so on.

[0045] S102, Determine the frame positions to be extracted from the video to be processed according to the frame extraction interval.

[0046] After the frame extraction interval for the video to be processed is determined, the frame positions to be extracted from the video to be processed according to the frame extraction interval can be determined.

[0047] The image frames corresponding to the frame positions to be extracted are the image frames to be extracted from the video to be processed.

[0048] The frame positions to be extracted can be the frame coordinates of the image frames to be extracted. For example, for a video to be processed with a duration of 100 s and N = 10 s, the frame time points to be extracted are 0 s, 10 s, 20 s, and so on.

[0049] S103, Slice the video to be processed based on the GOP length to obtain multiple frame slices to be extracted.

[0050] S104, Extract frames from each of the frame slices to be extracted according to the frame positions to be extracted from the video to be processed.

[0051] The first frame of the GOP is an I frame (key frame), and the decoding of the I frame does not require reference to other image frames, that is, the GOP can be decoded independently.

[0052] In the embodiments of the present disclosure, the frame slices to be extracted obtained by slicing the video to be processed based on the GOP length can be integer multiples of the GOP length. Briefly understood, integer multiples of the GOP can be used as one frame slice to be extracted.

[0053] For each frame slice to be extracted, extract the image frames corresponding to all the frame positions to be extracted in the frame slice to be extracted.

[0054] In an optional embodiment, S104 may include:

[0055] For each thread among multiple threads, according to the frame positions to be extracted from the video to be processed, after extracting frames from one frame slice to be extracted, load another frame slice to be extracted until frames are extracted from all the frame slices to be extracted.

[0056] For each frame extraction shard to be processed, frame extraction for the frame extraction shard to be processed is completed, that is, the image frames corresponding to all positions to be extracted in the frame extraction shard to be processed are extracted.

[0057] It can be understood that multiple threads execute asynchronously in parallel. Among them, the number of multiple threads can be determined according to the performance of the electronic device. For example, according to the performance of the device's Central Processing Unit (CPU), Y threads are established.

[0058] Simply understood, multiple frame extraction shards to be processed constitute a frame extraction pool, and multiple threads execute frame extraction in parallel. Specifically, for each thread, a frame extraction shard to be processed is taken out from the frame extraction pool for frame extraction. After the frame extraction for the frame extraction shard to be processed is completed, another frame extraction shard (a new unprocessed frame extraction shard) is taken out from the frame extraction pool, and frame extraction is continued for the other frame extraction shard. This process is repeated until the frame extraction for all frame extraction shards in the frame extraction pool is completed.

[0059] By using multiple threads to implement parallel asynchronous execution of multiple frame extraction shards to be processed, device resources are fully utilized, performance bottlenecks are reduced, frame extraction time consumption is reduced, and fast video frame extraction is achieved.

[0060] In an optional embodiment, S103 may include:

[0061] In response to the GOP length being equal to the frame extraction interval, the GOP length is used as the shard length; in response to the GOP length being greater than the frame extraction interval, the shard length is determined based on the ratio of the GOP length to the frame extraction interval; in response to the GOP length being less than the frame extraction interval, the product of the GOP length and the frame extraction interval is used as the shard length; the video to be processed is sliced using the shard length to obtain multiple frame extraction shards to be processed.

[0062] In the embodiments of the present disclosure, based on the different size relationships between the GOP length and the frame extraction interval, the shard length is determined respectively, and then the video to be processed is sliced using the shard length to obtain multiple frame extraction shards to be processed.

[0063] Among them, determining the shard length based on the ratio of the GOP length to the frame extraction interval includes:

[0064] Calculate the ratio of the GOP length to the frame extraction interval; round up the ratio to obtain the shard length.

[0065] The ratio of the GOP length to the frame extraction interval may be a decimal. By rounding up, the ratio can be conveniently adjusted to an integer, that is, an integer shard length is obtained, and then the video to be frame-extracted can be conveniently sliced based on the shard length.

[0066] For the case where the GOP length is equal to the frame extraction interval, for example, GOP length = time interval N, then the slice length X = GOP length, and slicing is performed according to the GOP length. As Figure 2 shown, each GOP in the video to be frame-extracted starts with an IDR frame (reference frame). If the GOP length = time interval N, then one GOP is sliced into one frame-extracted slice.

[0067] For the case where the GOP length is greater than the frame extraction interval, for example, GOP length > time interval N, then the slice length X = Ceil(GOP length / N), where Ceil means rounding up GOP length / N. As Figure 3 shown, GOP length = 3 seconds, N = 2 seconds;

[0068] then X = Ceil(3 / 2) * GOP = Ceil(1.5) * 3 = 2 * 3 = 6, where Fragment, that is, the slice length, can also be understood as the length of the frame-extracted slice.

[0069] For the case where the GOP length is less than the frame extraction interval, for example, GOP length < time interval N, then the slice length X = GOP * N seconds. As Figure 4 shown, GOP length = 2 seconds, N = 3 seconds, X = GOP * N = 6 seconds, where Fragment, that is, the slice length, can also be understood as the length of the frame-extracted slice.

[0070] Embodiments of the present disclosure can adaptively slice the video to be processed according to the different size relationships between the GOP length and the frame extraction interval.

[0071] In an optional embodiment, the video frame extraction method provided by the embodiments of the present disclosure may further include:

[0072] Obtain the minimum slice length.

[0073] S103 may include:

[0074] In response to the GOP length being greater than or equal to the minimum slice length, use the GOP length as the slice length; in response to the GOP length being less than the minimum slice length, use the product of the GOP length and a preset value as the slice length, where the preset value is used to ensure that the slice length is greater than the minimum slice length; use the slice length to slice the video to be processed to obtain multiple frame-extracted slices.

[0075] The minimum slice length can be manually configured. For example, the electronic device provides an input interface, and the minimum slice length is input through this input interface. In this way, the electronic device can receive the minimum slice length.

[0076] For example, the minimum slice length is M, and the value range of M is (0, video duration];

[0077] If the GOP length < M, then X = (GOP length * Count) > M; where Count is the above-mentioned preset value, and the preset value is a natural number, such as 2, 3, etc.

[0078] If the GOP length >= M, then X = GOP length.

[0079] The video frame extraction method provided by the embodiments of the present disclosure determines the slice length according to the size relationship between the GOP length and the minimum slice length, and then slices the video to be processed using the slice length to obtain multiple slices to be frame-extracted. According to different size relationships between the GOP length and the minimum slice length, the video to be processed is adaptively sliced.

[0080] In an optional embodiment, for the case where the GOP length is equal to the frame extraction interval, S104 may include:

[0081] For each slice to be frame-extracted, decode the reference frame IDR corresponding to the frame extraction position in the slice to be frame-extracted.

[0082] The GOP starts with an I frame, and the IDR is a type of I frame. In one case, the GOP starts with an IDR frame, and the decoding of the IDR frame does not require reference to other frames, or it can also be understood as not depending on other frames. The GOP length is equal to the frame extraction interval, and the image frame corresponding to the frame extraction position determined according to the frame extraction interval is an IDR frame. In this case, only the IDR frame needs to be decoded and the IDR frame is extracted to extract the image frame corresponding to the frame extraction position, without decoding other frames.

[0083] For example, during the video frame extraction process, when the GOP length is equal to the frame extraction interval, a GOP can be sliced into a slice to be frame-extracted, and only the IDR in each GOP needs to be decoded.

[0084] For the case where the GOP length is greater than or less than the frame extraction interval, S104 may include:

[0085] For each slice to be frame-extracted, if there is a reference frame IDR between the current frame extraction position and the next frame extraction position in the slice to be frame-extracted, after decoding and extracting the image frame corresponding to the current frame extraction position, jump to the IDR and decode the IDR.

[0086] If there is an IDR between the current frame extraction position and the next frame extraction position in the frame extraction slice to be processed, and the decoding of the IDR frame does not require reference to other frames, so it is possible to directly jump to this IDR and no longer decode the image frames between the current frame extraction position and this IDR. There is no frame extraction position between the current frame extraction position and this IDR, and the next frame extraction position can be decoded according to the decoding of this IDR and the image frames between this IDR and the next frame extraction position.

[0087] As Figure 5 shown, when the frame extraction interval is less than the GOP length, a frame extraction slice includes 2 GOPs, GOP1 and GOP2. Both GOP1 and GOP2 start with an IDR frame. For frame extraction of this frame extraction slice, (1) First, start decoding from the first I frame (the IDR in the figure) of the current slice (this frame extraction slice) and extract this frame after decoding; (2) Decode from this IDR to the frame extraction position determined according to the frame extraction interval. The image frame corresponding to the frame extraction position may be a B frame or a P frame, such as decoding to the current B frame or P frame in the figure; (3) There is no image frame to be extracted between the current frame extraction position and the next frame extraction position, that is, there is no point to be extracted in this part, and there is an IDR between the current frame extraction position and the next frame extraction position, so skip this part; (4) Start decoding from the next IDR (the IDR between the current frame extraction position and the next frame extraction position above); (5) Until decoding to the next frame extraction position, the image frame corresponding to the next frame extraction position may be a B frame or a P frame, such as decoding to the current B frame or P frame in the figure, and so on until the frame extraction of this frame extraction slice is completed, that is, all the image frames corresponding to the positions to be extracted in this frame extraction slice are extracted.

[0088] Start extracting the frame of the current IDR from the first frame; if all the frame extraction positions (PTS) in the current GOP have completed frame extraction, then directly jump to the next GOP to start decoding and frame extraction.

[0089] In the embodiments of the present disclosure, during the video frame extraction process, it is not necessary to decode the image frames in the video sequentially, but only to decode some of the image frames in the video to achieve video frame extraction. Compared with the related art that needs to decode all the image frames in the video sequentially, the embodiments of the present disclosure can further reduce the frame extraction time.

[0090] In one implementation, when the GOP length is equal to the frame extraction interval, each thread in multiple threads loads another frame extraction slice after completing the frame extraction of a frame extraction slice according to the frame extraction position of the video to be processed, until the frame extraction of all frame extraction slices is completed. Specifically, for each frame extraction slice, decode the reference frame IDR corresponding to the frame extraction position in the frame extraction slice.

[0091] In another implementable manner, when the GOP length is greater than or less than the frame extraction interval, each thread among multiple threads loads another frame extraction shard after completing frame extraction for a frame extraction shard to be processed according to the frame extraction position of the video to be processed until frame extraction is completed for all the frame extraction shards to be processed. Specifically, for each frame extraction shard, if there is a reference frame IDR between the current frame extraction position and the next frame extraction position in the frame extraction shard, after decoding and extracting the image frame corresponding to the current frame extraction position, it jumps to the IDR and decodes the IDR.

[0092] Corresponding to the video frame extraction method provided in the above embodiment, an embodiment of the present disclosure provides a video frame extraction device, as Figure 6 shown, which may include:

[0093] A first acquisition module 601, configured to acquire the frame extraction interval and the group of pictures (GOP) length of the video to be processed;

[0094] A determination module 602, configured to determine the frame extraction positions of the video to be processed according to the frame extraction interval;

[0095] A slicing module 603, configured to slice the video to be processed based on the GOP length to obtain a plurality of frame extraction shards;

[0096] A frame extraction module 604, configured to perform frame extraction on each of the frame extraction shards according to the frame extraction positions of the video to be processed.

[0097] Optionally, the frame extraction module 604 is specifically configured to, for the case where the GOP length is equal to the frame extraction interval, decode the reference frame IDR corresponding to the frame extraction position in each frame extraction shard; for the case where the GOP length is greater than or less than the frame extraction interval, for each frame extraction shard, if there is a reference frame IDR between the current frame extraction position and the next frame extraction position in the frame extraction shard, after decoding and extracting the image frame corresponding to the current frame extraction position, it jumps to the IDR and decodes the IDR.

[0098] Optionally, the slicing module 603 is specifically configured to, in response to the GOP length being equal to the frame extraction interval, use the GOP length as the shard interval; in response to the GOP length being greater than the frame extraction interval, determine the shard length based on the ratio of the GOP length to the frame extraction interval; in response to the GOP length being less than the frame extraction interval, use the product of the GOP length and the frame extraction interval as the shard length; and slice the video to be processed using the shard length to obtain a plurality of frame extraction shards.

[0099] Optionally, the slicing module 603 is specifically configured to calculate the ratio of the GOP length to the frame extraction interval; round up the ratio to obtain the shard length.

[0100] Optionally, as Figure 7 shown, the apparatus further includes:

[0101] A second acquisition module 701, configured to acquire the minimum slice length;

[0102] A slicing module 603, specifically configured to: in response to the GOP length being greater than or equal to the minimum slice length, use the GOP length as the slice length; in response to the GOP length being less than the minimum slice length, use the product of the GOP length and a preset value as the slice length, where the preset value is used to ensure that the slice length is greater than the minimum slice length; and slice the video to be processed using the slice length to obtain a plurality of slices to be frame-extracted.

[0103] Optionally, a frame extraction module 604, specifically configured to, for each thread among a plurality of threads, after frame extraction is completed for a slice to be frame-extracted of the video to be processed, load another slice to be frame-extracted until frame extraction is completed for all slices to be frame-extracted.

[0104] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0105] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.

[0106] Figure 8 FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0107] As Figure 8 shown, the device 800 includes a computing unit 801, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0108] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as a keyboard, mouse, etc.; output unit 807, such as various types of displays, speakers, etc.; storage unit 808, such as a disk, optical disc, etc.; and communication unit 809, such as a network card, modem, wireless communication transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0109] Computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 801 executes the various methods and processes described above, such as the video frame extraction method. For example, in some embodiments, the video frame extraction method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by computing unit 801, one or more steps of the video frame extraction method described above can be executed. Alternatively, in other embodiments, computing unit 801 can be configured to execute the video frame extraction method in any other suitable manner (e.g., by means of firmware).

[0110] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0111] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0112] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0113] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0114] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0115] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating blockchain.

[0116] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0117] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A video frame extraction method, comprising: obtaining a frame extraction interval and the length of a Group of Pictures (GOP) of a video to be processed; determining positions of frames to be extracted from the video to be processed according to the frame extraction interval; slicing the video to be processed based on the GOP length to obtain a plurality of frame extraction slices, including: in response to the GOP length being equal to the frame extraction interval, using the GOP length as the slice length; in response to the GOP length being greater than the frame extraction interval, determining the slice length based on the ratio of the GOP length to the frame extraction interval; in response to the GOP length being less than the frame extraction interval, using the product of the GOP length and the frame extraction interval as the slice length; and slicing the video to be processed using the slice length to obtain a plurality of frame extraction slices; extracting frames from each of the frame extraction slices respectively according to the positions of frames to be extracted from the video to be processed.

2. The method according to claim 1, wherein for the case where the GOP length is equal to the frame extraction interval, the step of extracting frames from each of the frame extraction slices respectively according to the positions of frames to be extracted from the video to be processed includes: for each of the frame extraction slices, decoding the reference frame IDR corresponding to the position of the frame to be extracted in the frame extraction slice; for the case where the GOP length is greater than or less than the frame extraction interval, the step of extracting frames from each of the frame extraction slices respectively according to the positions of frames to be extracted from the video to be processed includes: for each of the frame extraction slices, if there is a reference frame IDR between the current position of the frame to be extracted and the next position of the frame to be extracted in the frame extraction slice, after decoding and extracting the image frame corresponding to the current position of the frame to be extracted, jumping to the IDR and decoding the IDR.

3. The method according to claim 1, wherein the step of determining the slice length based on the ratio of the GOP length to the frame extraction interval includes: calculating the ratio of the GOP length to the frame extraction interval; rounding up the ratio to obtain the slice length.

4. The method according to claim 1, the method further comprises: obtaining a minimum slice length; the step of slicing the video to be processed based on the GOP length to obtain a plurality of frame extraction slices includes: in response to the GOP length being greater than or equal to the minimum slice length, using the GOP length as the slice length; in response to the GOP length being less than the minimum slice length, using the product of the GOP length and a preset value as the slice length, the preset value being used to ensure that the slice length is greater than the minimum slice length; slicing the video to be processed using the slice length to obtain a plurality of frame extraction slices.

5. The method according to claim 1, wherein the step of extracting frames from each of the frame extraction slices respectively according to the positions of frames to be extracted from the video to be processed includes: for each thread among a plurality of threads, according to the positions of frames to be extracted from the video to be processed, after extracting frames from one frame extraction slice, loading another frame extraction slice until frames are extracted from all the frame extraction slices.

6. A video frame extraction device, comprising: A first acquisition module, configured to acquire a frame extraction interval and a Group of Pictures (GOP) length of a video to be processed; A determination module, configured to determine a frame extraction position to be extracted from the video to be processed according to the frame extraction interval; A slicing module, configured to slice the video to be processed based on the GOP length to obtain a plurality of frame extraction sub - slices, including: in response to the GOP length being equal to the frame extraction interval, using the GOP length as the slice length; in response to the GOP length being greater than the frame extraction interval, determining the slice length based on the ratio of the GOP length to the frame extraction interval; in response to the GOP length being less than the frame extraction interval, using the product of the GOP length and the frame extraction interval as the slice length; and slicing the video to be processed using the slice length to obtain a plurality of frame extraction sub - slices; A frame extraction module, configured to perform frame extraction on each of the frame extraction sub - slices according to the frame extraction position to be extracted from the video to be processed.

7. The apparatus according to claim 6, wherein, the frame extraction module is specifically configured to, for the case where the GOP length is equal to the frame extraction interval, for each frame extraction sub - slice, decode the reference frame IDR corresponding to the frame extraction position in the frame extraction sub - slice; for the case where the GOP length is greater than or less than the frame extraction interval, for each frame extraction sub - slice, if there is a reference frame IDR between the current frame extraction position and the next frame extraction position in the frame extraction sub - slice, after decoding and extracting the image frame corresponding to the current frame extraction position, jump to the IDR and decode the IDR.

8. The apparatus according to claim 6, wherein, the slicing module is specifically configured to calculate the ratio of the GOP length to the frame extraction interval; round up the ratio to obtain the slice length.

9. The apparatus according to claim 6, the apparatus further comprises: A second acquisition module, configured to acquire a minimum slice length; the slicing module is specifically configured to, in response to the GOP length being greater than or equal to the minimum slice length, use the GOP length as the slice length; in response to the GOP length being less than the minimum slice length, use the product of the GOP length and a preset value as the slice length, where the preset value is used to ensure that the slice length is greater than the minimum slice length; and slice the video to be processed using the slice length to obtain a plurality of frame extraction sub - slices.

10. The apparatus according to claim 6, wherein, the frame extraction module is specifically configured to, through each thread in a plurality of threads, according to the frame extraction position to be extracted from the video to be processed, after completing frame extraction on a frame extraction sub - slice, load another frame extraction sub - slice until frame extraction is completed on all frame extraction sub - slices.

11. An electronic device, comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 - 5.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein, the computer instructions are for causing the computer to execute the method according to any one of claims 1-5.

13. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Method and device for generating image interchange format file

    CN106791918A

  • Video frame extraction method and device, electronic equipment and readable storage medium

    CN112565886A