Video frame extraction method and device, electronic equipment and storage medium

By dynamically adjusting the video frame extraction scheme based on the length of the storyboard segments and the frame extraction threshold, the information loss and redundancy problems caused by different video lengths are solved, and the efficiency and cost of video frame extraction are optimized.

CN120769110APending Publication Date: 2025-10-10BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511040268.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

In the existing technology, when videos have different lengths, the same frame extraction method causes the longer video to lose valid information and the shorter video to have information redundancy, which increases the computing cost.

Method used

By determining the storyboard segments of the video and their segment lengths, combined with the preset storyboard frame extraction threshold and the allocated frame extraction number, the frame extraction number is dynamically adjusted to ensure the minimum frame extraction number for each storyboard segment, and the frame extraction frequency is dynamically adjusted according to the segment length.

Benefits of technology

It achieves the dynamic determination of the number of video frames, which not only ensures that valid frames are extracted from videos with longer durations, but also reduces the information redundancy of videos with shorter durations and optimizes the computing cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120769110A_ABST
    Figure CN120769110A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a video frame extraction method and device, electronic equipment and a storage medium. The method comprises the following steps: determining a sub-lens segment included in a to-be-processed video and a segment duration of the sub-lens segment; the sum of a preset sub-lens frame extraction threshold value and the distribution frame extraction number determined based on the segment duration is determined as the number of frames to be extracted corresponding to the sub-lens segment, and the distribution frame extraction number is in positive correlation with the segment duration; based on the number of frames to be extracted, carrying out frame extraction on the lens-splitting lens segments to obtain a lens-splitting frame extraction sequence; and determining a video frame extraction sequence corresponding to the to-be-processed video based on the sub-mirror frame extraction sequence. By adopting the scheme of the invention, the frame extraction number is dynamically adjusted according to the duration of the sub-lens segment on the basis of ensuring the minimum frame extraction number of the sub-lens segment, and the dynamic determination of the video frame extraction number is realized, so that effective video frames can be extracted from a video with a relatively long duration, and the video frame extraction efficiency is improved. And the reduction of information redundancy caused by excessive frame extraction of the video with short duration is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of video processing technology, and in particular to a video frame extraction method, device, electronic device, and storage medium. Background Art

[0002] When performing video understanding tasks in vertical business fields, it is usually necessary to extract frames from the video and input the extracted video frame sequence into the model for training or inference.

[0003] Since the videos to be understood are not of fixed length and can range from 4s to 30s, the current frame extraction method used in related technologies extracts frames from videos of different lengths according to the same number of image frames. That is, the number of frames extracted for long and short videos is the same. This results in a lower frame extraction frequency for longer videos, which may result in the loss of valid video information, while shorter videos have a higher frame extraction frequency, in which a lot of information is redundant, increasing computational costs.

[0004] Therefore, for a known video, how to extract more effective video frames for analysis becomes an urgent problem to be solved. Summary of the Invention

[0005] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a video frame extraction method, device, electronic device and storage medium.

[0006] The present disclosure provides a video frame extraction method, the method comprising:

[0007] Determining the storyboard segments included in the video to be processed and the segment lengths of the storyboard segments;

[0008] Determining the number of frames to be extracted corresponding to the storyboard segment by summing a preset storyboard frame extraction threshold and a number of allocated frames determined based on the segment duration, wherein the number of allocated frames is positively correlated with the segment duration;

[0009] Extract frames from the storyboard segment based on the number of frames to be extracted to obtain a storyboard frame extraction sequence;

[0010] A video frame extraction sequence corresponding to the video to be processed is determined based on the storyboard frame extraction sequence.

[0011] The present disclosure also provides a video frame extraction device, the device comprising:

[0012] A first determining module is used to determine the storyboard segments included in the video to be processed and the segment lengths of the storyboard segments;

[0013] determine, as the number of frames to be extracted corresponding to the split shot segment, a sum of a preset split shot extraction threshold and an assigned number of extracted frames determined based on the segment duration, wherein the assigned number of extracted frames is positively correlated with the segment duration;

[0014] extract frames from the split shot segment based on the number of frames to be extracted, to obtain a split shot extracted frame sequence;

[0015] determine, as the video extracted frame sequence corresponding to the video to be processed, a split shot extracted frame sequence based on the split shot extracted frame sequence.

[0016] The embodiments of the present disclosure further provide an electronic device, which comprises a processor, a memory for storing executable instructions of the processor, and the processor is configured to read the executable instructions from the memory and execute the instructions to implement the video frame extraction method provided by the embodiments of the present disclosure.

[0017] The embodiments of the present disclosure further provide a computer readable storage medium, which stores a computer program for executing the video frame extraction method provided by the embodiments of the present disclosure.

[0018] The embodiments of the present disclosure further provide a computer program product for executing the video frame extraction method provided by the embodiments of the present disclosure.

[0019] The technical solutions provided by the embodiments of the present disclosure have the following advantages compared with the prior art.

[0020] The video frame extraction method provided by the embodiments of the present disclosure comprises the following steps: determining a split shot segment included in a video to be processed and a segment duration of the split shot segment; determining, as a number of frames to be extracted corresponding to the split shot segment, a sum of a preset split shot extraction threshold and an assigned number of extracted frames determined based on the segment duration, wherein the assigned number of extracted frames is positively correlated with the segment duration; extracting frames from the split shot segment based on the number of frames to be extracted, to obtain a split shot extracted frame sequence; and determining, as a video extracted frame sequence corresponding to the video to be processed, a split shot extracted frame sequence based on the split shot extracted frame sequence. By adopting the solutions of the present disclosure, the video to be processed is decomposed into a split shot segment and the segment duration of the split shot segment is obtained, the assigned number of extracted frames is determined based on the segment duration, the sum of the assigned number of extracted frames and the split shot extraction threshold is determined as the number of frames to be extracted corresponding to the split shot segment, and then the split shot segment is extracted according to the number of frames to be extracted, which realizes dynamic adjustment of the number of extracted frames according to the duration of the split shot segment on the basis of guaranteeing the minimum number of extracted frames of the split shot segment, realizes dynamic determination of the number of extracted frames of the video, helps to guarantee that a video with a longer duration can extract effective video frames, and helps to reduce information redundancy caused by excessive frame extraction of a video with a shorter duration. BRIEF DESCRIPTION OF DRAWINGS

[0021] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.

[0022] Figure 1 A flowchart of a video frame extraction method provided by an exemplary embodiment of the present disclosure;

[0023] Figure 2 A flowchart of a video frame extraction method provided by another exemplary embodiment of the present disclosure;

[0024] Figure 3 A flowchart of a video frame extraction method provided by another exemplary embodiment of the present disclosure;

[0025] Figure 4 A schematic structural diagram of a video frame extraction device provided in an embodiment of the present disclosure;

[0026] Figure 5 A schematic structural diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0027] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0028] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0029] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.

[0030] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0031] It should be noted that the modification of "one" and "multiple" mentioned in the present disclosure is illustrative but not restrictive, and those skilled in the art should understand that "one or more" should be understood unless otherwise explicitly indicated in the context.

[0032] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not used to limit the scope of the messages or information.

[0033] When performing a vertical field video content understanding task through a multi-modal large model, the machine performance and the inference performance of the model need to be considered comprehensively. In order to reduce the computational complexity of the model and reduce the consumption of machine resources, when performing a video understanding task, the model is usually input with a picture frame sequence extracted from a video instead of a complete video. Therefore, for a known video, how to extract more effective video frames for analysis becomes more important.

[0034] When performing a video understanding task in a vertical business field, the video length that needs to be understood can be a variable length from 4s to 30s. However, in the related art, for videos of different lengths, the same number of picture frames is usually extracted, which leads to a lower frame extraction frequency for a longer video, which may lose effective video information, and a higher frame extraction frequency for a shorter video, in which a lot of information is redundant, increasing the computational cost.

[0035] To solve the above problems, the present disclosure provides a video frame extraction scheme. For a video to be understood, the number of allocated frames is determined according to the segment length of the included shot segment, and the sum of the allocated frame number and a preset shot frame extraction threshold is determined as the frame number to be extracted for the shot segment. Then, the picture frames in the shot segment are selected according to the frame number to be extracted, so as to realize the frame extraction of the video. In this way, it is ensured that effective picture frames can be extracted for a longer shot, and information redundancy caused by excessive frame extraction for a shorter shot can be reduced.

[0036] The video frame extraction method, device, electronic equipment and storage medium of the present disclosure will be explained in detail below in combination with specific embodiments.

[0037] Figure 1 A flowchart of a video frame extraction method provided for an exemplary embodiment of the present disclosure is shown. The method can be executed by a video frame extraction device provided by an embodiment of the present disclosure, which can be implemented by software and / or hardware and generally integrated in an electronic device.

[0038] As shown in FIG. 1, the video frame extraction method includes: Figure 1

[0039] ​Step 101: Determine the storyboard segments included in the video to be processed and the segment lengths of the storyboard segments.

[0040] The video to be processed can be of any length and can be a video that requires model understanding during model training or model inference. A video typically contains at least one storyboard segment. In the disclosed embodiment, the storyboard segment and the duration of each storyboard segment contained in the video to be processed can be obtained.

[0041] As a kind of example, for some videos, it can store the transition point data corresponding to this video simultaneously when storing, transition point refers to the lens transition point, and transition point data are accurate to frame level, adopt the data with milliseconds to carry out the representation of time point position, i.e. the starting time and the ending time of each lens (storyboard).For each video, transition point data can be calculated by transition point algorithm, and write in the database.Therefore for the pending video having stored transition point data, pending video comprises transition point data, i.e. the starting time and the ending time of each storyboard, based on transition point data, can determine the storyboard segment of each storyboard, and can obtain the fragment duration of each storyboard segment.

[0042] As an example, the presence of split screens can be determined by analyzing the coherence between adjacent picture frames in the video to be processed. When the content between two adjacent picture frames changes, it is determined that a split screen has occurred. The first frame of the two picture frames belongs to one split screen, and the second frame belongs to another split screen. Therefore, different split screen segments can be determined based on the coherence between the picture frames in the video to be processed, and the segment length of each split screen segment can be obtained.

[0043] Step 102: The sum of the preset storyboard frame extraction threshold and the allocated frame extraction number determined based on the segment duration is determined as the number of frames to be extracted corresponding to the storyboard segment, wherein the allocated frame extraction number is positively correlated with the segment duration.

[0044] The frame extraction threshold represents the minimum number of frames required to be extracted for a storyboard segment. That is, the number of frames extracted for a storyboard segment must not be less than the frame extraction threshold. The specific value of the frame extraction threshold can be set according to actual needs, for example, it can be set to 3. By setting the frame extraction threshold, the number of frames extracted for a storyboard segment can be guaranteed.

[0045] In this embodiment, for each storyboard segment, the number of allocated frames corresponding to the storyboard segment is determined according to the segment duration of the storyboard segment, and the number of allocated frames is positively correlated with the segment duration.

[0046] As an example, different correspondences between durations and number of frames can be pre-set, where longer durations correspond to more number of frames. When determining the number of frames allocated for a storyboard segment based on its duration, the number of frames allocated for the storyboard segment can be determined by querying this correspondence.

[0047] In this embodiment, after the allocated number of frames to be extracted for a storyboard segment is determined, the allocated number of frames to be extracted is added to a preset storyboard frame extraction threshold, and the resulting sum is the number of frames to be extracted corresponding to the storyboard segment.

[0048] Step 103: extract frames from the storyboard segments based on the number of frames to be extracted to obtain a storyboard frame extraction sequence.

[0049] In this embodiment, after determining the number of frames to be extracted for each storyboard segment, image frames can be extracted for each storyboard segment based on the number of frames to be extracted, thereby obtaining a storyboard frame extraction sequence. For example, if the number of frames to be extracted for storyboard segment 1 is 3 and the number of frames to be extracted for storyboard segment 2 is 7, then 3 image frames are extracted from storyboard segment 1 and 7 image frames are extracted from storyboard segment 2.

[0050] As an example, when extracting picture frames from a storyboard segment based on the number of frames to be extracted in the storyboard segment, picture frames of the number of frames to be extracted can be evenly extracted from the storyboard segment to form a storyboard frame extraction sequence of the storyboard segment.

[0051] As an example, when extracting picture frames from a storyboard segment based on the number of frames to be extracted, the first picture frame and the last picture frame of the storyboard segment can be first obtained, and based on the number of frames to be extracted, the remaining number of frames to be extracted can be determined, that is, the remaining number of frames to be extracted = the number of frames to be extracted - 2 (because 2 picture frames have been extracted); then, based on the remaining number of frames to be extracted, the target picture frames for the remaining number of frames to be extracted can be extracted from the storyboard segment, wherein, when extracting frames, the target picture frames can be extracted in an evenly spaced frame extraction manner, for example, the frame extraction frequency (frame extraction frequency = total number of picture frames / remaining number of frames to be extracted) can be determined based on the total number of picture frames contained in the remaining storyboard segment after excluding the first picture frame and the last picture frame, and the remaining number of frames to be extracted, and then the target picture frames can be extracted from the remaining storyboard segment according to the frame extraction frequency; finally, based on the first picture frame, the target picture frame and the last picture frame, the storyboard frame extraction sequence corresponding to the storyboard segment is determined, wherein the picture frames in the storyboard frame extraction sequence retain the temporal relationship in the storyboard segment.

[0052] Step 104: Determine a video frame extraction sequence corresponding to the video to be processed based on the storyboard frame extraction sequence.

[0053] In this embodiment, after extracting the storyboard frame sequence of each storyboard segment, a video frame sequence can be constructed based on all the storyboard frame sequences, wherein each storyboard frame sequence in the video frame sequence retains the temporal relationship of the corresponding storyboard segment in the video to be processed. That is to say, assuming that the storyboard segments contained in the video to be processed are storyboard segment 1, storyboard segment 2 and storyboard segment 3 in sequence, and the extracted storyboard frame sequences are storyboard frame sequence 1, storyboard frame sequence 2 and storyboard frame sequence 3 in sequence, then the storyboard frame sequences of the video frame sequence are storyboard frame sequence 1, storyboard frame sequence 2 and storyboard frame sequence 3 in sequence.

[0054] The video frame extraction method of the embodiment of the present disclosure determines the storyboard segments included in the video to be processed and the segment lengths of the storyboard segments; determines the sum of a preset storyboard frame extraction threshold and the allocated frame extraction number determined based on the segment length as the number of frames to be extracted corresponding to the storyboard segments, wherein the allocated frame extraction number is positively correlated with the segment length; extracts the storyboard segments based on the number of frames to be extracted to obtain a storyboard frame extraction sequence; and determines a video frame extraction sequence corresponding to the video to be processed based on the storyboard frame extraction sequence. By adopting the scheme disclosed in the present invention, the video to be processed is decomposed into storyboard segments and the segment duration of the storyboard segments is obtained, the number of allocated frame extractions is determined based on the segment duration, the sum of the allocated frame extraction number and the storyboard frame extraction threshold is determined as the number of frames to be extracted corresponding to the storyboard segment, and then the storyboard segment is extracted according to the number of frames to be extracted. This achieves dynamic adjustment of the number of frame extractions according to the duration of the storyboard segment while ensuring the minimum number of frame extractions of the storyboard segment, and realizes dynamic determination of the number of video frame extractions, which not only helps to ensure that valid video frames can be extracted from videos with longer durations, but also helps to reduce information redundancy caused by excessive frame extraction from videos with shorter durations.

[0055] In order to ensure the balance of the number of frame extractions, in an optional embodiment of the present disclosure, as Figure 2 As shown, based on the above embodiment, the step of determining the number of allocated frames based on the segment duration in step 102 may include the following sub-steps:

[0056] Step 201: Obtain the number of shots in a storyboard segment and the duration of the video to be processed. In this embodiment, for a known video, the duration of the video can be determined, and the number of shots contained in the video can also be determined. For example, after determining the shot segments in the video to be processed in the aforementioned embodiment, the shot data of the shot segments can be statistically obtained.

[0057] As an example, when determining the video duration of a video to be processed, the video duration can be determined based on the total number of frames and frame rate of the video to be processed: video duration = total number of frames / frame rate. Parameters such as the total number of frames and frame rate of a video are typically recorded in the video's attribute information. Therefore, the total number of frames and frame rate of the video to be processed can be obtained from the attribute information of the video to calculate the video duration of the video to be processed.

[0058] As another example, assuming that the attribute information of the video to be processed records a parameter of video duration, the video duration can be directly obtained from the attribute information of the video to be processed.

[0059] Step 202: Determine the total number of frames to be extracted based on the number of shots and the duration of the video, wherein the total number of frames to be extracted is related to the larger value of a first number of frames to be extracted determined according to the duration of the video and a second number of frames to be extracted determined according to the number of shots.

[0060] In this embodiment, after determining the number of shots and the duration of the video to be processed, the total number of frames to be extracted from the video to be processed can be determined based on the number of shots and the duration of the video. The total number of frames to be extracted is positively correlated with the number of shots and the duration of the video.

[0061] As an example, a first mapping relationship between the number of storyboards and the number of frames to be extracted, and a second mapping relationship between the video duration and the number of frames to be extracted can be established in advance, wherein the number of storyboards is positively correlated with the number of frames to be extracted, and the video duration is positively correlated with the number of frames to be extracted. After determining the number of storyboards and the video duration of the video to be processed, the second mapping relationship is queried according to the video duration to determine the corresponding first number of frames to be extracted, and the first mapping relationship is queried according to the number of storyboards to determine the corresponding second number of frames to be extracted, and then the total number of frames to be extracted of the video to be processed is determined according to the larger value of the first number of frames to be extracted and the second number of frames to be extracted. For example, the larger value of the first number of frames to be extracted and the second number of frames to be extracted can be used as the total number of frames to be extracted. It should be noted that if the determined total number of frames to be extracted cannot guarantee the minimum number of frames extracted for the storyboard segment, that is, the total number of frames to be extracted is less than the product of the number of storyboards contained in the video to be processed and the storyboard frame extraction threshold, then the product of the number of storyboards and the storyboard frame extraction threshold is determined as the total number of frames to be extracted. In this way, it can be ensured that the number of picture frames extracted from each storyboard segment of the video to be processed is not less than the storyboard frame extraction threshold, so that the number of frames extracted for the storyboard segment meets the minimum frame extraction requirement.

[0062] Step 203: The difference obtained by subtracting the product of the number of storyboards and the storyboard frame extraction threshold from the total number of frames to be extracted is determined as the remaining number of frames to be allocated.

[0063] In this embodiment, the number of remaining frames to be allocated = the total number of frames to be extracted - the number of shots * the shot extraction threshold.

[0064] In step 204, the remaining frame numbers to be allocated are allocated to the subshots according to the time lengths of the subshots to obtain the allocated frame numbers of the subshots.

[0065] In this embodiment, for the total frame numbers to be extracted of the video to be processed, the remaining frame numbers to be allocated are allocated to the subshots according to the time lengths of the subshots on the premise that each subshot extracts at least the threshold of subshot frame numbers.

[0066] As an example, the allocated weights of the subshots can be determined according to the ratios of the time lengths of the subshots to the video length of the video to be processed, that is, the allocated weight is equal to the time length of the subshot divided by the video length of the video to be processed (the result can be kept to two decimal places), and then the remaining frame numbers to be allocated are allocated to the subshots according to the allocated weights of the subshots to obtain the allocated frame numbers of the subshots, for example, the product of the allocated weight and the remaining frame numbers to be allocated after rounding off can be obtained as the allocated frame number of the subshot. After the allocated frame numbers of the subshots are determined in the above manner, if the sum of the allocated frame numbers of all the subshots is less than the remaining frame numbers to be allocated, the difference between the two is calculated, and the allocated frame number of the subshot with the longest time length is increased by the difference, so that the subshot with the longer time length extracts more frame numbers. If the sum of the allocated frame numbers of all the subshots is greater than the remaining frame numbers to be allocated, the difference between the two is calculated, and subshots with the same number of the difference are selected in the order from short to long according to the time lengths of the subshots, and 1 is subtracted from the allocated frame numbers of the selected subshots respectively to obtain the final values of the allocated frame numbers of the subshots. For example, assuming that the sum of the allocated frame numbers of all the subshots is greater than the remaining frame numbers to be allocated by 2, the two subshots with the shortest time lengths are selected, and the allocated frame numbers of the two subshots are reduced by 1 respectively.

[0067] As an example, for the obtained remaining number of frames to be allocated, it can be determined whether the remaining number of frames to be allocated is 0. If it is 0, it is determined that the number of allocated frames corresponding to each storyboard segment is the storyboard frame withdrawal threshold; if the remaining number of frames to be allocated is not 0, then for each storyboard segment, the ratio of the segment length of the storyboard segment to the current number of allocated frames corresponding to itself is further determined to obtain the frame withdrawal interval, wherein the current number of allocated frames of a storyboard segment is the sum of the storyboard frame withdrawal threshold and the number of allocated frames, and the initial value of the number of allocated frames is 0; then, the number of allocated frames corresponding to the target storyboard segment with the largest frame withdrawal interval is accumulated by a preset value, so that the current number of allocated frames of the target storyboard segment is also accumulated by the preset value accordingly, and the remaining number of frames to be allocated is reduced by a preset value to update the remaining number of frames to be allocated, wherein the preset value can be set according to actual needs, for example, the preset value can be set to 1; thereby, the number of frames can be increased for storyboard segments with longer segment lengths. After determining the new number of remaining frames to be allocated, the process returns to the step of determining whether the remaining number of frames to be allocated is 0. If it is 0, the process ends and the number of allocated frames corresponding to the storyboard segment is obtained. If it is still not 0, the process continues with the step of determining the ratio of the segment duration to the currently allocated number of frames corresponding to the segment to obtain the frame extraction interval and subsequent steps, beginning a new round of allocation. It is understood that at this point, the currently allocated number of frames for some storyboard segments is greater than the storyboard frame extraction threshold. The above process is repeated until the remaining number of frames to be allocated is 0, at which point the number of allocated frames corresponding to the storyboard segment is obtained. Through the above process, the remaining number of frames to be allocated is recursively added to the eligible storyboard segments one by one, allowing longer storyboard segments to obtain a larger number of frames to be extracted, thereby extracting more frames, ensuring that valid frames are extracted for long storyboard segments, and reducing the number of frames extracted for short storyboard segments, thereby reducing the probability of information redundancy.

[0068] The video frame extraction method of the embodiment of the present invention obtains the number of storyboards in the storyboard segment and the video duration of the video to be processed, and determines the total number of frames to be extracted based on the number of storyboards and the video duration. The total number of frames to be extracted is related to the larger value of the first frame extraction number determined according to the video duration and the second frame extraction number determined according to the number of storyboards. Then, the difference obtained by subtracting the product of the number of storyboards and the storyboard frame extraction threshold from the total number of frames to be extracted is determined as the remaining number of frames to be allocated, and then the remaining number of frames to be allocated are allocated to the storyboard segments according to the segment duration to obtain the allocated number of frames corresponding to the storyboard segments. In this way, under the premise of ensuring the minimum number of frames for each storyboard segment, the remaining number of frames to be allocated are allocated to each storyboard segment according to the segment duration of the storyboard segment, so that the number of frames extracted of each storyboard segment matches the segment duration of the storyboard segment, thereby ensuring the balance of the number of frames extracted between storyboards.

[0069] In an optional embodiment of the present disclosure, Figure 3As shown, based on the above embodiment, step 202 may include the following sub-steps:

[0070] Step 301: Obtain a preset lower limit value of video duration, an upper limit value of video duration, a minimum number of extracted frames, and a maximum number of extracted frames.

[0071] Among them, the lower limit of video length, the upper limit of video length, the minimum number of frames extracted and the maximum number of frames extracted can all be pre-set according to actual needs. The lower limit of video length refers to the minimum length of the video that can be processed by this solution (for example, the lower limit of video length is set to 4 seconds). Correspondingly, the upper limit of video length refers to the maximum length of the video that can be processed (for example, the upper limit of video length is set to 30 seconds); the minimum number of frames extracted refers to the minimum number of picture frames for extracting a video, for example, the minimum number of frames extracted can be set to 16, that is, a video extracts at least 16 picture frames; the maximum number of frames extracted refers to the maximum number of picture frames for extracting a video, for example, the maximum number of frames extracted can be set to 32, that is, a video extracts at most 32 picture frames. The minimum number of frames extracted and the maximum number of frames extracted can be set according to the processing capability of the model used to process the extracted video frame sequence, and the present disclosure does not limit their specific values.

[0072] Step 302: Determine a ratio of a first difference between the video duration and the lower limit of the video duration divided by a second difference between the upper limit of the video duration and the lower limit of the video duration.

[0073] Step 303: Determine an increment of the number of frame extractions based on a product of a third difference between the maximum number of frame extractions and the minimum number of frame extractions and the ratio.

[0074] That is to say, the increment of the number of extracted frames = (video duration - video duration lower limit) / (video duration upper limit - video duration lower limit) * (maximum number of extracted frames - minimum number of extracted frames).

[0075] Step 304: Determine a first number of extracted frames based on the sum of the incremental number of extracted frames and the minimum number of extracted frames.

[0076] For example, the sum of the incremental number of extracted frames and the minimum number of extracted frames may be determined as the first number of extracted frames.

[0077] To ensure that the determined number of first extracted frames is an integer, the increment of the number of extracted frames can be obtained by rounding up the product of the third difference and the ratio. Assuming that the video duration is recorded as use_time, the minimum video duration is recorded as min_time, the maximum video duration is recorded as max_time, the maximum number of extracted frames is recorded as max_num, and the minimum number of extracted frames is recorded as min_num, the first number of extracted frames (recorded as img_num) can be calculated using the following formula:

[0078]

[0079] Step 305: Determine the total number of frames to be extracted based on the larger value of the first number of extracted frames and the second number of extracted frames obtained by multiplying the number of storyboards by the preset value of the number of extracted frames.

[0080] Among them, the preset value of the number of frame extractions is different from the storyboard frame extraction threshold described in the embodiment of the present disclosure. The preset value of the number of frame extractions is slightly larger than the storyboard frame extraction threshold. For example, the storyboard frame extraction threshold is set to 3, and the preset value of the number of frame extractions is set to 4.

[0081] In this embodiment, the second number of frames extracted = the number of storyboards * the preset value of the number of frames extracted. After obtaining the second number of frames extracted, the second number of frames extracted can be compared with the first number of frames extracted to determine the larger value of the two, and then the total number of frames to be extracted can be determined based on the larger value.

[0082] As an example, a larger value between the second number of extracted frames and the first number of extracted frames may be determined as the total number of frames to be extracted.

[0083] As an example, the larger value of the second number of extracted frames and the first number of extracted frames can be used as the temporary number of extracted frames, and then the total number of frames to be extracted can be determined based on the relationship between the temporary number of extracted frames and the maximum number of extracted frames. If the temporary number of extracted frames is not greater than the maximum number of extracted frames, the temporary number of extracted frames is determined as the total number of frames to be extracted; if the temporary number of extracted frames is greater than the maximum number of extracted frames, the maximum number of extracted frames is determined as the total number of frames to be extracted.

[0084] In the embodiment of the present disclosure, by setting a preset value for the number of frame extractions that is greater than the storyboard frame extraction threshold, the product of the preset value for the number of frame extractions and the number of storyboards is used as the second number of frame extractions, and the larger value of the second number of frame extractions and the first number of frame extractions is selected as the temporary number of frame extractions to determine the total number of frames to be extracted. This can appropriately increase the number of frame extractions of the video to be processed, thereby ensuring that some longer storyboard segments can add some frame extraction pictures, thereby increasing the number of extracted picture frames.

[0085] The video frame extraction method of the embodiment of the present invention obtains a preset lower limit value of video length, an upper limit value of video length, a minimum number of extracted frames and a maximum number of extracted frames, determines the ratio of the first difference between the video length and the lower limit value of video length divided by the second difference between the upper limit value of video length and the lower limit value of video length, determines the frame extraction increment based on the product of the third difference between the maximum number of extracted frames and the minimum number of extracted frames multiplied by the ratio, and determines the first number of extracted frames based on the sum of the frame extraction increment and the minimum number of extracted frames, and then determines the total number of frames to be extracted based on the larger value of the second number of extracted frames obtained by multiplying the first number of extracted frames and the number of storyboards by the preset value of the frame extraction. As a result, the determined total number of frames to be extracted matches the video length and the number of storyboards, and will not be lower than the minimum number of extracted frames nor exceed the maximum number of extracted frames, which meets the processing capability of the model for processing video understanding tasks.

[0086] When a video to be processed contains too many storyboards, it may not be possible to ensure that the number of frames to be extracted allocated to each storyboard segment is not less than the storyboard frame extraction threshold, that is, it cannot be guaranteed that each storyboard segment can extract image frames that meet the storyboard frame extraction threshold. In order to solve this problem, the video to be processed can be segmented, and each sub-video after segmentation can be extracted according to the video frame extraction scheme described in the aforementioned embodiment. Thus, in an optional embodiment of the present disclosure, after determining the number of shots contained in the video to be processed, it can be first determined whether the number of shots is greater than the shot number threshold, wherein the shot number threshold can be set according to actual needs; if the number of shots is not greater than the shot number threshold, the video frame extraction method described in the aforementioned embodiment is directly used to extract the frames to obtain a video frame extraction sequence; if the number of shots is greater than the shot number threshold, the video to be processed is divided into at least two sub-videos, each of which contains at least two sub-videos. The number of shots contained in each sub-video is not greater than the shot number threshold, and then the at least two sub-videos are subjected to frame extraction processing respectively. That is, for each sub-video, the number of frames to be extracted corresponding to the shot segments in the sub-video is determined according to the segment length of the shot segments included in the sub-video and the shot extraction threshold, and the shot segments in the sub-video are extracted according to the number of frames to be extracted corresponding to the shot segments in the sub-video to obtain a video frame extraction sequence corresponding to the sub-video. Finally, the video frame extraction sequences corresponding to the sub-videos are spliced ​​to obtain a video frame extraction sequence corresponding to the video to be processed.

[0087] As an example, when splitting the video to be processed, you can first split the video to be processed into two segments on average to obtain two sub-videos, and then for each sub-video, determine whether the number of shots contained in the sub-video is greater than the shot number threshold. If the number of shots contained in a sub-video is still greater than the shot number threshold, continue to split the sub-video into two segments on average again until the number of shots contained in each sub-video is no greater than the shot number threshold.

[0088] As an example, when splitting a video to be processed, the ratio of the number of shots in the video to be processed divided by the shot number threshold can be calculated, and the ratio is rounded up. The video to be processed is then evenly split into a corresponding number of sub-videos based on the rounded-up result. For example, if the ratio of the number of shots in the video to be processed divided by the shot number threshold is 2.4, and the rounded-up result is 3, the video to be processed is evenly split into 3 sub-videos. For each sub-video, it is then determined whether the number of shots in the sub-video is greater than the shot number threshold. If so, the video is split again until the number of shots in each sub-video is no greater than the shot number threshold.

[0089] As an example, when segmenting a video to be processed, segmentation can be performed based on a shot count threshold. For example, all shot segments in the video to be processed can be determined, and adjacent shot segments that meet the shot count threshold can be divided into a sub-video, resulting in at least two sub-videos. For example, if the shot count threshold is 3, and the video to be processed contains 7 shot segments, the first 3 shot segments can be divided into a sub-video, the 4th to 6th shot segments can be divided into a sub-video, and the last shot segment can be divided into a sub-video.

[0090] The video frame extraction scheme provided by the above-mentioned embodiments of the present disclosure effectively captures a video frame sequence that is fed into a large model for video content analysis. During model training and inference applications, the extracted video frame sequence is used instead of the original video input into the large model, ensuring better video description results with a smaller input size. Due to the reduced input size, a large model with a smaller number of parameters can be used to replace a larger model or a closed-source, paid model.

[0091] Assume that a large model that can be fine-tuned (such as the MiniCPMv2.6 7B model) is selected. During the training phase, the format of the training dataset (a training sample includes a sequence of video frames extracted from the video to be understood, the questions asked, and the corresponding answers) is:

[0092]

[0093] In the application reasoning phase, the format of the dataset to be reasoned (a piece of data includes a video frame sequence extracted from the video to be analyzed and understood and a question asked, and the expected output is the multimodal large model's understanding of the video content) is:

[0094]

[0095]

[0096] Figure 4 This is a schematic diagram of the structure of a video frame extraction device provided by an embodiment of the present disclosure, which can be implemented by software and / or hardware. Figure 4 As shown, the video frame extraction device 40 includes: a first determination module 410 , a second determination module 420 , a frame extraction module 430 and a third determination module 440 .

[0097] The first determining module 410 is configured to determine the storyboard segments and the duration of the storyboard segments included in the video to be processed;

[0098] The second determining module 420 is configured to determine, as the number of frames to be extracted for the shot segment, a sum of a preset shot extraction threshold and an allocated frame extraction number determined based on a segment duration, wherein the allocated frame extraction number is positively correlated with the segment duration.

[0099] The frame extraction module 430 is configured to extract frames from the shot segment based on the number of frames to be extracted, to obtain a shot frame extraction sequence.

[0100] The third determining module 440 is configured to determine, based on the shot frame extraction sequence, a video frame extraction sequence corresponding to the video to be processed.

[0101] Optionally, the second determining module 420 comprises:

[0102] The obtaining unit is configured to obtain a shot number of the shot segment and a video duration of the video to be processed.

[0103] The first determining unit is configured to determine, based on the shot number and the video duration, a total number of frames to be extracted, wherein the total number of frames to be extracted is related to a larger one of a first frame extraction number determined according to the video duration and a second frame extraction number determined according to the shot number.

[0104] The second determining unit is configured to determine, as a remaining number of frames to be allocated, a difference obtained by subtracting a product of the shot number and the shot extraction threshold from the total number of frames to be extracted.

[0105] The allocating unit is configured to allocate the remaining number of frames to be allocated to the shot segment according to the segment duration, to obtain the allocated frame extraction number corresponding to the shot segment.

[0106] Further optionally, the first determining unit is further configured to:

[0107] obtain a preset lower limit of the video duration, an upper limit of the video duration, a minimum frame extraction number, and a maximum frame extraction number;

[0108] determine a ratio of a first difference between the video duration and the lower limit of the video duration divided by a second difference between the upper limit of the video duration and the lower limit of the video duration;

[0109] determine, based on a product of a third difference between the maximum frame extraction number and the minimum frame extraction number multiplied by the ratio, a frame extraction number increment;

[0110] determine the first frame extraction number based on a sum of the frame extraction number increment and the minimum frame extraction number;

[0111] determine the total number of frames to be extracted based on a larger one of the first frame extraction number and a second frame extraction number obtained by multiplying the shot number and a preset frame extraction number.

[0112] Further optionally, the first determining unit is further configured to:

[0113] obtain the preset frame extraction number;

[0114] Determine the product of the preset value of the number of extracted frames and the number of storyboards to obtain a second number of extracted frames;

[0115] The larger value between the second frame extraction number and the first frame extraction number is used as the temporary frame extraction number;

[0116] Based on the size relationship between the temporary number of extracted frames and the maximum number of extracted frames, the total number of frames to be extracted is determined, wherein, when the temporary number of extracted frames is not greater than the maximum number of extracted frames, the temporary number of extracted frames is determined as the total number of frames to be extracted; when the temporary number of extracted frames is greater than the maximum number of extracted frames, the maximum number of extracted frames is determined as the total number of frames to be extracted.

[0117] Optionally, the allocation unit is further configured to:

[0118] Determine whether the number of remaining frames to be allocated is 0;

[0119] When the number of remaining frames to be allocated is not zero, the ratio of the duration of the storyboard segment to the currently allocated number of frames corresponding to the segment is determined to obtain the frame extraction interval. The currently allocated number of frames is the sum of the storyboard frame extraction threshold and the number of allocated frames. The initial value of the number of allocated frames is 0.

[0120] The number of allocated frames corresponding to the target storyboard with the largest frame extraction interval is increased by a preset value, and the number of remaining frames to be allocated is reduced by a preset value to update the number of remaining frames to be allocated;

[0121] Return to the step of determining whether the remaining number of frames to be allocated is 0, and when the remaining number of frames to be allocated is 0, obtain the number of allocated frames corresponding to the storyboard segment. Optionally, the frame extraction module 430 is further configured to:

[0122] Get the first and last frames of the storyboard;

[0123] Based on the number of frames to be extracted, determining the number of remaining extracted frames;

[0124] Based on the remaining number of frames, extract the remaining number of target picture frames from the storyboard;

[0125] Based on the first picture frame, the target picture frame and the last picture frame, the storyboard frame extraction sequence corresponding to the storyboard clip is determined.

[0126] Optionally, the video frame extraction device 40 further includes:

[0127] A judgment module, used to judge whether the number of storyboards is greater than a storyboard number threshold;

[0128] A video segmentation module is used to segment the video to be processed into at least two sub-videos when the number of storyboards is greater than a storyboard number threshold, so as to perform frame extraction processing on the at least two sub-videos respectively, wherein the number of frames to be extracted corresponding to the storyboard segments in the sub-video is determined according to the segment length of the storyboard segments included in the sub-video and the storyboard frame extraction threshold, and the storyboard segments in the sub-video are extracted according to the number of frames to be extracted corresponding to the storyboard segments in the sub-video to obtain a video frame extraction sequence corresponding to the sub-video.

[0129] The video frame extraction device provided in the embodiments of the present disclosure can execute the video frame extraction method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0130] In order to implement the above embodiments, the present disclosure further proposes a computer program product, including a computer program / instruction, which implements the video frame extraction method in the above embodiments when executed by a processor.

[0131] Figure 5 A schematic structural diagram of an electronic device provided in an embodiment of the present disclosure.

[0132] The following specific reference Figure 5 , which shows a schematic diagram of the structure of an electronic device 500 suitable for implementing the embodiment of the present disclosure. It can be understood that, Figure 5 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0133] like Figure 5 As shown, the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the electronic device 500 are also stored in the RAM 503. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0134] Typically, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 5The electronic device 500 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.

[0135] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the video frame extraction method of the embodiment of the present disclosure are performed.

[0136] It should be noted that the computer-readable medium disclosed herein may be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0137] In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the foregoing.

[0138] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0139] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device performs the functions defined in the video frame extraction method of the embodiment of the present disclosure.

[0140] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a Local Area Network (LAN) or a Wide Area Network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0141] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0142] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit does not necessarily limit the unit itself.

[0143] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0144] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0145] According to one or more embodiments of the present disclosure, the present disclosure provides an electronic device, including:

[0146] processor;

[0147] a memory for storing instructions executable by the processor;

[0148] The processor is configured to read the executable instructions from the memory and execute the instructions to implement any one of the video frame extraction methods provided in the present disclosure.

[0149] According to one or more embodiments of the present disclosure, the present disclosure provides a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer program is used to execute any of the video frame extraction methods provided by the present disclosure.

[0150] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0151] Moreover, while operations are depicted in a particular order, this should not be understood as requiring such an order nor infringing on the scope of the disclosure. Certain of the operations described in the discussion are combinable into a single operation, and certain operations can be separated into several operations. In some embodiments, the operations described in the discussion can be performed in an order different than presented in the discussion. In some embodiments, the operations described in the discussion can be performed concurrently. Also, while several specific implementation details are discussed in the discussion, these should not be interpreted as limiting the scope of the disclosure. Rather, certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.

[0152] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

Claims

1. A video frame extraction method, characterized in that: include: Determining the storyboard segments included in the video to be processed and the segment lengths of the storyboard segments; Determine the number of frames to be extracted corresponding to the storyboard segment by summing a preset storyboard frame extraction threshold and a number of allocated frames determined based on the segment duration, wherein the number of allocated frames is positively correlated with the segment duration; Extract frames from the storyboard segment based on the number of frames to be extracted to obtain a storyboard frame extraction sequence; A video frame extraction sequence corresponding to the video to be processed is determined based on the storyboard frame extraction sequence.

2. The video frame extraction method according to claim 1, wherein: Determining the number of allocated frames based on the segment duration includes: Obtaining the number of storyboards of the storyboard segment and the video length of the video to be processed; Determining a total number of frames to be extracted based on the number of shots and the duration of the video, wherein the total number of frames to be extracted is related to a larger value of a first number of frames to be extracted determined according to the duration of the video and a second number of frames to be extracted determined according to the number of shots; The difference obtained by subtracting the product of the number of storyboards and the storyboard frame extraction threshold from the total number of frames to be extracted is determined as the number of remaining frames to be allocated; The remaining number of frames to be allocated is allocated to the storyboard segment according to the length of the segment, so as to obtain the number of allocated frames corresponding to the storyboard segment.

3. The video frame extraction method according to claim 2, wherein: The determining the total number of frames to be extracted based on the number of storyboards and the length of the video includes: Get the preset video duration lower limit, video duration upper limit, minimum number of extracted frames, and maximum number of extracted frames; Determine a ratio of a first difference between the video duration and the video duration lower limit divided by a second difference between the video duration upper limit and the video duration lower limit; determining an increment of the number of frame extractions based on a product of a third difference between the maximum number of frame extractions and the minimum number of frame extractions and the ratio; Determining the first number of extracted frames based on the sum of the frame extraction number increment and the minimum number of extracted frames; The total number of frames to be extracted is determined based on the larger value of the first number of extracted frames and the second number of extracted frames obtained by multiplying the number of split mirrors and a preset value of the number of extracted frames.

4. The video frame extraction method according to claim 3, wherein: The determining the total number of frames to be extracted based on the larger value of the first number of extracted frames and the second number of extracted frames obtained by multiplying the number of split screens by a preset value of the number of extracted frames comprises: Obtaining the preset value of the number of frames extracted; Determine the product of the preset value of the number of extracted frames and the number of split screens to obtain the second number of extracted frames; The larger value between the second frame extraction number and the first frame extraction number is used as a temporary frame extraction number; Based on the size relationship between the temporary number of extracted frames and the maximum number of extracted frames, the total number of frames to be extracted is determined, wherein, when the temporary number of extracted frames is not greater than the maximum number of extracted frames, the temporary number of extracted frames is determined as the total number of frames to be extracted; when the temporary number of extracted frames is greater than the maximum number of extracted frames, the maximum number of extracted frames is determined as the total number of frames to be extracted.

5. The video frame extraction method according to claim 2, wherein: The allocating the remaining number of frames to be allocated to the storyboard segments according to the segment duration to obtain the number of allocated frames corresponding to the storyboard segments includes: Determine whether the remaining number of frames to be allocated is 0; When the remaining number of frames to be allocated is not zero, determining a ratio of the duration of the storyboard segment to the currently allocated number of frames corresponding to the storyboard segment, to obtain a frame extraction interval, wherein the currently allocated number of frames is the sum of the storyboard frame extraction threshold and the number of allocated frames, and an initial value of the number of allocated frames is 0; The number of allocated frames corresponding to the target storyboard with the largest frame extraction interval is added to a preset value, and the remaining number of frames to be allocated is reduced by the preset value to update the remaining number of frames to be allocated; Return to the step of determining whether the remaining number of frames to be allocated is 0, until the remaining number of frames to be allocated is 0, and obtain the number of allocated frames corresponding to the storyboard segment.

6. The video frame extraction method according to claim 1, wherein: Extracting frames from the storyboard segments based on the number of frames to be extracted to obtain a storyboard frame extraction sequence includes: Obtaining the first and last image frames of the storyboard segment; Determining the remaining number of frames to be extracted based on the number of frames to be extracted; Based on the remaining number of extracted frames, extract the remaining number of target picture frames from the storyboard segment; Based on the first picture frame, the target picture frame and the last picture frame, a storyboard frame extraction sequence corresponding to the storyboard segment is determined.

7. The method according to any one of claims 2 to 6, characterized in that: After determining the number of storyboards contained in the video to be processed, the method further includes: Determine whether the number of storyboards is greater than a storyboard number threshold; When the number of storyboards is greater than the storyboard number threshold, the video to be processed is divided into at least two sub-videos, wherein the number of frames to be extracted corresponding to the storyboard segments in the sub-video is determined according to the segment length of the storyboard segments included in the sub-video and the storyboard frame extraction threshold, and the storyboard segments in the sub-video are frame extracted according to the number of frames to be extracted corresponding to the storyboard segments in the sub-video to obtain a video frame extraction sequence corresponding to the sub-video.

8. A video frame extraction device, characterized in that: include: A first determining module is used to determine the storyboard segments included in the video to be processed and the segment lengths of the storyboard segments; A second determining module is configured to determine the number of frames to be extracted corresponding to the storyboard segment by taking the sum of a preset storyboard frame extraction threshold and a number of allocated frames determined based on the segment duration as the number of frames to be extracted, wherein the number of allocated frames is positively correlated with the segment duration; A frame extraction module is used to extract frames from the storyboard segment based on the number of frames to be extracted to obtain a storyboard frame extraction sequence; The third determining module is used to determine the video frame extraction sequence corresponding to the video to be processed based on the storyboard frame extraction sequence.

9. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing instructions executable by the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the video frame extraction method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and the computer program is used to execute the video frame extraction method according to any one of claims 1 to 7.