Video frame extraction method, device, apparatus and medium
By using an automated video frame extraction method to remove black screens and duplicate frames, the quality of extracted key frames is ensured, solving the problems of low efficiency and unstable accuracy in existing technologies, and achieving efficient video frame extraction and AI recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN DIANMAO TECH CO LTD
- Filing Date
- 2026-04-10
- Publication Date
- 2026-05-29
Smart Images

Figure CN122120403A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video processing technology, specifically to video frame extraction methods, apparatus, devices, and media. Background Technology
[0002] In the field of digital content processing, video, as a core information carrier, has seen its keyframe extraction and subsequent processing needs permeate multiple industries, including education, media, and enterprise services. With the rise of short video platforms, the widespread adoption of AI recognition technology, and the accelerated digital transformation of enterprises, various industries are placing higher demands on the automation level, processing efficiency, and result quality of video frame extraction. For example, the education industry needs to extract keyframes from classroom videos to assist in AI-based homework grading and courseware generation; content moderation requires frame extraction for rapid screening of inappropriate content; and media archiving scenarios face the challenge of generating and retrieving summaries from large volumes of videos.
[0003] Current traditional video frame extraction solutions mostly rely on manual intervention or single tools, which have problems such as fragmented processes, low processing efficiency, and unstable result quality: manual frame extraction is time-consuming and labor-intensive, making it difficult to meet the needs of large-scale video processing; simple tool-based processing lacks intelligent optimization mechanisms, which can easily result in extracted key frames that do not meet the requirements, affecting the accuracy of downstream AI recognition. Summary of the Invention
[0004] In view of this, this application provides a video frame extraction method, apparatus, device and medium to solve the problem that the quality and other defects of the extracted keyframes affect the subsequent processing effect.
[0005] In a first aspect, this application provides a video frame extraction method, the method comprising: Obtain a frame extraction request for extracting frames from the target video; the frame extraction request includes the target video or the resource address of the target video; Determine the frame extraction configuration parameters for the target video; the frame extraction configuration parameters include the total number of frames to be extracted; Extract the total number of video frames from the target video as candidate frames; Remove black screen frames and duplicate frames that overlap with other candidate frames from the candidate frames, and take the remaining candidate frames as valid frames to form a set of valid frames. If the number of frames in the set of valid frames does not meet the requirements, valid frames are extracted from the target video and added to the set of valid frames until the number of frames in the set of valid frames meets the requirements or it is determined that there are no video frames in the target video that can be used as valid frames.
[0006] Secondly, this application provides a video frame extraction device, the device comprising: The acquisition module is used to acquire a frame extraction request for extracting frames from a target video; the frame extraction request includes the target video or the resource address of the target video; The parameter determination module is used to determine the frame extraction configuration parameters of the target video; the frame extraction configuration parameters include the total number of frames to be extracted; A frame extraction module is used to extract video frames from the target video as candidate frames based on the total number of extracted frames. The processing module is used to remove black screen frames and duplicate frames that are the same as other candidate frames from the candidate frames, and to take the remaining candidate frames as valid frames to form a set of valid frames; if the number of frames in the set of valid frames does not meet the requirements, valid frames are extracted from the target video and added to the set of valid frames until the number of frames in the set of valid frames meets the requirements or it is determined that there are no video frames in the target video that can be used as valid frames.
[0007] Thirdly, this application provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the video frame extraction method described in the first aspect or any corresponding embodiment.
[0008] Fourthly, this application provides a computer-readable storage medium storing computer instructions for causing a computer to perform the video frame extraction method described in the first aspect or any corresponding embodiment thereof.
[0009] Fifthly, this application provides a computer program product, including computer instructions for causing a computer to execute the video frame extraction method described in the first aspect or any corresponding embodiment thereof.
[0010] The video frame extraction method provided in this application, upon receiving a user-initiated frame extraction request, first extracts a certain number of candidate frames based on the total number of frames to be extracted. Then, it performs black screen detection and duplicate detection on these candidate frames, retaining only those that pass the detection. By removing black screen and duplicate frames, the effectiveness and diversity of the frame extraction results are ensured, directly improving the accuracy of downstream AI recognition and content review. Furthermore, when the number of valid frames does not meet the requirements, frame extraction can be automatically repeated to ensure successful execution of the frame extraction task, improving the task success rate. It also enables end-to-end frame extraction, reducing manual intervention. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0012] Figure 1 This is a schematic diagram illustrating an application scenario according to an embodiment of this application; Figure 2 This is a schematic diagram of the first type of video frame extraction method according to an embodiment of this application; Figure 3 This is a schematic diagram of a second process for a video frame extraction method according to an embodiment of this application; Figure 4 This is a structural block diagram of a video frame extraction device according to an embodiment of this application; Figure 5 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] It is understood that before using the technical solutions disclosed in the various embodiments of this application, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this application in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0015] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0016] As one optional application scenario in the embodiments of this application, such as Figure 1 As shown, application 101 is installed in terminal device 110, and user 130 can interact with application 101 through terminal device 110 and / or access device of terminal device 110.
[0017] For example, application 101 can be any application. For instance, application 101 could be a classroom teaching application, a video review application, etc. Figure 1 In the application scenario shown, if application 101 is active, the terminal device 110 can display the interface 102 of application 101. The interface 102 may include various pages that application 101 can provide, such as interactive pages, settings pages, query pages, etc.
[0018] In some embodiments, terminal device 110 is communicatively connected to server 120 to provide services to application 101. Terminal device 110 may be a mobile terminal, fixed terminal, or portable terminal, including but not limited to mobile phones, desktop computers, laptop computers, multimedia tablets, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, terminal device 110 may also support any type of interface, and server 120 may be various types of computing systems or servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, and computing devices in cloud environments.
[0019] It should be noted that, Figure 1 This is merely an example of an application scenario and does not limit the scope of protection of this application.
[0020] According to an embodiment of this application, a video frame extraction method embodiment is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0021] This embodiment provides a video frame extraction method, which can be used in the aforementioned terminal device or server. Figure 2 This is a flowchart of a video frame extraction method according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps.
[0022] Step S201: Obtain a frame extraction request for extracting frames from the target video; the frame extraction request includes the target video or the resource address of the target video.
[0023] In this embodiment, when it is necessary to extract frames from a video, that video can be used as the target video, and a frame extraction request can be initiated for that target video. For example, the target video could be a classroom video during a teaching process, or a video awaiting review for any inappropriate content.
[0024] The frame extraction request can directly include the target video. For example, the frame extraction request can include the corresponding video file, supporting chunked uploads; large files can be uploaded in multiple chunks. Alternatively, the frame extraction request can include the resource address of the target video, without including the target video itself. For example, the frame extraction request can include the URL address of the target video, and the server can asynchronously pull the video through the downloader, supporting resumeable downloads and automatic retrying in case of network errors.
[0025] After the server obtains the target video, it can store it in a local temporary directory, which will be automatically cleaned up after the task is completed.
[0026] For any frame extraction request, a corresponding frame extraction task can be created and assigned a unique task identifier (e.g., task ID). Users can query the processing progress and results through the task identifier. Of course, the frame extraction request can directly contain the task identifier, and this embodiment does not limit this.
[0027] Step S202: Determine the frame extraction configuration parameters for the target video; the frame extraction configuration parameters include the total number of frames to be extracted.
[0028] In this embodiment, each frame extraction task requires the configuration of certain parameters, namely frame extraction configuration parameters. These frame extraction configuration parameters may specifically include, for example, the total number of frames to be extracted (N). Each frame extraction configuration parameter has a default value. After acquiring the target video, the system can automatically determine the optimal configuration parameters best suited for the target video through content-aware analysis. Alternatively, the user can also actively specify the frame extraction configuration parameters; for example, the frame extraction request may also include one or more frame extraction configuration parameters such as the total number of frames to be extracted (N).
[0029] Step S203: Extract the total number of video frames from the target video as candidate frames.
[0030] When extracting frames from the target video, the extraction is based on the total number of frames N, that is, N video frames are extracted from the target video. These N video frames are the content that needs to be processed later. For ease of description, the video frames extracted at this time are called "candidate frames". It can be understood that the number of candidate frames is also N.
[0031] For example, the target video can be uniformly sampled, or sampled at fixed intervals, to extract N candidate frames.
[0032] Step S204: Remove black screen frames and duplicate frames that are the same as other candidate frames from the candidate frames, and take the remaining candidate frames as valid frames to form a set of valid frames.
[0033] In this embodiment, the candidate frames initially selected from the target video may not meet the subsequent requirements. Specifically, during the video acquisition or generation process, some frames in the target video may be black due to faults or other reasons, and they basically do not contain any useful information. Therefore, it is necessary to filter out these black screen frames from the candidate frames and remove them.
[0034] Some candidate frames are highly similar to each other and are duplicate frames, so duplicate frames also need to be removed. Only one of the duplicate frames needs to be kept. After removing black screen frames and duplicate frames, the remaining candidate frames can be used as the final keyframes, i.e., the valid frames.
[0035] Furthermore, a set of valid frames is pre-set for recording valid frames. Each time a candidate frame is determined to be a valid frame, the candidate frame can be added to the set of valid frames.
[0036] Step S205: If the number of frames in the valid frame set does not meet the requirements, extract valid frames from the target video and add them to the valid frame set until the number of frames in the valid frame set meets the requirements or it is determined that there are no video frames in the target video that can be used as valid frames.
[0037] In this embodiment, a certain requirement is set for the number of valid frames. Generally, the number of valid frames needs to reach the total number of frames to be extracted, N. If the candidate frames do not contain black screen frames or duplicate frames, then all candidate frames will be considered valid frames, that is, the number of valid frames meets the requirement. At this time, the frame extraction task can be completed, and the downstream module (such as an AI recognition platform, content management system, etc.) will identify and detect the extracted valid frames.
[0038] Conversely, if the candidate frames contain black screen frames and / or duplicate frames, resulting in a preliminary number of valid frames being less than N, then valid frames need to be re-extracted from the target video until the number of frames in the valid frame set (i.e., the number of valid frames) meets the requirement. Alternatively, if all video frames in the target video have been traversed and there are no more video frames that can be used as valid frames, the frame extraction process stops even if the number of frames in the valid frame set is still less than N. For example, after a preset number of re-extraction processes (e.g., 3 times), regardless of whether the number of valid frames reaches N, it can be considered that there are no video frames in the target video that can be used as valid frames, and the processing stops. The valid frames in the valid frame set at this point are then used as the target for subsequent processing.
[0039] The video frame extraction method provided in this embodiment, upon receiving a user-initiated frame extraction request, first extracts a certain number of candidate frames based on the total number of frames to be extracted. Then, it performs black screen detection and duplicate detection on these candidate frames, retaining only those that pass the detection. By removing black screen and duplicate frames, the effectiveness and diversity of the frame extraction results are ensured, directly improving the accuracy of downstream AI recognition and content review. Furthermore, if the number of valid frames does not meet the requirements, frame extraction can be automatically repeated to ensure successful execution of the frame extraction task, improving the task success rate. It also enables end-to-end frame extraction, reducing manual intervention.
[0040] This embodiment provides a video frame extraction method, which can be used in the aforementioned terminal device or server. Figure 3 This is a flowchart of a video frame extraction method according to an embodiment of this application, such as... Figure 3 As shown, the process includes the following steps.
[0041] Step S301: Obtain a frame extraction request for extracting frames from the target video; the frame extraction request includes the target video or the resource address of the target video.
[0042] Please see details Figure 2 Step S201 of the illustrated embodiment will not be described again here.
[0043] Optionally, the method further includes: performing format verification on the target video; and if the target video passes the format verification, extracting basic attribute information of the target video. The basic attribute information includes: video duration, frame rate, and resolution. This basic attribute information can be used to subsequently determine frame extraction configuration parameters.
[0044] In this embodiment, after acquiring the target video, the format is first validated. For example, the FFmpeg interface can be called to open the input target video file and perform format validation. If the format is not supported (e.g., format incompatibility), the task failure status is returned directly. For target videos with valid formats, they can pass the format validation, and then the basic attribute information of the target video can be extracted.
[0045] The basic attribute information includes: video duration. (Unit: seconds), Frame Rate At least one of the following attributes: frames per second, resolution, etc., video preprocessing can be performed based on this fundamental attribute information. Among these, video duration... The density of sparse sampling can be determined (e.g., longer videos are sampled more sparsely), and the resolution is used to determine whether downsampling is needed, for example, automatically recommending a suitable compression size based on the original resolution to improve analysis efficiency.
[0046] Specifically, the total number of video frames in the target video can be determined based on basic attribute information. ,and , This refers to the rounding function, such as the rounding function.
[0047] Furthermore, during the frame extraction process, the number of processed frames can be divided by the total number of video frames. It calculates the percentage of processing progress, performs progress estimation, and can display a progress bar on the front end.
[0048] Step S302: Determine the frame extraction configuration parameters for the target video; the frame extraction configuration parameters include the total number of frames to be extracted.
[0049] Please see details Figure 2 Step S202 of the illustrated embodiment will not be described again here.
[0050] In some optional implementations, step S302 above, which determines the frame extraction configuration parameters of the target video, includes: extracting the frame extraction configuration parameters of the target video from the frame extraction request; or, performing content-aware analysis on the target video to determine the frame extraction configuration parameters of the target video.
[0051] In this embodiment, the frame extraction configuration parameters may specifically include: the total number of frames N, the brightness threshold T, the similarity threshold S, image output parameters, and the distribution density parameter α. Each frame extraction configuration parameter has a default value. After acquiring the target video, the system can automatically determine the optimal configuration parameters best suited for the target video through content-aware analysis. Alternatively, the user can actively specify the frame extraction configuration parameters; for example, the frame extraction request may include one or more frame extraction configuration parameters.
[0052] For example, the total number of frames N: the number of valid keyframes to be extracted, with a default value of 10 and a configurable range of 1 to 1000.
[0053] Brightness threshold T: The average pixel brightness threshold for judging black screen frames. Default value: 20, range: 0~255.
[0054] Similarity threshold S: The threshold for determining the inter-frame hash similarity of duplicate frames. Default value: 0.95, range: 0~1.
[0055] Image output parameters: configuration parameters such as output format, quality, and size, which can specifically be image compression parameters.
[0056] Distribution density parameter α: controls the temporal distribution density of the frame. Default value: 1.0, at which point the distribution density can conform to a linear distribution.
[0057] Optionally, if the user does not specify the distribution density parameter α, the system can automatically determine a suitable distribution density parameter α. Specifically, the process of determining the distribution density parameter includes the following steps a1 to a5.
[0058] Step a1: Perform sparse sampling on the target video to extract sample frames of the sample frame number; there is a positive correlation between the sample frame number and the video duration of the target video.
[0059] In this embodiment, the number of sample frames used for content-aware analysis can be determined. Specifically, it is based on the video duration of the target video. Determine the number of sample frames M, and establish a positive correlation between the two. For example, ,in, This represents the floor function.
[0060] To avoid the performance overhead of full-frame analysis, the system first performs sparse sampling on the target video, extracting M video frames at uniform intervals as sample frames for subsequent content analysis. Sample Frame Set It can be represented as: , Let i be the i-th sample frame.
[0061] Step a2: For each sample frame, determine the motion intensity and information density corresponding to the sample frame.
[0062] Specifically, for each sample frame, the average amplitude of its optical flow field is calculated and used as the motion intensity of the sample frame at the corresponding time point; and its information density is calculated based on Laplacian edge detection.
[0063] If the time point corresponding to the sample frame is t, then its motion intensity Information density They are respectively: ; ; Where W and H are the width and height of the sample frame, Indicates the position coordinates within the sample frame. , These are the components of optical flow in the x and y directions, respectively. The Laplacian operator for the sample frame.
[0064] Step a3: Determine the first average motion intensity and the first average information density based on the motion intensity and information density of each sample frame in the first time period; determine the second average motion intensity and the second average information density based on the motion intensity and information density of each sample frame in the second time period; the first time period is located in the first half of the target video, and the second time period is located in the second half of the target video.
[0065] In this embodiment, the target video can be divided into two parts, namely the first half and the second half, to analyze which part of the target video is more important, thereby extracting a larger number of video frames from the more important part. For example, the time period corresponding to the first half is designated as the first time period, and the time period corresponding to the second half is designated as the second time period. For example, if the length of the target video is 30 seconds, then the first time period corresponds to 0~15 seconds, and the second time period corresponds to 15 seconds~30 seconds.
[0066] Furthermore, based on the motion intensity and information density of the sample frames in each time period, the relevant information for each time period can be comprehensively determined.
[0067] As the names suggest, the first average motion intensity is the average motion intensity of all sample frames within the first time period, and the first average information density is the average information density of all sample frames within the first time period. The second average motion intensity is the average motion intensity of all sample frames within the second time period, and the second average information density is the average information density of all sample frames within the second time period.
[0068] Step a4: Determine the exercise intensity ratio based on the first average exercise intensity and the second average exercise intensity; determine the information density ratio based on the first average information density and the second average information density.
[0069] Among them, the second average exercise intensity With the first average exercise intensity The ratio of the two is used as the ratio of exercise intensity. ,Right now Similarly, the second average information density... Compared with the first average information density The ratio of information density is used as the information density ratio. ,Right now .
[0070] Step a5: Determine the distribution density parameter based on the exercise intensity ratio and the information density ratio; the distribution density parameter is positively correlated with both the exercise intensity ratio and the information density ratio.
[0071] In this embodiment, a larger motion intensity ratio or information density ratio indicates that the motion intensity or information content in the latter half of the target video is greater, meaning the latter half is more important. Therefore, a larger distribution density parameter can be set. Among them, the distribution density parameter It is the power of the distribution of valid frames, that is, the distribution of valid frames is represented by a power function, and this distribution density parameter is... It is the exponent of the power distribution.
[0072] Optionally, step a5 "determines the distribution density parameter based on the motion intensity ratio and information density ratio, including steps a51 to a53."
[0073] Step a51: For any two adjacent sample frames, determine the scene switching time points in the target video based on the similarity between the two sample frames, and determine the distribution of scene switching based on each scene switching time point.
[0074] Specifically, sample frames With sample frames The structural similarity index (SSIM) between adjacent elements can be expressed as: The similarity between the two is represented by their SSIM scores.
[0075] like Less than the first preset threshold (For example, If a scene change occurs, it can be determined based on the sample frames. Sample frames The corresponding timestamp determines the time position of this scene switch, that is, the scene switch time point.
[0076] The distribution of scene transitions can be determined by the time when each scene transition occurs.
[0077] Specifically, the total number of scene switching times can be counted. If the number of scene transitions within the first time period is greater than the total number of transitions... The ratio is greater than the ratio of the number of scene transitions during the second time period to the total number of transitions. If the ratio is positive, it means that the scene switching mainly occurs in the first half; otherwise, it means that it mainly occurs in the second half.
[0078] Step a52: Determine the basic distribution parameters based on the ratio of motion intensity and the ratio of information density.
[0079] Step a53: Correct the basic distribution parameters according to the distribution of scene switching to obtain the distribution density parameters.
[0080] If a scene switch exceeding a preset proportion occurs in the first time period, the distribution density parameter is: .
[0081] If scene switching exceeding the preset proportion occurs in the second time period, the distribution density parameter is: The preset ratio is greater than 0.5.
[0082] In other cases, .in, For the distribution density parameter, Based on the basic distribution parameters, The ratio of exercise intensity Information density ratio; , All are preset coefficients, and , .
[0083] In this embodiment, the basic distribution parameters It can be represented as: Furthermore, it can be adjusted based on the distribution of scene transitions. Among these, To adjust the coefficient, ,For example wait.
[0084] Specifically, if more than a preset proportion (greater than 0.5, such as 0.7 or 0.8) of scene changes occur in the first time period, for example, more than 70% of the scene changes occur in the first half of the video (i.e., the number of scene changes in the first time period exceeds the total number of scene changes), the following criteria apply: If the ratio is greater than 70%, then the basic distribution parameter Specifically: , b is the parameter adjustment value, and For example, b=0.3.
[0085] Conversely, if a pre-defined proportion of scene changes occurs in the second time period, for example, more than 70% of scene changes occur in the latter half of the video, then the basic distribution parameter... Specifically: .
[0086] Apart from the above situations, the basic distribution parameters can be directly used. As a distribution density parameter .
[0087] Furthermore, the distribution density parameter can also be... Apply constraints to confine it to the range [c, d]; where c and d are preset coefficients, and 0 < c < d.<c≤0.5,d> 1.5, for example, c=0.5, d=2.
[0088] Optionally, the distribution density parameter α can be automatically determined based on the type of the target video, facilitating targeted processing by downstream business systems. Target videos can specifically include course videos, news videos, movies / TV shows, surveillance footage, etc.
[0089] For example, for course videos: there are few scene changes, low exercise intensity, and high content density in the second half, so α≈1.3~1.8.
[0090] For news videos: scene changes are frequent and content distribution is relatively even, so α≈0.9~1.1.
[0091] For film and television dramas: the scene transitions are moderate and there are plot climaxes, so α can adapt to the content.
[0092] For surveillance video: there are few scene changes and the motion intensity is uneven, so α can adapt to the motion distribution.
[0093] Optionally, the total number of frames N can be determined based on the scene switching frequency. There is a positive correlation between the two; that is, the higher the scene switching frequency, the larger the total number of frames N.
[0094] For example, if scene switching is frequent (e.g., >10 times / minute), then N = min(30, number of scene switching times × 2). If scene switching is frequent, then N = 15; if scene switching is infrequent (e.g., <2 times / minute), then N = 10.
[0095] Step S303: Extract the total number of video frames from the target video as candidate frames.
[0096] Most existing solutions employ uniform frame skipping or simple fixed-interval frame skipping, which fails to adapt to the content distribution characteristics of different video types. Even when non-uniform frame skipping is attempted, it often requires manual pre-setting of rules and cannot be automatically adjusted according to the video content, resulting in a simplistic and rigid frame skipping distribution strategy.
[0097] In this embodiment, instead of directly using uniform frame dropping, an appropriate distribution density parameter is set based on the actual situation of the target video. The distribution density parameter As an exponent of the power distribution, it is used for frame extraction based on the power distribution.
[0098] Specifically, as shown above, the frame extraction configuration parameters include distribution density parameters; and step S303, "extracting video frames from the target video with the total number of extracted frames as candidate frames", may include steps S3031 to S3032.
[0099] Step S3031: Determine the frame index of the total number of frames to be sampled based on the distribution density parameter; the frame index conforms to a power distribution with the distribution density parameter as the exponent.
[0100] Step S3032: Select the video frames corresponding to each frame index in the target video as candidate frames.
[0101] In this embodiment, the distribution density parameter is power-distributed, and the exponent is used to determine the frame index of the total number of frames N based on this power-distribution exponent, such that all frame indices conform to the exponent of the distribution density parameter. The power distribution allows for better selection of suitable candidate frames based on the key content of the target video.
[0102] In some alternative implementations, if the key content of the target video is in the first half, then the distribution density parameter... If the key content of the target video is less than 1, then the distribution density parameter is less than 1. Greater than 1; if the key content of the target video is in the center, then a Gaussian distribution is further fused, and the distribution density parameter at this time is... It can be near 1, for example Generally, you can set it directly. .
[0103] In this embodiment, based on the video duration of the target video Determine the time point corresponding to each frame index i = 1, 2, ..., N, where N is the total number of frames drawn.
[0104] Specifically, if the keyframes of the target video are located in the first or second half, meaning the important content of the target video is in the first or second half, then... ,or, .
[0105] Understandable, At that time, the time points corresponding to each frame index are linearly distributed and uniformly distributed on the time axis.
[0106] like The frame index is sparse at the beginning and dense at the end, with more frames concentrated in the second half of the video, which is suitable for scenarios where the key content is in the second half (such as course videos).
[0107] like If the frame index is denser at the beginning and sparser at the end, more frames will be concentrated in the first half of the video, which is suitable for scenarios where the key content is in the first half (such as news videos).
[0108] When the keyframe of the target video is located in the middle, As shown above, the distribution density parameter at this time It can be 1, that is .
[0109] in, For the time point corresponding to the index of the i-th frame, Let α be the video duration of the target video, and α be the distribution density parameter. To adjust the coefficient, and ,For example, To ensure the timing exist between. It is the inverse function of the standard normal distribution (quantile function).
[0110] Based on this, each time point can be determined. corresponding Frame index, and: .
[0111] in, For the index of the i-th frame, For the time point corresponding to the index of the i-th frame, The frame rate of the target video. Let i be the integer function; i = 1, 2, ..., N, where N is the total number of frames extracted. Based on this, N frame indices can be determined. Subsequently, based on these N frame indices, N candidate frames can be initially extracted from the target video. The positional distribution of the candidate frames in the target video conforms to the distribution density parameter. It is a power-law distribution of the exponent.
[0112] It is understandable that each frame index... It will not exceed the total number of video frames. .
[0113] In addition, if the key parts of the target video have a multi-peak distribution, such as a course video with multiple chapters, a power distribution or a Gaussian distribution can be used separately for each region, which will not be elaborated here.
[0114] Step S304: Remove black screen frames and duplicate frames that are the same as other candidate frames from the candidate frames, and take the remaining candidate frames as valid frames to form a set of valid frames.
[0115] Please see details Figure 2 Step S204 of the illustrated embodiment will not be described again here.
[0116] Optionally, as shown above, the frame extraction configuration parameters further include a brightness threshold and a similarity threshold. Furthermore, step S304, "removing black screen frames and duplicate frames that overlap with other candidate frames," may specifically include steps b1 to b5.
[0117] Step b1: For each candidate frame, determine the average brightness value of the candidate frame.
[0118] Step b2: Select candidate frames with an average brightness value greater than the brightness threshold as intermediate frames; among them, black screen frames are candidate frames with an average brightness value less than the brightness threshold.
[0119] In this embodiment, for the selected N candidate frames, the average brightness value L is calculated based on the color information of all pixels; for example: Where W and H are the width and height of the candidate frame, These are the RGB components of the pixel (x,y) in the candidate frame, namely the red component, green component, and blue component.
[0120] For a candidate frame, if its average brightness value L is less than a preset brightness threshold, the candidate frame is determined to be a black screen frame and is directly discarded. Furthermore, black screen detection can continue within a certain range (e.g., 5 frames before and after) around the black screen frame. If other video frames that are not black screen frames are detected, these other video frames are used as new candidate frames, replacing the candidate frame detected as a black screen frame. If all the frames within a certain range around the black screen frame are also black screen frames, then that index is skipped, and new frame extraction points are added later.
[0121] If the average brightness value L of the candidate frame is greater than the preset brightness threshold, the candidate frame can be considered a normal video frame and can be used as an intermediate frame for black screen detection and subsequent processing.
[0122] By removing black screen frames, invalid frames can be prevented from occupying storage resources, which can reduce cloud storage costs. Furthermore, black screen frames can be eliminated from interfering with downstream AI recognition, improving recognition accuracy. They can also reduce the amount of invalid computation in downstream systems and reduce overall computing power consumption.
[0123] Step b3: For each intermediate frame, determine the perceptual hash value between the intermediate frame and the valid frames in the valid frame set; the valid frame set is initially empty.
[0124] Step b4: If the perceptual hash value between the intermediate frame and at least one valid frame in the set of valid frames is greater than the similarity threshold, then the intermediate frame is a duplicate frame and is removed.
[0125] Step b5: If the perceptual hash value between the intermediate frame and any valid frame in the set of valid frames is less than the similarity threshold, then the intermediate frame is considered a valid frame.
[0126] For each intermediate frame that passes the black screen detection, it can be compared with the valid frames in the valid frame set to determine their similarity. Here, the similarity is represented by a perceptual hash value. Initially, the valid frame set is empty, so the first intermediate frame can be directly used as a valid frame and added to the valid frame set.
[0127] For any intermediate frame, if its perceptual hash value with a certain valid frame is greater than the similarity threshold, it indicates that the two are highly similar. This intermediate frame contains a large amount of duplicate information and should not be considered a valid frame; that is, the intermediate frame is a duplicate frame and can be directly removed. Conversely, if the intermediate frame is not similar to any valid frame, it can be retained and added to the set as a valid frame.
[0128] One approach is to set a sliding window for intermediate frames and compare them only with valid frames within the sliding window, such as comparing them with the five most recently determined valid frames, in order to improve comparison efficiency.
[0129] Step S305: If the number of frames in the valid frame set does not meet the requirements, extract valid frames from the target video and add them to the valid frame set until the number of frames in the valid frame set meets the requirements or it is determined that there are no video frames in the target video that can be used as valid frames.
[0130] Please see details Figure 2 Step S205 of the illustrated embodiment will not be described again here.
[0131] Optionally, step S305, "re-extracting valid frames from the target video and adding them to the valid frame set", includes steps c1 to c2.
[0132] Step c1 involves setting a preset number of frames before and after the retained valid frames and / or the selected candidate frames. Each video frame within the range is marked as used; for example, This is to avoid subsequent supplementary frames being too similar to already processed frames.
[0133] Step c2 involves extracting frames from the target video, excluding those marked as used.
[0134] Step c3: Perform black screen frame detection and duplicate frame detection on the extracted video frames, and take the video frames that are not black screen frames or duplicate frames as valid frames.
[0135] For example, frame skipping can be performed based on the distribution density parameter α mentioned above, which will not be elaborated upon in this embodiment. Furthermore, the number of frames skipped can remain the total number of frames N, or it can be determined according to the actual situation; no limitation is made here.
[0136] For example, appropriate valid frames can be selected by analyzing the content of the target video. These could include positions near scene transition points (priority 1), positions with high motion intensity (priority 2), positions with high content density (priority 3), and positions far from already retained frames (priority 4). It's understood that for repeatedly extracted video frames, black screen detection and duplicate frame detection will also be performed. Only video frames that pass the detection will be considered valid frames and added to the valid frame set until the number of valid frames reaches N or there are no more video frames to choose from.
[0137] After extracting valid frames, compression and format conversion are performed according to preset image output parameters, and then the frames are uploaded in batches to designated cloud storage. After the upload is complete, local temporary video files and temporary frame files are automatically deleted to free up storage space. By uploading to cloud storage and automatically cleaning up local temporary files, processing efficiency and storage space usage are balanced; an automatic retry mechanism for task failures ensures task reliability in large-scale scenarios.
[0138] In addition, corresponding task notifications can be sent to users. Specifically, after a task is successfully processed, information such as the task processing status, the list of cloud storage URLs for valid keyframes, the task duration, automatically determined parameter values (e.g., distribution density parameter α), and the identified video type are encapsulated into a standardized message and pushed to downstream business systems via a message queue. Simultaneously, the task status in the database is updated to success. If the task processing fails, it is automatically retried according to the configured number of retries; if the retry count is exceeded, a failure notification is pushed. By connecting with downstream systems through a message queue, the coupling between modules can be reduced, allowing for flexible reuse in multiple business scenarios such as education, content moderation, and media archiving, resulting in strong scalability.
[0139] In some alternative implementations, when it is necessary to modify the frame extraction configuration parameters such as the total number of frames N, image output parameters (e.g., compression parameters), and storage path, it is generally necessary to restart the system. It does not support cloud hot updates, has high operation and maintenance costs, and cannot quickly respond to business needs adjustments.
[0140] In related solutions, frame extraction configuration parameters are stored in static memory as configuration objects, and are not reloaded during runtime. However, resources such as thread pools, database connection pools, and FFmpeg instances depend on configuration parameters (e.g., thread count, connection timeout, buffer size) during initialization. These resources cannot be dynamically adjusted after configuration changes and must be destroyed and rebuilt. Some configuration parameters are often hard-coded directly into the business logic as constants or static variables, or cached in multiple places, lacking a unified configuration management mechanism. Furthermore, the absence of a configuration version management mechanism makes it impossible to track change history, making rollback difficult after problems occur. This embodiment solves these problems through a distributed configuration center architecture, allowing offline configuration modifications without restarting the system.
[0141] Specifically, the system divides configuration into three levels: global default configuration, business scenario configuration, and task-level configuration. The global default configuration is stored in the configuration center and is the basic configuration shared by all tasks; the business scenario configuration is a preset configuration template for different business scenarios (education, media, auditing, etc.); and the task-level configuration is the personalized configuration parameters that can be covered by a single frame extraction task.
[0142] When the frame extraction service starts, it pulls the initial configuration (such as the global default configuration) from the configuration center and establishes a long connection to listen for configuration change events. When a user modifies the configuration in the management backend and publishes it, the configuration center pushes the change event to all subscribed service nodes; after receiving the configuration change notification, the service nodes atomically update the configuration object in memory without restarting the service.
[0143] The new configuration takes effect immediately on newly submitted tasks, while tasks that are currently running can choose to continue using the old configuration or switch smoothly.
[0144] In this embodiment, resources such as thread pools and connection pools are encapsulated in ResourceManager. After the configuration is changed, ResourceManager starts a new resource pool instance, the old resource pool continues to process the submitted tasks, new tasks use the new resource pool, and the old resource pool is automatically destroyed when it is idle.
[0145] Furthermore, each frame extraction task takes a snapshot of the current configuration at the start. This snapshot is used throughout the task execution, unaffected by mid-task configuration changes; new frame extraction tasks automatically use the latest configuration. This achieves a smooth transition from old configurations for old tasks to new configurations for new tasks.
[0146] In addition, it supports configuration version management, generating a new version number with each change. Service nodes verify the configuration version before processing tasks to ensure the latest configuration is used. Furthermore, it provides a one-click rollback function, allowing for quick restoration to the previous version in case of problems. All configuration changes are recorded in audit logs, including the person making the change, the time of the change, the content of the change, and the reason for the change.
[0147] The video frame extraction method provided in this embodiment can realize a fully automated frame extraction chain: task submission → video processing → content-aware analysis → adaptive frame extraction → result optimization → cloud storage upload → downstream notification. Through multi-dimensional analysis such as scene change detection, motion intensity analysis, and content density assessment, it can automatically perceive an appropriate distribution density parameter α based on the video content, adapting to the content distribution characteristics of different videos without manual intervention. Applying the distribution density parameter α to the temporal distribution of video keyframes allows for flexible adjustment of the distribution density of effective frames based on content-aware results or user configuration, thereby improving the coverage of key information. Combining black screen detection (based on pixel brightness threshold) and duplicate frame removal (based on inter-frame similarity calculation) ensures the effectiveness and diversity of frame extraction results, directly improving the accuracy of downstream AI recognition and content review.
[0148] This embodiment also provides a video frame extraction device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0149] This embodiment provides a video frame extraction device, such as... Figure 4 As shown, the device includes: The acquisition module 401 is used to acquire a frame extraction request for extracting frames from a target video; the frame extraction request includes the target video or the resource address of the target video; The parameter determination module 402 is used to determine the frame extraction configuration parameters of the target video; the frame extraction configuration parameters include the total number of frames to be extracted; The frame extraction module 403 is used to extract video frames from the target video as candidate frames based on the total number of extracted frames. The processing module 404 is used to remove black screen frames and duplicate frames that are the same as other candidate frames from the candidate frames, and to take the remaining candidate frames as valid frames to form a set of valid frames; if the number of frames in the set of valid frames does not meet the requirements, valid frames are extracted from the target video and added to the set of valid frames until the number of frames in the set of valid frames meets the requirements or it is determined that there are no video frames in the target video that can be used as valid frames.
[0150] In some optional implementations, the frame-skipping configuration parameters include distribution density parameters; The step of extracting video frames from the target video according to the total number of extracted frames as candidate frames includes: The frame index of the total number of frames is determined based on the distribution density parameter; the frame index conforms to a power distribution with the distribution density parameter as the exponent. The video frames corresponding to each frame index in the target video are used as candidate frames.
[0151] In some optional implementations, the frame index satisfies:
[0152] in, For the index of the i-th frame, For the time point corresponding to the index of the i-th frame, The frame rate of the target video. The function is the floor function; i = 1, 2, ..., N, where N is the total number of frames drawn; When the keyframe of the target video is located in the first or second half, ,or, ; When the keyframe of the target video is located in the middle position. ; in, Let α be the video duration of the target video, and α be the distribution density parameter. It is the inverse function of the standard normal distribution. To adjust the coefficient, and .
[0153] In some alternative implementations, the distribution density parameter is determined based on the following method: The target video is sparsely sampled to extract sample frames of a certain number; the number of sample frames is positively correlated with the video duration of the target video. For each sample frame, determine the motion intensity and information density corresponding to the sample frame; Based on the motion intensity and information density of each sample frame within the first time period, a first average motion intensity and a first average information density are determined; based on the motion intensity and information density of each sample frame within the second time period, a second average motion intensity and a second average information density are determined; the first time period is located in the first half of the target video, and the second time period is located in the second half of the target video. The motion intensity ratio is determined based on the first average motion intensity and the second average motion intensity; the information density ratio is determined based on the first average information density and the second average information density. The distribution density parameter is determined based on the motion intensity ratio and the information density ratio; the distribution density parameter is positively correlated with both the motion intensity ratio and the information density ratio.
[0154] In some optional implementations, determining the distribution density parameter based on the motion intensity ratio and the information density ratio includes: For any two adjacent sample frames, the scene switching time points in the target video where scene switching exists are determined based on the similarity between the two sample frames, and the distribution of scene switching is determined based on each scene switching time point. The basic distribution parameters are determined based on the motion intensity ratio and the information density ratio. The basic distribution parameters are corrected based on the distribution of scene switching to obtain the distribution density parameters; Wherein, if scene switching exceeding a preset proportion occurs in the first time period, the distribution density parameter is: ; If scene switching exceeding a preset proportion occurs during the second time period, the distribution density parameter is: The preset ratio is greater than 0.5. In other cases, ; in, The distribution density parameter is... The basic distribution parameters are... The ratio of the stated exercise intensity The information density ratio; , All are preset coefficients, and , .
[0155] In some optional implementations, the frame-skipping configuration parameters further include: a brightness threshold and a similarity threshold; The removal of black screen frames and duplicate frames that overlap with other candidate frames from the candidate frames includes: For each candidate frame, determine the average brightness value of the candidate frame; Candidate frames with an average brightness value greater than the brightness threshold are used as intermediate frames; wherein, black screen frames are candidate frames with an average brightness value less than the brightness threshold. For each intermediate frame, determine the perceptual hash value between the intermediate frame and the valid frames in the valid frame set; the valid frame set is initially empty. If the perceptual hash value between the intermediate frame and at least one valid frame in the set of valid frames is greater than the similarity threshold, then the intermediate frame is a duplicate frame and is removed. If the perceptual hash value between the intermediate frame and any valid frame in the set of valid frames is less than the similarity threshold, then the intermediate frame is considered a valid frame.
[0156] In some optional implementations, the step of re-extracting valid frames from the target video and adding them to the valid frame set includes: Mark all video frames within a preset number of frames before and after the retained valid frames and / or the selected candidate frames as used; Frame extraction is performed on all video frames in the target video except those marked as used. The extracted video frames are subjected to black screen frame detection and duplicate frame detection. Video frames that are not black screen frames or duplicate frames are considered valid frames.
[0157] The video frame extraction apparatus provided in this disclosure can execute the video frame extraction method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the various modules and units described above are the same as in the corresponding embodiments described above, and will not be repeated here.
[0158] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0159] The following is a detailed reference. Figure 5The diagram illustrates a structural schematic suitable for implementing the electronic device described in the embodiments of this application. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from memory 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device. The processor 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0160] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0161] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a memory 508, or installed from a ROM 502. When the computer program is executed by the processor 501, it performs the functions defined in the video frame extraction method of embodiments of this application.
[0162] Figure 5 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0163] This application also provides a computer-readable storage medium. The methods described in this application can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the video frame extraction method shown in the above embodiments is implemented.
[0164] A portion of this application can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to this application through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0165] Although embodiments of this application have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of this application, and all such modifications and variations fall within the scope defined by the appended claims.
Claims
1. A video frame extraction method, characterized in that, The method includes: Obtain a frame extraction request for extracting frames from the target video; the frame extraction request includes the target video or the resource address of the target video; Determine the frame extraction configuration parameters for the target video; the frame extraction configuration parameters include the total number of frames to be extracted; Extract the total number of video frames from the target video as candidate frames; Remove black screen frames and duplicate frames that overlap with other candidate frames from the candidate frames, and take the remaining candidate frames as valid frames to form a set of valid frames. If the number of frames in the set of valid frames does not meet the requirements, valid frames are extracted from the target video and added to the set of valid frames until the number of frames in the set of valid frames meets the requirements or it is determined that there are no video frames in the target video that can be used as valid frames.
2. The method according to claim 1, characterized in that, The frame extraction configuration parameters include distribution density parameters; The step of extracting video frames from the target video according to the total number of extracted frames as candidate frames includes: The frame index of the total number of frames is determined based on the distribution density parameter; the frame index conforms to a power distribution with the distribution density parameter as the exponent. The video frames corresponding to each frame index in the target video are used as candidate frames.
3. The method according to claim 2, characterized in that, The frame index satisfies: in, For the index of the i-th frame, For the time point corresponding to the index of the i-th frame, The frame rate of the target video. The function is the floor function; i = 1, 2, ..., N, where N is the total number of frames drawn; When the keyframe of the target video is located in the first or second half, ,or, ; When the keyframe of the target video is located in the middle position. ; in, Let α be the video duration of the target video, and α be the distribution density parameter. It is the inverse function of the standard normal distribution. To adjust the coefficient, and .
4. The method according to claim 2, characterized in that, The distribution density parameter is determined based on the following method: The target video is sparsely sampled to extract sample frames of a certain number; the number of sample frames is positively correlated with the video duration of the target video. For each sample frame, determine the motion intensity and information density corresponding to the sample frame; Based on the motion intensity and information density of each sample frame within the first time period, a first average motion intensity and a first average information density are determined; based on the motion intensity and information density of each sample frame within the second time period, a second average motion intensity and a second average information density are determined; the first time period is located in the first half of the target video, and the second time period is located in the second half of the target video. The motion intensity ratio is determined based on the first average motion intensity and the second average motion intensity; the information density ratio is determined based on the first average information density and the second average information density. The distribution density parameter is determined based on the motion intensity ratio and the information density ratio; the distribution density parameter is positively correlated with both the motion intensity ratio and the information density ratio.
5. The method according to claim 4, characterized in that, The step of determining the distribution density parameter based on the motion intensity ratio and the information density ratio includes: For any two adjacent sample frames, the scene switching time points in the target video where scene switching exists are determined based on the similarity between the two sample frames, and the distribution of scene switching is determined based on each scene switching time point. The basic distribution parameters are determined based on the motion intensity ratio and the information density ratio. The basic distribution parameters are corrected based on the distribution of scene switching to obtain the distribution density parameters; Wherein, if scene switching exceeding a preset proportion occurs in the first time period, the distribution density parameter is: ; If scene switching exceeding a preset proportion occurs during the second time period, the distribution density parameter is: The preset ratio is greater than 0.
5. In other cases, ; in, The distribution density parameter is... The basic distribution parameters are... The ratio of the stated exercise intensity The information density ratio; , All are preset coefficients, and , .
6. The method according to claim 1, characterized in that, The frame-skipping configuration parameters also include: a brightness threshold and a similarity threshold; The removal of black screen frames and duplicate frames that overlap with other candidate frames from the candidate frames includes: For each candidate frame, determine the average brightness value of the candidate frame; Candidate frames with an average brightness value greater than the brightness threshold are used as intermediate frames; wherein, black screen frames are candidate frames with an average brightness value less than the brightness threshold. For each intermediate frame, determine the perceptual hash value between the intermediate frame and the valid frames in the valid frame set; the valid frame set is initially empty. If the perceptual hash value between the intermediate frame and at least one valid frame in the set of valid frames is greater than the similarity threshold, then the intermediate frame is a duplicate frame and is removed. If the perceptual hash value between the intermediate frame and any valid frame in the set of valid frames is less than the similarity threshold, then the intermediate frame is considered a valid frame.
7. The method according to claim 1, characterized in that, The step of re-extracting valid frames from the target video and adding them to the valid frame set includes: Mark all video frames within a preset number of frames before and after the retained valid frames and / or the selected candidate frames as used; Frame extraction is performed on all video frames in the target video except those marked as used. The extracted video frames are subjected to black screen frame detection and duplicate frame detection. Video frames that are not black screen frames or duplicate frames are considered valid frames.
8. A video frame extraction device, characterized in that, The device includes: The acquisition module is used to acquire a frame extraction request for extracting frames from a target video; the frame extraction request includes the target video or the resource address of the target video; The parameter determination module is used to determine the frame extraction configuration parameters of the target video; the frame extraction configuration parameters include the total number of frames to be extracted; A frame extraction module is used to extract video frames from the target video as candidate frames based on the total number of extracted frames. The processing module is used to remove black screen frames and duplicate frames that are the same as other candidate frames from the candidate frames, and to take the remaining candidate frames as valid frames to form a set of valid frames; if the number of frames in the set of valid frames does not meet the requirements, valid frames are extracted from the target video and added to the set of valid frames until the number of frames in the set of valid frames meets the requirements or it is determined that there are no video frames in the target video that can be used as valid frames.
9. An electronic device, characterized in that, include: A memory and a processor are communicatively connected, the memory stores computer instructions, and the processor executes the video frame extraction method according to any one of claims 1 to 7 by executing the computer instructions.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the video frame extraction method according to any one of claims 1 to 7.