Video specific frame determination method and device, equipment, storage medium and program product

By determining the expression category of face images in the video frame sequence and determining the keyframe based on the expression change conditions, the problem of ignoring video content in the prior art is solved, and the video processing efficiency is improved.

CN120017835APending Publication Date: 2025-05-16BIGO TECH PTE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510164660.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art ignores the importance of video content when determining specific frames of video, resulting in poor selection accuracy and reducing the effectiveness of specific frames during video encoding and decoding.

Method used

By obtaining the video frame sequence, the expression category of the face image in each video image frame is determined, and the expression category sequence is generated. The video image frame that meets the expression change conditions is determined based on the expression category recorded in the expression category sequence, and it is determined as a video keyframe.

Benefits of technology

It improves the selection accuracy of specific frames, reduces the complexity of subsequent video encoding and decoding, and improves the overall video processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017835A_ABST
    Figure CN120017835A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video specific frame determination method and device, equipment, a storage medium and a program product, and the method comprises the steps: obtaining an input video frame sequence, the video frame sequence comprises a plurality of video image frames, each video image frame comprises a face image, determining the expression type of the face image in each video image frame, and determining the expression type of the face image in each video image frame; and generating an expression category sequence, determining video image frames meeting expression change conditions based on expression categories recorded in the expression category sequence, and determining the video image frames meeting the expression change conditions as video key frames. The key frame is determined according to the expression change condition of the video image frame, so that the technical problems of poor selection accuracy, low effectiveness of the specific frame in the video encoding and decoding process and the like caused by neglecting the importance of the video content when the specific frame is determined in the related technology are solved, the selection accuracy of the specific frame is improved, and the user experience is improved. The complexity of subsequent video coding and decoding is reduced, and the overall video processing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of video processing technology, and in particular, to a method, device, equipment, storage medium and program product for determining a specific frame of a video. Background Art

[0002] In the process of video processing, it is usually necessary to determine some specific frames for video encoding and decoding, such as key frames and reference frames in a video frame sequence. Among them, key frames are usually frames in a video frame sequence that record important changes or actions, while reference frames refer to frames used to predict other frames when encoding a video.

[0003] In the related art, taking key frames as an example, most of the determination methods adopt an equally spaced key frame selection strategy or a method of determining key frames after clustering based on histogram features. Among them, when using the equally spaced key frame strategy to determine key frames, it uses a fixed time interval as the basis for selecting key frames, ignoring the importance of video content, and in the method of using histograms as features to cluster and determine key frames, the histogram cannot effectively express the video content. The above methods of determining specific frames in the video all ignore the importance of the video content, and the accuracy of the selection is poor, which reduces the effectiveness of specific frames in the video encoding and decoding process. Summary of the invention

[0004] The embodiments of the present application provide a method, apparatus, device, storage medium and program product for determining specific frames of a video, which solves the technical problems that the related technology ignores the importance of video content when determining specific frames, resulting in poor selection accuracy and low effectiveness of specific frames in the video encoding and decoding process, improves the selection accuracy of specific frames, reduces the complexity of subsequent video encoding and decoding, and improves the overall video processing efficiency.

[0005] In a first aspect, an embodiment of the present application provides a method for determining a specific frame of a video, including:

[0006] Obtaining an input video frame sequence, wherein the video frame sequence includes a plurality of video image frames, each of the video image frames includes a face image;

[0007] Determining the expression category of the face image in each of the video image frames, and generating an expression category sequence;

[0008] The video image frames satisfying the expression change condition are determined based on the expression categories recorded in the expression category sequence, and the video image frames satisfying the expression change condition are determined as video key frames.

[0009] In a second aspect, an embodiment of the present application further provides a video specific frame determination device, including:

[0010] An acquisition module is configured to acquire an input video frame sequence, wherein the video frame sequence includes a plurality of video image frames, and each of the video image frames includes a face image;

[0011] An expression determination module, configured to determine the expression category of the face image in each of the video image frames and generate an expression category sequence;

[0012] The key frame determination module is configured to determine the video image frame that meets the expression change condition based on the expression category recorded in the expression category sequence, and determine the video image frame that meets the expression change condition as the video key frame.

[0013] In a third aspect, an embodiment of the present application further provides a video specific frame determination device, the device comprising:

[0014] memory and one or more processors;

[0015] The memory is used to store one or more programs;

[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the method for determining a specific video frame described in the embodiment of the present application.

[0017] In a fourth aspect, an embodiment of the present application further provides a non-volatile storage medium storing computer executable instructions, wherein the computer executable instructions, when executed by a computer processor, are used to execute the method for determining a specific video frame described in the embodiment of the present application.

[0018] In a fifth aspect, an embodiment of the present application further provides a computer program product, including a computer program, which, when executed by a processor, implements the method for determining a specific video frame described in an embodiment of the present application.

[0019] In an embodiment of the present application, after obtaining an input video frame sequence, the expression category of the face image in each video image frame is determined, and an expression category sequence is generated. Based on the expression category recorded in the expression category sequence, the video image frame that meets the expression change condition is determined, and the video image frame that meets the expression change condition is determined as a video key frame. The key frame is determined by the expression change of the video image frame. The selection method of the key frame is related to the video content itself, rather than selecting the key frame by a fixed time interval. This solves the technical problems in the related art of ignoring the importance of the video content when determining a specific frame, resulting in poor selection accuracy and low effectiveness of specific frames in the video encoding and decoding process. It improves the selection accuracy of specific frames, reduces the complexity of subsequent video encoding and decoding, and improves the overall video processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1A flowchart of a method for determining a specific video frame provided in an embodiment of the present application;

[0021] Figure 2 A flowchart of another method for determining a specific video frame provided in an embodiment of the present application;

[0022] Figure 3 A flowchart of another method for determining a specific video frame provided in an embodiment of the present application;

[0023] Figure 4 A flowchart of another method for determining a specific video frame provided in an embodiment of the present application;

[0024] Figure 5 A schematic diagram of marking feature points obtained after extracting facial features from a video frame image provided in an embodiment of the present application;

[0025] Figure 6 A flowchart of another method for determining a specific video frame provided in an embodiment of the present application;

[0026] Figure 7 A flowchart of another method for determining a specific video frame provided in an embodiment of the present application;

[0027] Figure 8 A block diagram of the module structure of a device for determining a specific video frame provided in an embodiment of the present application;

[0028] Fig. 9 A schematic diagram of the structure of a video specific frame determination device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0029] The embodiments of the present application are further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the embodiments of the present application, rather than to limit the embodiments of the present application. It should also be noted that, for ease of description, only parts related to the embodiments of the present application are shown in the accompanying drawings, rather than all structures.

[0030] The terms "first", "second", etc. in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and or or" in the specification and claims represents at least one of the connected objects, and the character "or" generally indicates that the objects associated before and after are in an "or" relationship.

[0031] The method for determining a specific video frame provided in the embodiment of the present application can be applied to scenarios containing facial images, such as live video broadcast, video call, and video conference. The aforementioned business scenarios are only exemplary and explanatory. In actual applications, it can also be applied to any other scenarios containing facial images that require video processing.

[0032] In the method for determining a specific video frame provided in the embodiment of the present application, the executor of each step may be a computer device, which refers to any electronic device with data calculation, processing and storage capabilities, such as mobile phones, PCs (Personal Computers), tablet computers and other terminal devices, and may also be servers and other devices, which are not limited in the embodiment of the present application.

[0033] Figure 1 A flowchart of a method for determining a specific video frame provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, specifically including:

[0034] Step S101: obtaining an input video frame sequence, wherein the video frame sequence includes a plurality of video image frames, each of the video image frames includes a face image.

[0035] Among them, the video frame sequence is a sequence containing multiple continuous video image frames. In one embodiment, each video image frame contains a face image. For example, in a live video broadcast, a live video screen containing the host's face will be generated. For the live video, each frame of the image, that is, each video image frame, contains the host's face. In one embodiment, when determining a specific video frame, it is determined in units of a video frame sequence, and the video frame sequence can be composed of each frame of the input video stream in sequence. The specific video frame is some special types of frames that need to be determined during the video processing process for video encoding and decoding, such as key frames and reference frames in a video frame sequence.

[0036] Step S102: determine the expression category of the face image in each video image frame, and generate an expression category sequence.

[0037] In one embodiment, for each video image frame of the input video frame sequence, the expression category of the face image in each video image frame can be determined in turn. The expression category is the classification of facial expressions displayed by the face, for example, angry, disgust, fear, happy, sad, surprise, neutral and other expression categories. The corresponding expression category can be determined by identifying the face image in each video image frame. Optionally, the face image can be subjected to expression recognition by a set expression category recognition algorithm to determine the expression category. After obtaining the expression category corresponding to each video image frame, the expression category sequence is recorded accordingly, and the expression category sequence records the expression category corresponding to each video image frame. For example, the expression category sequence is recorded as {S1, S2, S3, S4, ..., Sn-1, Sn}, where n is a natural number, S1 is the expression category corresponding to the first video image frame, S2 is the expression category corresponding to the second video image frame, and so on, and Sn is the expression category corresponding to the nth video image frame. Optionally, after determining the expression category of the face image in the current video image frame, the expression category is stored in a cache sequence, and an expression category sequence is generated when the expression category of the face image in each video image frame is stored in the cache sequence in sequence.

[0038] Step S103: determining a video image frame that meets an expression change condition based on the expression category recorded in the expression category sequence, and determining the video image frame that meets the expression change condition as a video key frame.

[0039] In one embodiment, taking the specific video frame to be determined as a video key frame as an example, it is determined by the change of the expression category recorded in the expression category sequence. Among them, the expression change condition is used to characterize the condition for judging whether the expression category of the face image in the current video image frame has changed. The video key frame can be determined using the expression change condition. The video key frame can be a frame in the video frame sequence that records important changes in facial expressions.

[0040] In one embodiment, a video image frame satisfies the expression change condition when the expression category of the current video image frame recorded in the expression category sequence is different from the expression categories of the previous a consecutive video image frames. At this time, it is determined that the expression category has changed, and it satisfies the expression change condition. Optionally, the value a can be set according to actual conditions, such as 3, 5 or 10. It should be noted that the above is an exemplary description of satisfying the expression change condition, and no specific limitation is made.

[0041] In one embodiment, after determining a video image frame that meets the expression change condition in a video frame sequence, the video image frame is determined as a video key frame, and the video key frame can be used for subsequent video encoding and decoding processing.

[0042] As can be seen from the above, an input video frame sequence is obtained, the video frame sequence includes multiple video image frames, each video image frame contains a face image, the expression category of the face image in each video image frame is determined, and an expression category sequence is generated. The video image frame that meets the expression change condition is determined based on the expression category recorded in the expression category sequence, and the video image frame that meets the expression change condition is determined as a video key frame. The present application determines the key frame by the expression change of the video image frame, which solves the technical problems in the related art that the importance of the video content is ignored when determining a specific frame, resulting in poor selection accuracy and low effectiveness of specific frames in the video encoding and decoding process, improves the selection accuracy of specific frames, reduces the complexity of subsequent video encoding and decoding, and improves the overall video processing efficiency.

[0043] Figure 2 A flowchart of another method for determining a specific video frame provided in an embodiment of the present application is provided, which provides a method for optionally determining a video image frame that satisfies an expression change condition, such as Figure 2 As shown, including:

[0044] Step S201: Obtain an input video frame sequence, where the video frame sequence includes a plurality of video image frames, each of which contains a face image.

[0045] Step S202: determine the expression category of the face image in each of the video image frames, and generate an expression category sequence.

[0046] Step S203, based on the expression categories recorded in the expression category sequence, determine a video image frame whose expression category of the recorded first video image frame is inconsistent with the expression categories of a first preset number of consecutive video image frames before the first video image frame at a ratio greater than a preset ratio, and determine the video image frame as a video key frame.

[0047] The first video image frame may be a video image frame in a video frame sequence. The first preset number is used to represent a value preset according to actual needs. The first preset number can be used to determine each video image frame for expression category comparison with the first video image frame. The preset ratio is used to represent a ratio value preset according to actual conditions. The preset ratio can be used to determine whether the first video image frame meets the expression change condition. Exemplarily, the first video image frame is the 20th video image frame in the video frame sequence, the first preset number is 10, the preset ratio is 70%, the expression category of the 20th video image frame recorded in the expression category sequence is angry, the expression categories of the 10 consecutive video image frames before the 20th video image frame are surprise, surprise, surprise, surprise, neutral, neutral, neutral, angry, angry, and the expression category of the 20th video image frame is inconsistent with the expression category of the 10 consecutive video image frames before it is 80%, which is greater than the preset ratio, then it is determined that the 20th video image frame meets the expression change condition, and the 20th video image frame is determined as a video key frame.

[0048] Optionally, when determining the video image frames that meet the expression change conditions in the video frame sequence, the current frames can be judged in sequence starting from the first frame of the video frame sequence according to the expression change conditions provided in the above embodiment, and the video image frames that are determined to meet the expression change conditions are determined as video key frames.

[0049] From the above, it can be seen that when setting the expression change conditions, the ratio of the inconsistency between the expression category of the first video image frame recorded in the expression category sequence and the expression categories of the first preset number of consecutive video image frames before the first video image frame is greater than the preset ratio is used as the judgment condition. This can more accurately and reasonably determine the video image frames with expression changes in the video frame sequence, further improve the determination accuracy of the video key frames, realize the association between the selection of video key frames and the video content, and ensure the efficiency of subsequent video processing.

[0050] Figure 3 A flowchart of another method for determining a specific video frame provided in an embodiment of the present application is provided, which provides a method for optionally determining a key video frame, such as Figure 3 As shown, including:

[0051] Step S301: obtaining an input video frame sequence, wherein the video frame sequence includes a plurality of video image frames, each of the video image frames includes a face image, determining the expression category of the face image in each of the video image frames, and generating an expression category sequence.

[0052] Step S302: determining a video image frame that meets an expression change condition based on the expression category recorded in the expression category sequence, and determining the video image frame that meets the expression change condition as a video key frame.

[0053] Step S303: if the number of video image frames that do not meet the expression change condition is determined to reach a set number based on the expression categories recorded in the expression category sequence, the video image frames that reach the set number are determined as video key frames.

[0054] Among them, the set number can be a value that is custom-set according to specific needs, and there is no unified standard value. In one embodiment, when determining a video key frame in a video frame sequence, if it is determined in sequence whether the video image frame satisfies the expression category change condition, if the video image frame that does not satisfy the expression change condition reaches a set number, then the video image frame that reaches the set number is correspondingly determined as a video key frame. This avoids the problem of reduced encoding efficiency caused by the absence of key frames in more frame images during subsequent encoding. Optionally, after the video key frame is determined in the above manner, when the next video image frame is determined to satisfy the expression change condition, the above-mentioned cumulative number is correspondingly cleared to zero for recounting.

[0055] In one embodiment, a counting mechanism can be implemented by setting a counter, that is, when it is determined that the current video image frame does not meet the expression change condition, the statistical value of the non-key frame counter is increased by one, and the current statistical value of the non-key frame counter is compared with the set number. When the current statistical value of the non-key frame counter reaches the set number, the current video image frame is determined as a video key frame, and the statistical value of the non-key frame counter is cleared to re-count the number of non-key frames. Among them, the non-key frame counter is a counter that counts the number of video image frames (i.e., non-key frames) that continuously do not meet the expression change condition. For example, when the set number is 15 and the current statistical value of the non-key frame counter is 15, it is determined that the current video image frame does not meet the expression change condition, that is, the video key frame is not determined in 15 consecutive frames of images, then the current video image frame, i.e., the 15th video image frame, is determined as a video key frame, and the statistical value of the non-key frame counter is cleared to re-count the number of non-key frames.

[0056] In another embodiment, when it is determined that the current video image frame does not meet the expression change condition, the current video image frame can also be stored in a corresponding storage list, and the number of video image frames in the storage list is counted, and the number is compared with a set number. When the number reaches the set number, the current video image frame is determined as a video key frame, and the video image frames in the storage list are cleared to re-count the number of non-key frames.

[0057] From the above, it can be seen that for a video frame sequence, when a longer video image frame is not determined as a video key frame, a forced method is used to directly determine the current frame as a key frame, which can avoid the problem of not selecting a key frame when the video image frame has no expression changes for a long time.

[0058] Figure 4 A flowchart of another method for determining a specific video frame provided in an embodiment of the present application is provided, which provides a method for optionally determining the expression category of a face image in a video image frame, such as Figure 4 As shown, including:

[0059] Step S401: Obtain an input video frame sequence, wherein the video frame sequence includes a plurality of video image frames, and each of the video image frames includes a face image.

[0060] Step S402: perform face detection and feature extraction on each of the video image frames to obtain a corresponding face feature map, wherein the face feature map includes feature points of different parts of the face extracted, and input the face feature map corresponding to each of the video image frames into an expression classification network model to obtain a corresponding expression category, and generate an expression category sequence.

[0061] In one embodiment, face detection is implemented using a lightweight convolutional neural network. Optionally, MobileNet is used as the skeleton network, a center point-based target prediction method is used, with modules such as SE and Hard-Swish Activation, and the loss function of the entire convolutional neural network consists of a heat map loss and a position coordinate offset (BoundingBox) loss function.

[0062] In one embodiment, feature extraction uses a deep learning network to detect key points of the face in the video frame image, and locates the key area positions of the face, including dozens to hundreds of key point information such as eyebrows, eyes, nose, mouth, and facial contours. In the process of feature extraction, face normalization processing is performed simultaneously, and a local constraint model is optionally used to solve the transformation parameters related to rotation, scaling, and translation between the detected face model and the standard model, and the optimal transformation parameters are obtained through iterative update of the parameters. Finally, the image is cropped and normalized to a uniform size to eliminate image inconsistencies caused by changes in face posture and angle.

[0063] Optionally, in the feature extraction process, a facial feature extraction algorithm with 68 feature points is used, which is implemented using a cascaded convolutional neural network, and the entire face area image is used as input. The convolutional neural network includes multiple stages, and the input of each stage is a corrected image, a key point heat map, and a feature map generated by a fully connected layer, and the output is the face shape. For example, Figure 5 As shown, Figure 5 This is a schematic diagram of marking feature points obtained after extracting facial features from a video frame image provided by an embodiment of the present application. In this figure, 68 feature points are used to mark different feature areas of the face.

[0064] Among them, the face feature map is a face image containing the feature points obtained by the above feature extraction, and the face feature map is input into the expression classification network model to obtain the corresponding expression category. Optionally, the expression classification network model uses a pre-trained EfficientNet network, which is divided into 9 stages in total. The first stage is a common convolution layer with a convolution kernel size of 3x3 and a step size of 2 (including a batch normalization layer and an activation function Swish), Stage2 to Stage8 are repeatedly stacked MBConv structures, and Stage9 consists of an ordinary 1x1 convolution layer (including a batch normalization layer and an activation function Swish), an average pooling layer, and a fully connected layer.

[0065] Step S403: determining a video image frame that meets an expression change condition based on the expression category recorded in the expression category sequence, and determining the video image frame that meets the expression change condition as a video key frame.

[0066] From the above, it can be seen that in the process of identifying expression categories, face detection and feature extraction are performed on each video image frame to obtain the corresponding face feature map. The face feature map includes the extracted feature points of different parts of the face. The face feature map corresponding to each video image frame is input into the expression classification network model to obtain the corresponding expression category. Through the specially set model structure and recognition method, the accuracy and robustness of expression category recognition of facial images are significantly improved.

[0067] In one embodiment, in the case where a specific frame includes a video reference frame, in order to solve the problem of large amount of calculation caused by determining the reference frame of a video image frame by motion search and estimation method for an input video frame sequence when determining the video reference frame in the related art, a flowchart of another method for determining a specific video frame is proposed, as shown in FIG. Figure 6 As shown, Figure 6 A flowchart of another method for determining a specific video frame provided in an embodiment of the present application, the method comprising:

[0068] Step S501: Obtain an input video frame sequence, wherein the video frame sequence includes a plurality of video image frames, and each of the video image frames includes a face image.

[0069] Step S502: perform face detection and feature extraction on each of the video image frames to obtain a corresponding face feature map, wherein the face feature map includes feature points of different parts of the face extracted, and input the face feature map corresponding to each of the video image frames into a face feature classifier to obtain corresponding face feature information, and determine a video reference frame based on the face feature information corresponding to each of the video image frames.

[0070] The facial feature classifier is used to represent the classification algorithm module for different facial features of the facial feature map. The facial feature classifier can be used to determine the facial feature information of the video image frame. The facial feature information can be related information of the facial features in the facial feature map. The video reference frame of the video frame sequence can be determined using the facial feature information corresponding to each video image frame. The video reference frame refers to a frame used to predict other frames when performing video encoding.

[0071] Optionally, a method for determining facial feature information may be to input the facial feature map corresponding to each video image frame into different sub-feature classifiers to obtain corresponding facial sub-features, wherein the facial sub-features include facial angle features and facial area features, and the facial area features include at least one of eye features, mouth features, nose features and eyebrow features. Correspondingly, a method for determining a video reference frame may be to determine a video reference frame based on the facial sub-features corresponding to each video image frame. Among them, the sub-feature classifier may be a classification algorithm module for features of different categories in the facial feature map, such as eye feature classifiers, mouth feature classifiers and other sub-feature classifiers. The facial angle feature is used to characterize the angle feature information of the face in the facial feature map, such as the rotation angle of the face about the x-axis, the rotation angle about the y-axis, the rotation angle about the z-axis or the three rotation angle feature information about the x, y and z axes. The eye feature is used to characterize the feature information of the eye in the facial feature map, such as the area of ​​the left and right eyes, the longitudinal length of the left and right eyes and other feature information. Mouth features are used to represent the feature information of the mouth in the face feature map, such as the area, size, thickness, etc. of the mouth. Nose features are used to represent the feature information of the nose in the face feature map, such as the width and area of ​​the nose. Eyebrow features are used to represent the feature information of the eyebrows in the face feature map, such as the width, length, area, etc. of the eyebrows.

[0072] In one embodiment, taking the sub-feature classifiers as face angle feature classifier, eye feature classifier and mouth feature classifier as examples, the face feature map corresponding to the current video image frame is input into the face angle feature classifier to obtain three rotation angles of the face about the x, y, and z axes, which are a, b, and c respectively, the face feature map corresponding to the current video image frame is input into the eye feature classifier to obtain the left eye area S1 and the right eye area S2, and the face feature map corresponding to the current video image frame is input into the mouth feature classifier to obtain the mouth area S3.

[0073] When performing area calculations, optionally, Figure 5 Taking the feature point diagram shown in the figure as an example, the state of the mouth (degree of opening) can be Figure 5 The area C61-C68 (the area formed by the open mouth) enclosed by the feature points 61-68 is calculated by dividing the square root by the distance L(61,65) between the feature points 61 and 65 (the width of the mouth). The calculation formula is as follows:

[0074]

[0075] Among them, M n Indicates the mouth feature information corresponding to the current frame n.

[0076] Similarly, the state of the left eye is calculated by dividing the square root of the area C43-C68 (the area of ​​the left eye) enclosed by the feature points 43-48 by the distance L(43,46) (the width of the left eye) between the feature points 43 and 46, as follows:

[0077]

[0078] in, Indicates the left eye feature information corresponding to the current frame n.

[0079] The state of the right eye is calculated by taking the square root of the area C37-C42 (the area formed by the left eye opening) enclosed by the feature points 37-42 and dividing it by the distance L(37,40) between the feature points 37 and 40 (the width of the left eye), as follows:

[0080]

[0081] in, Indicates the right eye feature information corresponding to the current frame n.

[0082] In one embodiment, after obtaining the facial feature information corresponding to each video image frame, the video reference frame can be determined based on the facial feature information. Optionally, for the current video image frame in the video frame sequence, when determining its reference frame, the video image frame with the facial feature information closest to the current video image frame can be determined as its video reference frame. Taking the facial feature information as the mouth feature as an example, the way to calculate the matching degree between two frames of images (such as the i-th frame image and the j-th frame image, i and j are natural numbers, i≠j) can be to subtract the mouth feature information Mi corresponding to the i-th frame image from the mouth feature information Mj corresponding to the j-th frame image and take the absolute value, and use the calculation result as the matching degree of the two frames of images. Optionally, when calculating the eye feature matching degree, the matching degrees of the left eye and the right eye can be calculated separately and then added to obtain the eye feature matching degree. The specific calculation method of the matching degree of the left eye and the right eye separately belongs to the calculation method of the mouth feature matching degree, that is, to subtract the feature information corresponding to the two frames of images and take the absolute value.

[0083] Step S503, input the facial feature map corresponding to each of the video image frames into the expression classification network model to obtain the corresponding expression category, and generate an expression category sequence, determine the video image frames that meet the expression change conditions based on the expression categories recorded in the expression category sequence, and determine the video image frames that meet the expression change conditions as video key frames.

[0084] It should be noted that in the process of determining specific frames of a video, taking key frames and reference frames as an example, the order of determination is not limited, and both can be determined at the same time or one can be determined first and then the other.

[0085] From the above, it can be seen that when determining the video reference frame, after performing face detection and feature extraction on each video image frame to obtain the corresponding face feature map, the face feature map corresponding to each video image frame is input into the face feature classifier to obtain the corresponding face feature information, and the video reference frame is determined based on the face feature information corresponding to each of the video image frames. This can solve the problem of large amount of calculation caused by determining the reference frame of the video image frame through motion search and estimation methods for the input video frame sequence when determining the video reference frame in the related technology. The present application can significantly improve the accuracy of the video reference frame.

[0086] Figure 7 A flowchart of another method for determining a specific video frame provided in an embodiment of the present application, which provides another method for optionally determining a reference video frame, such as Figure 7 As shown, including:

[0087] Step S601: Obtain an input video frame sequence, where the video frame sequence includes a plurality of video image frames, and each of the video image frames includes a face image.

[0088] Step S602: Perform face detection and feature extraction on each of the video image frames to obtain a corresponding face feature map, wherein the face feature map includes feature points of different parts of the face extracted.

[0089] Step S603, input the facial feature map corresponding to each of the video image frames into different sub-feature classifiers respectively to obtain corresponding facial sub-features, wherein the facial sub-features include facial angle features and facial area features, and the facial area features include at least one of eye features, mouth features, nose features and eyebrow features, and determine the video reference frame based on the facial sub-features corresponding to each of the video image frames.

[0090] In one embodiment, a method for optionally determining a video reference frame is provided. Taking the second video image frame of the current video reference frame to be determined in the video frame sequence as an example, a second preset number of video image frames with the closest facial angle features are screened out from the video image frames before the second video image frame, and a video image frame with the closest facial region features to the second video image frame is screened out from the second preset number of video image frames as the video reference frame. The second video image frame is used to represent the video image frame in the video frame sequence currently waiting to determine the video reference frame. The second preset number can be a value pre-set according to actual conditions. Optionally, in the case where the facial region features include at least two, the reference frames to be selected that are closest to each facial region feature of the second video image frame are screened out from the second preset number of video image frames, and the video reference frame is determined from the at least two reference frames to be selected. For the obtained at least two reference frames to be selected, the corresponding matching degrees recorded by each of them can be used to screen out the final video reference frame by multiplying the matching degree by the weight value of the corresponding facial region feature. For example, the weight of the eye feature is 2, and the weight of the mouth feature is 1. At this time, the matching degree of the candidate reference frame of the corresponding eye feature is 0.6, and the matching degree of the candidate reference frame of the mouth feature is 0.7. After multiplying their respective weights, the final score of the eye feature is 1.2, and the mouth feature is 0.7. At this time, the candidate reference frame corresponding to the eye feature is selected as the video reference frame.

[0091] In another embodiment, the method for determining the video reference frame can also be that, for the second video image frame that is currently to be determined as the video reference frame in the video frame sequence, the video image frame that is closest to the facial angle feature of the second video image frame is screened out from the video image frames before the second video image frame, and the video image frame that is closest to the facial area feature of the second video image frame is screened out from the video image frames before the second video image frame, and a random selection is made between the video image frame that is closest to the facial angle feature and the video image frame that is closest to the facial area feature to determine as the video reference frame.

[0092] Step S604, input the facial feature map corresponding to each of the video image frames into the expression classification network model to obtain the corresponding expression category, and generate an expression category sequence, determine the video image frames that meet the expression change conditions based on the expression categories recorded in the expression category sequence, and determine the video image frames that meet the expression change conditions as video key frames.

[0093] As can be seen from the above, the facial feature map corresponding to each video image frame is input into different sub-feature classifiers to obtain corresponding facial sub-features, which include facial angle features and facial area features. The facial area features include at least one of eye features, mouth features, nose features and eyebrow features. The video reference frame is determined based on the facial sub-features corresponding to each video image frame. The present application selects the video reference frame by the facial sub-features of the facial image in the video image frame, which can solve the problem of large amount of calculation generated when selecting the reference frame by motion search and other methods through multiple reference frame sequences, and improve the accuracy of the reference frame.

[0094] Figure 8 This is a module structure diagram of a video specific frame determination device provided in an embodiment of the present application. The device is used to execute a video specific frame determination method provided in the above embodiment, and has the corresponding functional modules and beneficial effects of the execution method. Figure 8 As shown, the device specifically includes:

[0095] An acquisition module 101 is configured to acquire an input video frame sequence, wherein the video frame sequence includes a plurality of video image frames, each of which includes a face image;

[0096] The expression determination module 102 is configured to determine the expression category of the face image in each of the video image frames and generate an expression category sequence;

[0097] The key frame determination module 103 is configured to determine the video image frames that meet the expression change condition based on the expression categories recorded in the expression category sequence, and determine the video image frames that meet the expression change condition as video key frames.

[0098] It can be seen from the above scheme that by obtaining an input video frame sequence, the video frame sequence includes multiple video image frames, each video image frame contains a face image, the expression category of the face image in each video image frame is determined, and an expression category sequence is generated. Based on the expression category recorded in the expression category sequence, the video image frame that meets the expression change condition is determined, and the video image frame that meets the expression change condition is determined as a video key frame. The present application determines the key frame by the expression change of the video image frame, which solves the technical problem that the related technology ignores the importance of the video content when determining a specific frame, resulting in poor selection accuracy and low effectiveness of specific frames in the video encoding and decoding process, improves the selection accuracy of specific frames, reduces the complexity of subsequent video encoding and decoding, and improves the overall video processing efficiency.

[0099] In a possible embodiment, the key frame determination module 103 is configured as follows:

[0100] The proportion of inconsistency between the expression category of the first video image frame recorded in the expression category sequence and the expression categories of a first preset number of consecutive video image frames before the first video image frame is greater than a preset proportion.

[0101] In a possible embodiment, the key frame determination module 103 is further configured to:

[0102] If the number of video image frames that do not meet the expression change condition is determined to reach a set number in sequence based on the expression categories recorded in the expression category sequence, the video image frames that reach the set number are determined as video key frames.

[0103] In a possible embodiment, the expression determination module 102 is configured as follows:

[0104] Performing face detection and feature extraction on each of the video image frames to obtain a corresponding face feature map, wherein the face feature map includes feature points of different parts of the face extracted;

[0105] The facial feature map corresponding to each video image frame is input into the expression classification network model to obtain the corresponding expression category.

[0106] In a possible embodiment, it further includes a reference frame determination module configured to:

[0107] Inputting the facial feature map corresponding to each of the video image frames into a facial feature classifier to obtain corresponding facial feature information;

[0108] A video reference frame is determined based on the facial feature information corresponding to each of the video image frames.

[0109] In a possible embodiment, the reference frame determination module is further configured to:

[0110] Inputting the facial feature map corresponding to each of the video image frames into different sub-feature classifiers to obtain corresponding facial sub-features, wherein the facial sub-features include facial angle features and facial region features, and the facial region features include at least one of eye features, mouth features, nose features, and eyebrow features;

[0111] A video reference frame is determined based on the facial sub-features corresponding to each of the video image frames.

[0112] In a possible embodiment, the reference frame determination module is further configured to:

[0113] For a second video image frame in the video frame sequence, the video image frame before the second video image frame is selected to obtain a second preset number of video image frames that are closest to the facial angle features.

[0114] A video image frame closest to the facial region feature of the second video image frame is selected from the second preset number of video image frames as a video reference frame.

[0115] In a possible embodiment, the reference frame determination module is further configured to:

[0116] In the case where the facial region features include at least two, selecting reference frames to be selected that are closest to each facial region feature of the second video image frame from the second preset number of video image frames;

[0117] A video reference frame is determined from among the at least two reference frames to be selected.

[0118] Fig. 9 A schematic diagram of the structure of a video specific frame determination device provided in an embodiment of the present application, such as Fig. 9 As shown, the device includes a processor 201, a memory 202, an input device 203 and an output device 204; the number of processors 201 in the device can be one or more. Fig. 9 A processor 201 is taken as an example; the processor 201, memory 202, input device 203 and output device 204 in the device can be connected by a bus or other means. Fig. 9The example of connecting through a bus is taken. The memory 202, as a computer-readable storage medium, can be used to store software programs, computer executable programs and modules, such as program instructions or modules corresponding to a method for determining a specific video frame in an embodiment of the present application. The processor 201 executes various functional applications and data processing of the device by running the software programs, instructions and modules stored in the memory 202, that is, implements the above-mentioned method for determining a specific video frame. The input device 203 can be used to receive input digital or character information, and generate key signal input related to user settings and function controls of the device. The output device 204 may include a display device such as a display screen.

[0119] The embodiment of the present application further provides a non-volatile storage medium containing computer executable instructions, wherein the computer executable instructions are used to perform a method for determining a specific video frame when executed by a computer processor, the method comprising:

[0120] Obtaining an input video frame sequence, wherein the video frame sequence includes a plurality of video image frames, each of the video image frames includes a face image;

[0121] Determining the expression category of the face image in each of the video image frames, and generating an expression category sequence;

[0122] The video image frames satisfying the expression change condition are determined based on the expression categories recorded in the expression category sequence, and the video image frames satisfying the expression change condition are determined as video key frames.

[0123] The embodiment of the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, a method for determining a specific video frame is implemented. The method includes:

[0124] Obtaining an input video frame sequence, wherein the video frame sequence includes a plurality of video image frames, each of the video image frames includes a face image;

[0125] Determining the expression category of the face image in each of the video image frames, and generating an expression category sequence;

[0126] The video image frames satisfying the expression change condition are determined based on the expression categories recorded in the expression category sequence, and the video image frames satisfying the expression change condition are determined as video key frames.

[0127] It is worth noting that in the embodiment of the above-mentioned video specific frame determination method system, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present application.

[0128] In some possible implementations, various aspects of the method provided in this application may also be implemented in the form of a program product, which includes a program code. When the program product is run on a computer device, the program code is configured to enable the computer device to execute the steps of the method according to various exemplary embodiments of the present application described above in this specification. For example, the computer device may execute the application publishing method recorded in the embodiment of this application. The program product may be implemented in any combination of one or more readable media.

Claims

1. A method for determining a specific frame of a video, characterized in that: include: Obtaining an input video frame sequence, wherein the video frame sequence includes a plurality of video image frames, each of the video image frames includes a face image; Determining the expression category of the face image in each of the video image frames, and generating an expression category sequence; The video image frames satisfying the expression change condition are determined based on the expression categories recorded in the expression category sequence, and the video image frames satisfying the expression change condition are determined as video key frames.

2. The method for determining a specific video frame according to claim 1, wherein: The expression change condition is satisfied, including: The proportion of inconsistency between the expression category of the first video image frame recorded in the expression category sequence and the expression categories of a first preset number of consecutive video image frames before the first video image frame is greater than a preset proportion.

3. The method for determining a specific video frame according to claim 1, wherein: After generating the expression category sequence, the method further includes: If the number of video image frames that do not meet the expression change condition is determined to reach a set number in sequence based on the expression categories recorded in the expression category sequence, the video image frames that reach the set number are determined as video key frames.

4. The method for determining a specific video frame according to any one of claims 1 to 3, characterized in that: The step of determining the expression category of the facial image in each of the video image frames comprises: Performing face detection and feature extraction on each of the video image frames to obtain a corresponding face feature map, wherein the face feature map includes feature points of different parts of the face extracted; The facial feature map corresponding to each video image frame is input into the expression classification network model to obtain the corresponding expression category.

5. The method for determining a specific video frame according to claim 4, characterized in that: After performing face detection and feature extraction on each of the video image frames to obtain a corresponding face feature map, the method further includes: Inputting the facial feature map corresponding to each of the video image frames into a facial feature classifier to obtain corresponding facial feature information; A video reference frame is determined based on the facial feature information corresponding to each of the video image frames.

6. The method for determining a specific video frame according to claim 5, characterized in that: The step of inputting the facial feature map corresponding to each of the video image frames into a facial feature classifier to obtain corresponding facial feature information includes: Inputting the facial feature map corresponding to each of the video image frames into different sub-feature classifiers to obtain corresponding facial sub-features, wherein the facial sub-features include facial angle features and facial region features, and the facial region features include at least one of eye features, mouth features, nose features, and eyebrow features; Accordingly, determining the video reference frame based on the facial feature information corresponding to each of the video image frames includes: A video reference frame is determined based on the facial sub-features corresponding to each of the video image frames.

7. The method for determining a specific video frame according to claim 6, wherein: The determining of the video reference frame based on the facial sub-features corresponding to each of the video image frames comprises: For a second video image frame in the video frame sequence, the video image frame before the second video image frame is selected to obtain a second preset number of video image frames that are closest to the facial angle features. A video image frame closest to the facial region feature of the second video image frame is selected from the second preset number of video image frames as a video reference frame.

8. The method for determining a specific video frame according to claim 7, wherein: The step of selecting a video image frame closest to the face region feature of the second video image frame from the second preset number of video image frames as a video reference frame includes: In the case where the facial region features include at least two, selecting reference frames to be selected that are closest to each facial region feature of the second video image frame from the second preset number of video image frames; A video reference frame is determined from among the at least two reference frames to be selected.

9. A device for determining a specific frame of a video, characterized in that: include: An acquisition module is configured to acquire an input video frame sequence, wherein the video frame sequence includes a plurality of video image frames, and each of the video image frames includes a face image; An expression determination module, configured to determine the expression category of the face image in each of the video image frames and generate an expression category sequence; The key frame determination module is configured to determine the video image frame that meets the expression change condition based on the expression category recorded in the expression category sequence, and determine the video image frame that meets the expression change condition as the video key frame.

10. A device for determining a specific frame of a video, characterized in that: include: memory and one or more processors; The memory is used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method for determining a specific video frame as described in any one of claims 1 to 8.

11. A non-volatile storage medium storing computer executable instructions, characterized in that: When the computer executable instructions are executed by a computer processor, they are used to perform the method for determining a specific video frame according to any one of claims 1 to 8.

12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for determining a specific video frame according to any one of claims 1 to 8 is implemented.