A method, apparatus and electronic device for extracting high-brightness video
By extracting and matching the target avatar and body image matching the preset highlight state in the video and performing video scoring methods, the problem of inefficient automatic extraction of highlight videos in the prior art is solved, and efficient and automatic highlight video generation and quality assurance are achieved.
Patent Information
- Application Number
- CN202111600465.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-12-24
AI Technical Summary
The prior art is inefficient in automatically extracting live highlight videos from longer videos, requiring a lot of manpower to perform manual editing.
By extracting the target avatar and body images in each frame of the video that match the preset highlight avatar status and actions, and matching and tracking, highlight videos are generated. The specific steps include using the Yolov5 model to identify avatar and body images, combining ResNext, CSPNet and SENet neural network models to identify avatar status and actions, perform similarity comparison, and perform video scoring based on the highlight attribute degree, pixel change degree, avatar completeness, displacement degree and picture quality.
The process of automatically extracting highlight videos in long videos is realized, editing efficiency is improved, manpower investment is reduced, and the quality of highlight videos is ensured through multi-dimensional scoring.
Smart Images

Figure CN114399704B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video processing, and particularly to a method, device and electronic equipment for extracting highlight videos of living bodies. Background Art
[0002] When shooting videos of living bodies such as people or animals on a stage, in a movie, or in a documentary, relatively long videos are obtained. Usually, for purposes such as publicity, it is necessary to clip highlight videos of representative actions and representative expressions of the living bodies from the relatively long videos. The existing technologies for such videos often adopt the method of manually watching the videos and then manually clipping the highlight moments of the living bodies in the videos to obtain the highlight videos. However, this method has low clipping efficiency and consumes a large amount of manpower. Therefore, how to automatically extract the highlight videos of living bodies from long videos is an urgent problem to be solved. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a method, device and electronic equipment for extracting highlight videos, thereby realizing the automatic extraction of highlight videos of living bodies from long videos.
[0004] According to a first aspect, the present invention provides a method for extracting highlight videos, the method comprising: extracting a target head and a target body image that match a preset highlight head state and a preset highlight action from each frame of a sample video, the preset highlight head state being used to characterize an expression, a head orientation, and a head pitch state that meet the highlight criteria; matching the target head and the target body image extracted from each frame of the video based on whether they belong to the same living body, and obtaining a target living body image that is successfully matched in each frame of the video; tracking the target living body in each frame of the sample video based on the target living body image, and extracting the frames where the target living body exists to form a highlight video.
[0005] Optionally, the extracting a head and a body image that match a preset highlight expression and a preset highlight action from each frame of the sample video comprises: identifying the head and the body image in each frame of the video based on the Yolov5 model for reducing convolutional operations; respectively identifying the head state and the action of the head and the body image based on a preset neural network model, the preset neural network model being generated by training a network structure composed of ResNext, CSPNet, and SENet through head samples and body image samples; comparing the identified head state and action with the preset highlight head state and the preset highlight action respectively in terms of similarity; and extracting a target head and a target body image from the head and the body image, the similarity between the head state of the target head and the preset highlight head state being higher than a first preset threshold, and the similarity between the action of the target body image and the preset highlight action being higher than a second preset threshold.
[0006] Optionally, the steps of generating the Yolov5 model with reduced convolution operations include: obtaining the number n of convolution operations in the bottleneck layer of the backbone network of the initial Yolov5 model; replacing the n convolution operations with a superposition operation of m convolution operations and s linear transformations to generate the Yolov5 model with reduced convolution operations, where n = m × s.
[0007] Optionally, the steps of generating the Yolov5 model with reduced convolution operations further include: respectively replacing the upsampling convolution operations and downsampling convolution operations in the feature pyramid network and the pixel aggregation network of the NECK structure in the Yolov5 model with bilinear interpolation operations to generate the Yolov5 model with reduced convolution operations.
[0008] Optionally, the method further includes: scoring the highlight video based on the highlight attribute degree, pixel change degree, head portrait integrity, displacement degree of the target live body in the highlight video, and the picture quality of the highlight video, where the highlight attribute degree is used to characterize the matching degree between the head portrait and body image of the target live body and the preset highlight head portrait state and preset highlight action respectively; when the score of the highlight video is greater than a preset threshold, retaining the highlight video, otherwise deleting the highlight video.
[0009] Optionally, the scoring the highlight video based on the highlight attribute degree, pixel change degree, head portrait integrity, displacement degree of the target live body in the highlight video, and the picture quality of the highlight video includes: obtaining all the head portrait states and all the actions of the target live body in the highlight video, and taking the mean of the first similarity and the second similarity as the highlight attribute degree, where the first similarity is the mean of the similarity calculations between all the head portrait states and the preset highlight head portrait state respectively, and the second similarity is the mean of the similarity calculations between all the actions and the preset highlight action respectively; taking the mean of the change amounts between the pixels of each frame in the highlight video and the average pixels of all the frames as the pixel change degree; performing identity recognition on the head portraits of each frame of the target live body in the highlight video, and taking the identity recognition success rate as the head portrait integrity, where the identity recognition success rate is the ratio of the number of head portraits whose identities can be recognized to the total number of head portraits; taking the mean of the change values of the head portrait center points of each frame of the target live body in the highlight video as the displacement degree, where the change value of the head portrait center point is the change amount between the head portrait center point of the target live body in each frame and the average head portrait center point of all the frames; characterizing the picture quality based on the weighted calculation result of the picture clip amount, aspect ratio, brightness, and clarity of the highlight video; performing a weighted calculation on the highlight attribute degree, pixel change degree, head portrait integrity, displacement degree, and the picture quality to obtain the score of the highlight video.
[0010] Optionally, before retaining the highlight video when the score of the highlight video is greater than a preset threshold and deleting the highlight video otherwise, the method further includes: adjusting the score of the highlight video by subtracting points based on the out-of-frame rate of the target live body in the highlight video, where the out-of-frame rate is the ratio of the average size of the target live body out of the video frame in the highlight video to the size of the video frame.
[0011] According to a second aspect, the present invention provides a highlight video extraction device, including: a structured detection module, configured to extract a target head image and a target body image that match a preset highlight head state and a preset highlight action from each frame of a sample video, where the preset highlight head state is used to define an expression, a head orientation, and a head pitch state that meet the highlight criteria; a target matching module, configured to match the target head image and the target body image extracted from each frame of the video based on whether they belong to the same live body, and obtain a target live body image that is successfully matched in each frame of the video; a video generation module, configured to track the target live body in each frame of the sample video based on the target live body image, and extract the frames with the target live body to form a highlight video.
[0012] According to a third aspect, an embodiment of the present invention provides an electronic device, including: a memory and a processor, which are communicatively connected to each other, where the memory stores computer instructions, and the processor executes the computer instructions to execute the method according to the first aspect or any optional implementation manner of the first aspect.
[0013] According to a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, which stores computer instructions for causing the computer to execute the method according to the first aspect or any optional implementation manner of the first aspect.
[0014] The technical solution provided by this application has the following advantages:
[0015] The technical solution provided by the present application first analyzes each frame of the sample video separately, performs target recognition on the head portrait and body image of the living body in each frame, and then determines whether the head portrait and body image recognized in each frame match the preset highlight state, and retains the target head portrait and body image that meet the highlight state. After that, the target head portrait and target body image extracted from the picture are matched to form a complete living target, so as to determine the identity information of the living body that needs to be tracked, and finally the target living body is tracked in each frame, the pictures containing the target living body are retained, and the pictures not including the target living body are discarded, and finally the remaining pictures are combined into a video, so as to obtain a complete video for the target living body, and the video contains the highlight state of the target living body, which is a highlight video of the target living body.
[0016] In addition, the present invention also uses a lightweight Yolov5 model to identify the head and body images in each frame, which greatly improves the target recognition speed. In addition, the highlight video is scored based on the highlight attribute degree, pixel change degree, head image integrity, displacement degree and image quality of the target living body, further ensuring the quality of the highlight video. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The features and advantages of the present invention will be more clearly understood by referring to the accompanying drawings, which are schematic and should not be construed as limiting the present invention in any way. In the accompanying drawings:
[0018] Figure 1 A schematic diagram showing the steps of a highlight video extraction method in one embodiment of the present invention is shown;
[0019] Figure 2 A schematic diagram of quantization of a head image and a body image of a highlight video extraction method in one embodiment of the present invention is shown;
[0020] Figure 3 A schematic diagram of an operation method for generating a lightweight yolov5 model in one embodiment of the present invention is shown;
[0021] Figure 4 The structure diagram of FPN+PAN in the prior art is shown;
[0022] Figure 5 A schematic structural diagram of a highlight video extraction device in one embodiment of the present invention is shown;
[0023] Figure 6 A schematic structural diagram of an electronic device in one embodiment of the present invention is shown. DETAILED DESCRIPTION
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0025] Please refer to Figure 1 , in one embodiment, a method for extracting high-light videos specifically includes the following steps:
[0026] Step S101: Extract the target head and target body images in each frame of the sample video that match the preset high-light head state and preset high-light action. The preset high-light head state is used to represent the expression, head orientation, and head pitch state that meet the high-light standard.
[0027] Step S102: Match the target head and target body images extracted from each frame of the video based on whether they belong to the same living body, and obtain the target living body images that match successfully in each frame of the video.
[0028] Step S103: Track the target living body in each frame of the sample video based on the target living body image, and extract the frames with the target living body to form a high-light video.
[0029] Specifically, in order to clip a highlight video of a live subject's representative actions and representative expressions from a long video, first, the single-frame images of a video are processed to separate the frame images of the video. Then, based on the object recognition algorithm, the head and body images of the live subject in the images are recognized, and algorithms including but not limited to Yolov1, Yolov2, Yolov3, etc. can be used for recognition. After that, the similarity between the recognized head and body images and the preset highlight head states and preset highlight actions is calculated. The methods for similarity calculation include but are not limited to inner product and Euclidean distance. If the similarity is above the preset threshold, it is considered that the currently matched head or body image has a representative expression or action. In the embodiment of the present invention, in order to improve the accuracy of live subject highlight state recognition, the similarity matching of the head and body images is performed respectively from the aspects of the expression of the head, the head orientation, the head pitch state, and multiple action dimensions. Since the tracking of the highlight video is the tracking of a complete target, and the target head and target body images with representative expressions and representative actions recognized in each frame image are usually not unique, it is also necessary to match multiple target heads and target body images, match the target head and target body images belonging to the same live subject together, and then search for the complete target image in the live subject identity information database to confirm the identity of the target to be tracked. In this embodiment, it is determined whether the face and the body match by calculating whether the center coordinates of the face are within the body coordinates and calculating the maximum iou value between the face and the body. Finally, the live subject is tracked in each frame image, the images containing the live subject are retained, and the images not including the target live subject are discarded. Finally, the retained images are combined into a video, and the complete video of the target live subject can be obtained, and the highlight state of the target live subject is included in the video, which is the highlight video of the target live subject.
[0030] In this embodiment, the tracking algorithm for tracking multiple targets in each frame is implemented using the SORT multi-object tracking algorithm. For example, the classification and positions of all potential high-brightness live targets in the first frame are obtained, and a unique ID is assigned to each target. A Kalman filter tracker is initialized for each target to predict the position of each target in the next frame. The target detection model is used for target detection in the second frame to obtain the classification and positions of all targets in the second frame. The IoU between each pair of the M targets in the first frame and the N targets in the second frame is calculated to establish a cost matrix. The Hungarian matching algorithm is used to obtain the unique match with the largest IoU, and then the matching pairs with a matching value less than the IoU threshold are removed. The positions of the targets matched in the second frame are used to update the Kalman tracker, and the Kalman gain, state estimate value, and estimated error covariance at the second frame are calculated, and the state estimate value is output for calculating the predicted position in the next frame. For the targets that are not matched, the Kalman filter tracker is re-initialized. Similar processing is performed for each subsequent frame of the image according to the methods of the first frame and the second frame, so as to achieve the tracking of the targets in each frame of the image. This is only an example, and the tracking algorithm adopted is not limited to this.
[0031] Specifically, in one embodiment, the above step S101 specifically includes the following steps:
[0032] Step 1: Identify the head and body images in each frame based on the Yolov5 model that reduces convolutional operations. Specifically, in the embodiment of the present invention, the head and body images are identified based on the Yolov5 target recognition algorithm, which performs well in target recognition in the prior art, to improve the recognition efficiency. Moreover, by reducing the convolutional operations in the Yolov5 model, the algorithm complexity is further reduced, and the calculation speed is increased. It still has a relatively fast speed for long videos.
[0033] Step 2: Based on a preset neural network model, respectively identify the head states and actions of the head and body images. The preset neural network model is generated by training a network structure composed of ResNext, CSPNet, and SENet using head samples and body image samples.
[0034] Specifically, after identifying the avatars and human body images in each frame of the video, in order to improve the accuracy of calculating the similarity of the highlight state, it is first necessary to identify the specific states of the avatars and body images in each video. In this embodiment, it specifically includes the expression, head orientation, and head pitch state of the avatar, as well as the standing body size and movement, standing with an object in hand, sitting body size and movement, and sitting with an object in hand. The expressions include happy, calm, sad, and expressionless. The avatar orientations include front, back, left, and right. In this embodiment, the above-mentioned quantified avatar states and body image states are realized by a network structure composed of ResNext, CSPNet, and SENet for machine recognition of avatars and body images. ResNext, CSPNet, and SENet are convolutional neural networks that have emerged in recent years, and their analysis and classification of images are more accurate than traditional networks such as CNN. Before training the above network structure, analyze multiple attributes of the images in the sample set to generate a one-dimensional vector with binary labels, thereby completing the quantification of avatars and human body images, as Figure 2 shown, where each attribute corresponds to a different label. Then, input the labeled image sample set into the neural network to generate a one-dimensional label probability vector for each attribute; according to the one-dimensional vector with binary labels and the one-dimensional label probability vector, calculate the loss value of the label, and update the parameter weights of the neural network according to the partial derivative of the loss value with respect to the network parameters. When the number of updates is greater than the first preset number of times, a trained child highlight attribute recognition model is obtained. Finally, use the child highlight attribute recognition model to analyze the avatar and body images, generate multi-label probability values, combine the one-dimensional label probability vectors of the human head and body, generate a probability value matrix corresponding to multiple attributes and multiple labels, and determine the specific state of each avatar or body image based on the values in the probability matrix. Through the above steps, the specific states of the avatar and body image in multiple dimensional attributes are identified, providing a reliable sample for the subsequent discrimination of whether the avatar and body image have highlight attributes, and further improving the discrimination accuracy. In this embodiment, the formula for calculating the loss value of the label is as follows:
[0035]
[0036]
[0037] where ∈ is a constant, which is taken as 0.1 during the implementation process, N represents the number of all categories after multi-label fusion, i is the correct classification, j is the wrong classification, ce represents the original cross-entropy calculation formula, K is the number of categories, y k represents the true label of the kth category, and p k is the predicted probability of the network.
[0038] Step 3: Compare the recognized avatar status and actions with the preset high-light avatar status and preset high-light actions respectively for similarity.
[0039] Step 4: Extract the target avatar and target body image from the avatar and body images, where the similarity between the avatar status of the target avatar and the preset high-light avatar status is higher than the first preset threshold, and the similarity between the action of the target body image and the preset high-light action is higher than the second preset threshold.
[0040] Specifically, after obtaining the specific status of each avatar and body image through the above Step 1 and Step 2, compare them with the preset high-light avatar status and preset high-light actions for similarity. For example, the preset high-light avatar status includes: smiling, laughing, facing forward, looking up, and the preset high-light actions include hands open, standing, etc. Calculate the similarity between each avatar status and body image status and the above status. Assuming that the similarity results of the current avatar status and body image status comparison are respectively greater than their preset thresholds, it indicates that the current avatar status and body image status have high-light attributes and are thus used as the target avatar and target body image for subsequent analysis.
[0041] Specifically, in one embodiment, the specific steps of generating the Yolov5 model with reduced convolution operations in the above Step 1 include:
[0042] Step 5: Obtain the number n of convolution operations in the bottleneck layer of the backbone network of the initial Yolov5 model.
[0043] Step 6: Replace the n convolution operations with the superposition operation of m convolution operations and s linear transformations to generate the Yolov5 model with reduced convolution operations, where n = m × s.
[0044] Specifically, Yolov5 is an object recognition model improved based on Yolov4, with very fast object recognition speed. Its network structure includes three parts: backbone network, NECK, and prediction network. Its backbone network is a composite network composed of CBL structure and multi-level Resnet structure, and the CBL structure is composed of a convolutional layer (Conv), a bottleneck layer (Bottleneck, BN), and an activation function Leakyrelu. In the bottleneck layer, there are also a large number of convolutional operations. When processing images, there will be a large number of redundant feature pairs in the middle feature extraction part of the backbone network. For example, some similar feature maps often require expensive convolutional operations to obtain, and this part of the operations is mainly concentrated in the bottleneck layer. Therefore, in this embodiment, for these similar feature maps, consider using some simple operations to replace complex convolutional operations, which can not only reduce the convolutional kernels used to generate intermediate feature maps, reduce the computational amount, so as to achieve the purpose of optimizing the number of parameters, but also does not affect the effect of feature extraction. In principle, the improvement of reducing the computational amount cannot affect the number of feature maps to accurately connect to subsequent network processing. In the traditional way, n feature maps can be obtained through n convolutional operations. Therefore, in this embodiment, first obtain the number of convolutional operations n in the bottleneck layer, as Figure 3 shown, replace the n convolutional operations with a superposition operation of first performing m convolutional operations and then performing s linear transformation operations (such as addition, subtraction, translation operations), so that m×s = n, thus achieving the effect of greatly reducing the computational amount while keeping the output quantity unchanged. For example: after the input image undergoes one convolutional operation, m original feature maps are obtained where h′ and w′ are the length and width of the input image respectively, and the operation of any convolutional layer for generating n feature maps can be expressed as:
[0045] Y′ = X*f′ + b′
[0046] where * is the convolutional operation, b is the bias term, is the convolutional kernel used, c is the kernel size of the convolutional kernel f′, k is the number of channels, and m is the number of convolutional kernels in each channel. Apply a series of lightweight linear operations to each original feature in Y′ to generate s similar feature maps:
[0047]
[0048] where y′ i is the i-th original feature map in Y′, Φ ij is the j-th linear operation for generating the j-th similar feature map y ij , that is, y′ i can have one or more similar feature maps By using lightweight convolutional operations, we can obtain n = m×s feature maps Y = [y11 , y 12 , …, y ms as the output data of the lightweight convolution module.
[0049] Specifically, in one embodiment, in the above step one, the specific steps of generating the Yolov5 model with reduced convolution operations further include:
[0050] Step seven: Replace the upsampling convolution operation and the downsampling convolution operation in the feature pyramid network and the pixel aggregation network of the NECK structure in the Yolov5 model respectively with bilinear interpolation operations to generate the Yolov5 model with reduced convolution operations.
[0051] Specifically, in the NECK structure part of Yolov5, there are two parts: the feature pyramid network (Feature Pyramid Network, FPN) and the pixel aggregation network (Pixel Aggregation Network, PAN). Its main purpose is to deepen the image features for three different size specifications when Yolov5 identifies three targets of different sizes. In the process of the convolutional neural network, the deeper the network layer, the stronger the target feature information, and the better the model's prediction of the target. However, at the same time, the position information of the target will become weaker and weaker, and in the continuous convolution process, the information of small targets is easily lost. As Figure 4 shown, the structure of FPN + PAN is that in the PAN path, the image is convolved multiple times to obtain better target feature information, and in the FPN layer path, the convolution of the bottom layer is upsampled multiple times to expand the image pixels and obtain better position information. Then, the images with the same size specifications in the FPN path and the PAN path are added horizontally to obtain a feature map with relatively strong position information and feature information for each size specification, so as to more accurately identify targets of different sizes. In this process, the computational complexity of the downsampling convolution operation in the PAN path and the upsampling convolution operation in the FPN path is relatively large, which affects the target recognition efficiency. Therefore, in this embodiment, the computational complexity of the FPN + PAN structure is reduced by replacing the upsampling convolution operation and the downsampling convolution operation with bilinear interpolation operations, thereby further lightweighting the Yolov5 model. Based on bilinear interpolation for image scaling, the quality of the scaled image is high because the correlation influence of the four direct neighboring points around the sampling point to the sampling point is considered, basically overcoming the disadvantage of discontinuous nearest neighbor interpolation. Although the accuracy is slightly decreased compared with the convolution operation, the Yolov5 model has undergone a large number of convolutional processes in the backbone network, and the overall effect is not significant, but the computational speed of the Yolov5 model is greatly improved. The specific operation process of bilinear interpolation is prior art and will not be elaborated here.
[0052] Specifically, in one embodiment, a method for extracting high-light videos further includes the following steps:
[0053] Step Eight: Score the high-light video based on the high-light attribute degree, pixel change degree, head integrity, displacement degree of the target live body in the high-light video, and the picture quality of the high-light video.
[0054] Step Nine: When the score of the high-light video is greater than the preset threshold, retain the high-light video; otherwise, delete the high-light video.
[0055] Specifically, in this embodiment, the high-light video is scored from multiple dimensions including the high-light attribute degree, pixel change degree, head integrity, displacement degree of the target live body in the high-light video, and the picture quality of the high-light video. Only when the score of the high-light video is higher than the preset threshold, the current high-light video is considered a high-light video that meets the requirements and is retained. This further improves the accuracy and reliability of the high-light video. In this embodiment, the specific operations for the above scoring include the following steps:
[0056] 1. Obtain all the head states and all the actions of the target live body in the high-light video, and take the average of the first similarity and the second similarity as the high-light attribute degree. The first similarity is the average of the similarity calculations between all the head states and the preset high-light head states respectively, and the second similarity is the average of the similarity calculations between all the actions and the preset high-light actions respectively. Among them, for the specific operations of similarity calculation, existing technologies including but not limited to Euclidean distance, inner product, etc. can be used, which will not be elaborated here. The calculation of head states and body movement degrees is the preset high-light head states and preset high-light actions such as smiling, laughing, facing forward, looking up, hands open, standing, etc. described in the above Step Four. When the high-light attribute degree is larger, it indicates that the target live body in the video shows more representative expressions or actions, and thus the video effect is better.
[0057] 2. Take the average value of the change amount between the pixels of each frame in the high-light video and the average pixels of all the frames as the pixel change degree. Specifically, first calculate the average pixels and pixel variance of all the frames of the high-light video, and then perform the following operations on each frame:
[0058] Single-frame pixel change = (single-frame pixel - average pixel) / pixel variance
[0059] After that, sum and then average the single-frame pixel changes of each frame, which is the pixel change degree. When the pixel change degree is larger, it means that the actions of the live body in the video change more frequently, and thus the video effect is better.
[0060] 3. Identify the identity of the head in each frame of the high-light video of the target living body, and use the identity recognition success rate as the head integrity. The identity recognition success rate is the ratio of the number of heads whose identities can be recognized to the total number of heads. Specifically, in this embodiment, based on the identity integrity of the target living body picture in each frame as the head integrity, perform image search by image in the preset living body image database. Taking the target living body image as a human as an example, when the face is a frontal face or not blocked, the probability of matching a registered photo with an identity label from the database is relatively high. When the proportion is larger, it means that the proportion of the frontal face without occlusion in the picture is larger, so the visual effect of this video is better.
[0061] 4. Use the mean value of the head center point change values in each frame of the high-light video of the target living body as the displacement degree. The head center point change value is the change amount between the head center point of the target living body in each frame and the average head center point of all the frames. Specifically, first obtain the head center point of the target living body in each frame, then calculate the mean value of the above-mentioned multiple head center points to obtain the average center point, and then calculate the center point variance of the multiple head center points. Then perform the following operations on each frame:
[0062] Single-frame center point change value = (single-frame head center point - average center point) / center point variance
[0063] After that, sum up and then calculate the mean value of the single-frame center point change values of each frame, which is the displacement degree. When the displacement degree is larger, it indicates that the action changes of the living body in the video are more frequent, so the video effect is better.
[0064] 5. Characterize the picture quality based on the weighted calculation results of the picture clip amount, aspect ratio, brightness, and clarity of the high-light video. Specifically, in this embodiment, calculate the picture quality through the following formula: Picture quality = f1 * (cropped length / original length + cropped width / original width) + f2 * (aspect ratio) + f3 * brightness + f4 * clarity, where f1, f2, f3, and f4 are empirical weights. Through this step, when the cropping amount of the video is smaller, the aspect ratio is larger, and the brightness and clarity are higher, the corresponding picture gives a better visual experience, and at the same time, the picture quality is higher.
[0065] 6. Calculate the weighted sum of the high - light attribute degree, pixel change degree, avatar integrity, displacement degree, and picture quality to obtain the score of the high - light video. Specifically, calculate the final score of the high - light video through weighted calculation of all the above scores, so as to realize the multi - dimensional evaluation of the high - light video and further improve the reliability and accuracy of the high - light video. For example: High - light video score = N1 * High - light attribute degree+N2 * Pixel change degree+N3 * Avatar integrity+N4 * Displacement degree+N5 * Picture quality, where the specific values of the weights N1 - N5 corresponding to the high - light attribute degree, pixel change degree, avatar integrity, displacement degree, and picture quality are usually determined according to expert experience and will not be elaborated here.
[0066] Specifically, in one embodiment, a method for extracting a high - light video further includes the following steps:
[0067] Step ten: Adjust the score of the high - light video by subtracting points based on the out - of - frame rate of the target living body in the high - light video. The out - of - frame rate is the ratio of the average size of the target living body out of the picture to the size of the picture in the high - light video.
[0068] Specifically, the target living body may be out of the frame in each frame of the picture. For example, only half of the target living body is in the picture, and being out of the frame will seriously affect the visual effect of the high - light video. Therefore, in this embodiment, adjust the score of the high - light video obtained in steps eight to nine by subtracting points based on the out - of - frame rate of the target living body, so as to further ensure the reliability and accuracy of the generation of the high - light video. In this embodiment, the out - of - frame rate = out - of - frame size / actual size. When the out - of - frame rate is greater than 50%, subtract 40 points from the comprehensive score; when the out - of - frame rate is greater than 0 and less than 50%, calculate according to the following formula:
[0069] Deducted score = 10 2x ×4,
[0070] where x is the out - of - frame rate. When x = 10%, about 6.3 points are deducted; when x = 20%, about 10.0 points are deducted; when x = 30%, about 15.9 points are deducted; when x = 40%, about 25.2 points are deducted. When the out - of - frame rate is equal to 0, no point - deduction operation is performed.
[0071] Through the above steps, the technical solution provided by this application first analyzes each frame of the sample video separately, performs object recognition on the living head and living body parts in each frame, then determines whether the recognized head and body images in each frame match the preset highlight state, and retains the head and body images that meet the highlight state. Then, the target head and target body image extracted from the picture are matched to form a complete living target, so as to determine the identity information of the living body to be tracked. Finally, the target living body is tracked in each frame of the picture, the pictures containing the target living body are retained, and the pictures not including the target living body are discarded. Finally, the retained pictures are combined into a video to obtain a complete video of the target living body, and this video contains the highlight state of the target living body, which is the highlight video of the target living body.
[0072] In addition, the present invention also uses a lightweight Yolov5 model to identify the head and body images in each frame, which greatly improves the object recognition speed. And based on the highlight attribute degree, pixel change degree, head integrity, displacement degree of the target living body and the picture quality of the highlight video, a score is given to the highlight video, which further ensures the quality of the highlight video.
[0073] As Figure 5 shown, this embodiment also provides a highlight video extraction device, which includes:
[0074] A structured detection module 101, which is used to extract the target head and target body images that match the preset highlight head state and preset highlight action in each frame of the sample video. The preset highlight head state is used to define the expression, head orientation and head pitch state that meet the highlight standard. For the detailed content, refer to the relevant description of step S101 in the above method embodiment, and will not be elaborated here.
[0075] A target matching module 102, which is used to match the target head and target body images extracted from each frame according to whether they belong to the same living body, and obtain the target living body images that are successfully matched in each frame. For the detailed content, refer to the relevant description of step S102 in the above method embodiment, and will not be elaborated here.
[0076] A video generation module 103, which is used to track the target living body in each frame of the sample video based on the target living body image, and extract the pictures with the target living body to form a highlight video. For the detailed content, refer to the relevant description of step S103 in the above method embodiment, and will not be elaborated here.
[0077] A highlight video extraction device provided by an embodiment of the present invention is used to execute a highlight video extraction method provided by the above embodiment, and its implementation manner and principle are the same. For the detailed content, refer to the relevant description of the above method embodiment, and will not be elaborated here.
[0078] Through the collaborative cooperation of the above-mentioned various components, the technical solution provided by this application first analyzes each frame of the sample video separately, performs target recognition on the living head and living body parts in each frame of the video, then determines whether the recognized head and body images in each frame match the preset highlight state, and retains the head and body images that meet the highlight state. After that, the target head and target body images extracted from the video are matched to form a complete living target, so as to determine the identity information of the living body to be tracked. Finally, the target living body is tracked in each frame of the video, the frames containing the target living body are retained, and the frames not including the target living body are discarded. Finally, the retained frames are combined into a video, and a complete video of the target living body can be obtained. Moreover, this video contains the highlight state of the target living body and is the highlight video of the target living body.
[0079] In addition, the present invention also uses a lightweight Yolov5 model to identify the head and body images in each frame of the video, which greatly improves the target recognition speed. Moreover, based on the highlight attribute degree, pixel change degree, head integrity, displacement degree of the target living body, and the picture quality of the highlight video, a score is given to the highlight video, which further ensures the quality of the highlight video.
[0080] Figure 6 An electronic device according to an embodiment of the present invention is shown. The device includes a processor 901 and a memory 902, which can be connected through a bus or other means. Figure 6 Taking the connection through the bus as an example.
[0081] The processor 901 can be a central processing unit (CPU). The processor 901 can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or combinations of the above types of chips.
[0082] The memory 902, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the methods in the above method embodiments. The processor 901 executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory 902, that is, implements the methods in the above method embodiments.
[0083] The memory 902 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created by the processor 901 and the like. In addition, the memory 902 may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 902 may optionally include a memory remotely provided with respect to the processor 901, and these remote memories can be connected to the processor 901 through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0084] One or more modules are stored in the memory 902 and, when executed by the processor 901, execute the methods in the above method embodiments.
[0085] For the specific details of the above electronic device, reference can be made to the corresponding relevant descriptions and effects in the above method embodiments for understanding, and details will not be described herein again.
[0086] Those skilled in the art can understand that to implement all or part of the processes in the above method embodiments, it can be completed by instructing relevant hardware through a computer program. The implemented program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the above method embodiments. Among them, the storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above types of memories.
[0087] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for extracting high - light videos, characterized in that, the method includes: extracting target head images and target body images that match a preset high - light head state and a preset high - light action from each frame of the sample video, where the preset high - light head state is used to characterize expressions, head orientations, and head pitch states that meet the high - light criteria; matching the target head images and target body images extracted from each frame of the video based on whether they belong to the same living body, and obtaining target living - body images that match successfully in each frame of the video; the matching of the target head images and target body images extracted from each frame of the video based on whether they belong to the same living body includes: determining whether the face and the body match by calculating whether the central coordinates of the face are within the body coordinates and calculating the maximum IoU value between the face and the body; tracking the target living body in each frame of the sample video based on the target living - body images, and extracting the frames with the target living body to form a high - light video.
2. The method according to claim 1, characterized in that, the extracting of the head images and body images that match a preset high - light expression and a preset high - light action from each frame of the sample video includes: identifying the head images and body images in each frame based on the Yolov5 model that reduces convolutional operations; respectively identifying the head state and action of the head image and the body image based on a preset neural network model, where the preset neural network model is generated by training a network structure composed of ResNext, CSPNet, and SENet with head samples and body - image samples; performing a similarity comparison between the identified head state and action and the preset high - light head state and preset high - light action respectively; extracting target head images and target body images from the head images and the body images, where the similarity between the head state of the target head image and the preset high - light head state is higher than a first preset threshold, and the similarity between the action of the target body image and the preset high - light action is higher than a second preset threshold.
3. The method according to claim 2, characterized in that, the steps of generating the Yolov5 model that reduces convolutional operations include: obtaining the number n of convolutional operations in the bottleneck layer of the backbone network of the initial Yolov5 model; replacing the n convolutional operations with a superposition operation of m convolutional operations and s linear transformations to generate the Yolov5 model that reduces convolutional operations, where n = m×s.
4. The method according to claim 3, characterized in that, the steps of generating the Yolov5 model that reduces convolutional operations further include: respectively replacing the up - sampling convolutional operations and down - sampling convolutional operations in the feature pyramid network and pixel aggregation network of the NECK structure in the Yolov5 model with bilinear interpolation operations to generate the Yolov5 model that reduces convolutional operations.
5. The method according to claim 1, characterized in that, the method further includes: Score the highlight video based on the highlight attribute degree, pixel change degree, head integrity, displacement degree of the target living body in the highlight video, and the picture quality of the highlight video, where the highlight attribute degree is used to characterize the matching degrees of the target living body's head image and body image with the preset highlight head state and preset highlight action respectively; When the score of the highlight video is greater than a preset threshold, retain the highlight video; otherwise, delete the highlight video.
6. The method according to claim 5, wherein, the scoring the highlight video based on the highlight attribute degree, pixel change degree, head integrity, displacement degree of the target living body in the highlight video, and the picture quality of the highlight video includes: Obtain all the head states and all the actions of the target living body in the highlight video, and use the average of the first similarity and the second similarity as the highlight attribute degree, where the first similarity is the average of the similarity calculations of all the head states with the preset highlight head state respectively, and the second similarity is the average of the similarity calculations of all the actions with the preset highlight action respectively; Use the average of the change amounts between the pixels of each frame in the highlight video and the average pixels of all the frames as the pixel change degree; Perform identity recognition on the heads of each frame in the highlight video of the target living body, and use the identity recognition success rate as the head integrity, where the identity recognition success rate is the ratio of the number of heads whose identities can be recognized to the total number of heads; Use the average of the change values of the head center points of each frame in the highlight video of the target living body as the displacement degree, where the change value of the head center point is the change amount between the head center point of the target living body in each frame and the average head center point of all the frames; Characterize the picture quality based on the weighted calculation result of the picture clip amount, aspect ratio, brightness, and clarity of the highlight video; Perform a weighted calculation on the highlight attribute degree, pixel change degree, head integrity, displacement degree, and the picture quality to obtain the score of the highlight video.
7. The method according to claim 6, wherein, before the when the score of the highlight video is greater than a preset threshold, retain the highlight video; otherwise, delete the highlight video, the method further includes: Perform a score deduction adjustment on the score of the highlight video based on the out-of-frame rate of the target living body in the highlight video, where the out-of-frame rate is the ratio of the average out-of-frame size of the target living body in the highlight video to the frame size.
8. A highlight video extraction device, wherein, the device includes: A structured detection module, configured to extract the target head image and target body image that match the preset highlight head state and preset highlight action from each frame of the sample video, where the preset highlight head state is used to define the expression, head orientation, and head pitch state that meet the highlight standard; A target matching module, configured to match the target head images and target body images extracted from the respective frame images based on whether they belong to the same living body, and obtain the target living body images that are successfully matched in each frame image; the matching of the target head images and target body images extracted from the respective frame images based on whether they belong to the same living body includes: determining whether the face and the body are matched by calculating whether the central coordinates of the face are within the body coordinates and calculating the maximum IoU value between the face and the body; A video generation module, configured to track the target living body in each frame image of the sample video based on the target living body image, and extract the frames with the target living body to form a highlight video.
9. An electronic device, characterized in that, it includes: a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.
Citation Information
Patent Citations
Human body target optimization method and device based on multiple evaluation indexes and storage medium
CN110765913A
Athlete style recognition system and method
US20200394413A1
Method and system of clipping a video, computing device, and computer storage medium
US20210125639A1