Video processing method and device, electronic equipment and storage medium

CN120264073APending Publication Date: 2025-07-04BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510477747.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, manual video ratio adjustment is inefficient and cannot meet the batch processing needs, resulting in an imbalance in the video footage proportion and affecting the viewing experience.

Method used

By obtaining the video frames to be processed, detecting the human body's key point information, calculating the video proportional correction parameters, and performing proportional correction processing, including video content feature analysis, key point detection, proportional correction parameter calculation and video enhancement optimization.

Benefits of technology

It realizes efficient and accurate automatic adjustment of video proportions, which is suitable for batch processing of massive content on the video platform, significantly improving the video viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120264073A_ABST
    Figure CN120264073A_ABST
Patent Text Reader

Abstract

The invention provides a video processing method and device, electronic equipment and a storage medium. The method comprises the steps of obtaining a to-be-processed video; extracting a plurality of to-be-processed video frames from the to-be-processed video; for each to-be-processed video frame, detecting first human body key point information in the to-be-processed video frame, and determining a first key part proportion according to the first human body key point information; determining a video proportion correction parameter according to a first key part proportion and a standard human body proportion corresponding to each to-be-processed video frame; and performing proportion correction processing on the to-be-processed video according to the video proportion correction parameter to obtain a target video. Through the method and the device, the video proportion can be efficiently, accurately and automatically adjusted, the method and the device are suitable for batch processing of mass contents of a video platform, the processing efficiency is greatly improved while the correction precision is ensured, and the video watching experience is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of video processing, and in particular, to a video processing method, apparatus, electronic device, and storage medium. Background Art

[0002] With the rapid development of applications such as short video platforms, online education, and video surveillance, the production and dissemination of video content have shown explosive growth. However, due to factors such as shooting equipment, encoding formats, and uploading platforms, the problem of video aspect ratio imbalance has become increasingly prominent. For example, the videos uploaded by users may be incorrectly stretched or compressed, resulting in distorted characters and warped images, seriously affecting the viewing experience. In professional fields such as film and television production, advertising placement, and online courses, maintaining the correct video aspect ratio is particularly important. Currently, the commonly used method for adjusting the video aspect ratio is the manual adjustment method, specifically: adjusting the aspect ratio of the image frame by frame through professional video editing software (such as Adobe Premiere, Final Cut Pro).

[0003] However, this manual adjustment method is inefficient and cannot meet the requirements of batch processing. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a video processing method, apparatus, electronic device, and storage medium to solve the problem that the manual adjustment method is inefficient and cannot meet the requirements of batch processing. The specific technical solutions are as follows:

[0005] In a first aspect, this application provides a video processing method, including:

[0006] Obtain a video to be processed;

[0007] Extract a plurality of video frames to be processed from the video to be processed;

[0008] For each video frame to be processed, detect the first human body key point information in the video frame to be processed, and determine the first key part ratio according to the first human body key point information;

[0009] Determine a video aspect ratio correction parameter according to the first key part ratio corresponding to each video frame to be processed and the standard human body ratio;

[0010] Perform aspect ratio correction processing on the video to be processed according to the video aspect ratio correction parameter to obtain a target video.

[0011] In a possible implementation manner, the extracting a plurality of video frames to be processed from the video to be processed includes:

[0012] Analyze the motion features and scene change features of the video content in the video to be processed. The motion features are used to characterize the intensity distribution of pixel displacement between video frames, and the scene change features are used to characterize the spatio-temporal distribution characteristics of video shot transitions;

[0013] Determine a sampling strategy according to the motion features and scene change features. The sampling strategy is used to characterize the mapping relationship between the frame sampling frequency and the change of video content;

[0014] Extract multiple video frames to be processed from the video to be processed according to the sampling strategy.

[0015] In a possible implementation manner, the determining the sampling strategy according to the motion features and scene change features includes:

[0016] When the intensity of pixel displacement between frames represented by the motion features of the video content is greater than the motion intensity threshold, determine that the sampling frequency corresponding to the video content is the first sampling frequency, and the first sampling frequency is less than the reference sampling frequency;

[0017] When the shot transition frequency represented by the scene change features of the video content is greater than the scene transition threshold, determine that the sampling frequency corresponding to the video content is the second sampling frequency, and the second sampling frequency is greater than the reference sampling frequency;

[0018] When the intensity of pixel displacement between frames represented by the motion features of the video content is less than or equal to the motion intensity threshold, and the shot transition frequency represented by the scene change features of the video content is less than or equal to the scene transition threshold, determine that the sampling frequency corresponding to the video content is the reference sampling frequency;

[0019] Generate the sampling strategy according to the sampling frequencies corresponding to all video content.

[0020] In a possible implementation manner, the detecting the first human key point information in the video frame to be processed includes:

[0021] Perform a preprocessing operation on the video frame to be processed;

[0022] Use a pre-trained key point detection model to identify key points in the video frame to be processed after the preprocessing operation, and obtain human key point coordinate data;

[0023] Verify the credibility of the human key point coordinate data according to the topological constraint relationship between key points;

[0024] Store the verified human key point coordinate data in time series to obtain the first human key point information.

[0025] In a possible implementation, determining the video ratio correction parameter according to the first key part ratio corresponding to each of the to-be-processed video frames and the standard human body ratio includes:

[0026] Performing statistical analysis on the first key part ratio corresponding to each of the to-be-processed video frames to obtain a statistical result;

[0027] Removing outliers from the statistical result to obtain valid data;

[0028] Calculating the reference human body ratio according to the valid data;

[0029] Comparing the reference human body ratio with the standard human body ratio to obtain the correction parameter.

[0030] In a possible implementation, performing ratio correction processing on the to-be-processed video according to the video ratio correction parameter to obtain a target video includes:

[0031] Calculating the target video size according to the video ratio correction parameter;

[0032] Performing resolution resampling on the to-be-processed video according to the target video size to obtain a first video that conforms to the target video size;

[0033] Performing edge processing on the first video to obtain a second video;

[0034] Performing quality optimization on the second video through a video enhancement algorithm to obtain the target video.

[0035] In a possible implementation, the method further includes:

[0036] Extracting verification frames from the target video;

[0037] Detecting second human body key point information in the verification frames, and determining a second key part ratio according to the second human body key point information;

[0038] Calculating the matching degree between the second key part ratio and the standard human body ratio;

[0039] In the case where the matching degree is lower than a preset threshold, recalculating the video ratio correction parameter.

[0040] In a second aspect, the present application provides a video processing device, including:

[0041] An acquisition module, configured to acquire a to-be-processed video;

[0042] An extraction module, configured to extract a plurality of to-be-processed video frames from the to-be-processed video;

[0043] A detection module, configured to detect first human key-point information in each video frame to be processed, and determine a first key part ratio according to the first human key-point information;

[0044] A determination module, configured to determine a video ratio correction parameter according to the first key part ratio corresponding to each video frame to be processed and a standard human body ratio;

[0045] A correction module, configured to perform ratio correction processing on the video to be processed according to the video ratio correction parameter to obtain a target video.

[0046] In a possible implementation manner, the extraction module is specifically configured to:

[0047] Analyze the motion feature and scene change feature of the video content in the video to be processed, where the motion feature is used to characterize the intensity distribution of pixel displacement between video frames, and the scene change feature is used to characterize the spatio-temporal distribution characteristic of video shot switching;

[0048] Determine a sampling strategy according to the motion feature and scene change feature, where the sampling strategy is used to characterize the mapping relationship between the frame sampling frequency and the video content change;

[0049] Extract a plurality of video frames to be processed from the video to be processed according to the sampling strategy.

[0050] In a possible implementation manner, the extraction module is further configured to:

[0051] When the intensity of pixel displacement between frames represented by the motion feature of the video content is greater than a motion intensity threshold, determine that the sampling frequency corresponding to the video content is a first sampling frequency, and the first sampling frequency is less than a reference sampling frequency;

[0052] When the shot switching frequency represented by the scene change feature of the video content is greater than a scene switching threshold, determine that the sampling frequency corresponding to the video content is a second sampling frequency, and the second sampling frequency is greater than the reference sampling frequency;

[0053] When the intensity of pixel displacement between frames represented by the motion feature of the video content is less than or equal to the motion intensity threshold, and the shot switching frequency represented by the scene change feature of the video content is less than or equal to the scene switching threshold, determine that the sampling frequency corresponding to the video content is the reference sampling frequency;

[0054] Generate the sampling strategy according to the sampling frequencies corresponding to all video contents.

[0055] In a possible implementation manner, the detection module is specifically configured to:

[0056] Perform preprocessing operations on the video frame to be processed;

[0057] Use a pre-trained key point detection model to identify key points in the preprocessed video frame to be processed, and obtain human key point coordinate data;

[0058] Verify the credibility of the human key point coordinate data according to the topological constraint relationship between key points;

[0059] Store the verified human key point coordinate data in time series to obtain the first human key point information.

[0060] In a possible implementation manner, the determining module is specifically configured to:

[0061] Perform statistical analysis on the first key part ratios corresponding to each of the video frames to be processed to obtain a statistical result;

[0062] Remove outliers from the statistical result to obtain valid data;

[0063] Calculate a reference human body ratio according to the valid data;

[0064] Compare the reference human body ratio with the standard human body ratio to obtain the correction parameter.

[0065] In a possible implementation manner, the correction module is specifically configured to:

[0066] Calculate a target video size according to the video ratio correction parameter;

[0067] Perform resolution resampling on the video to be processed according to the target video size to obtain a first video that meets the target video size;

[0068] Perform edge processing on the first video to obtain a second video;

[0069] Optimize the quality of the second video through a video enhancement algorithm to obtain a target video.

[0070] In a possible implementation manner, the device further includes a verification module, configured to:

[0071] Extract verification frames from the target video;

[0072] Detect second human key point information in the verification frames, and determine a second key part ratio according to the second human key point information;

[0073] Calculate the matching degree between the second key part ratio and the standard human body ratio;

[0074] In the case where the matching degree is lower than a preset threshold, recalculate the video ratio correction parameter.

[0075] In a third aspect, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;

[0076] The memory is used to store a computer program;

[0077] The processor, when executing the program stored on the memory, implements the method steps described in any one of the first aspects.

[0078] In a fourth aspect, a computer-readable storage medium is provided, characterized in that a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method steps described in any one of the first aspects are implemented.

[0079] In a fifth aspect, a computer program product containing instructions is provided, which when running on a computer causes the computer to execute the video processing method described above.

[0080] Beneficial effects of the embodiments of the present application:

[0081] The embodiments of the present application provide a video processing method, device, electronic device, and storage medium. In the embodiments of the present application, first, obtain a video to be processed, and extract a plurality of video frames to be processed from the video to be processed. Then, for each video frame to be processed, detect the first human body key point information in the video frame to be processed, and determine the first key part ratio according to the first human body key point information. Furthermore, according to the first key part ratio corresponding to each video frame to be processed and the standard human body ratio, determine the video ratio correction parameter. Finally, perform ratio correction processing on the video to be processed according to the video ratio correction parameter to obtain the target video. Through the present application, the video ratio can be automatically adjusted efficiently and accurately, which is applicable to batch processing of a large amount of content on video platforms. While ensuring the correction accuracy, the processing efficiency is greatly improved, and the video viewing experience is significantly improved.

[0082] Of course, it is not necessary for any product or method implementing the present application to achieve all the above advantages simultaneously. Description of the Drawings

[0083] The drawings here are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.

[0084] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0085] One or more embodiments are exemplarily illustrated by the pictures in the corresponding accompanying drawings. These exemplary illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, unless otherwise stated, and the drawings in the figures do not constitute a scale limitation.

[0086] Figure 1 It is a flowchart of a video processing method provided by an embodiment of the present application;

[0087] Figure 2 It is a flowchart of another video processing method provided by an embodiment of the present application;

[0088] Figure 3 It is a flowchart of yet another video processing method provided by an embodiment of the present application;

[0089] Figure 4 It is a schematic structural diagram of a video processing device provided by an embodiment of the present application;

[0090] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0091] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.

[0092] The following disclosure provides many different embodiments or examples for implementing different structures of the present invention. To simplify the disclosure of the present invention, the components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present invention. In addition, the present invention may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed.

[0093] Figure 1The flowchart of a video processing method provided by an embodiment of the present application. This method can be applied to one or more electronic devices such as smart phones, laptops, desktop computers, portable computers, servers, etc. In addition, the execution subject of this method can be hardware or software. When the above execution subject is hardware, the execution subject can be one or more of the above electronic devices. For example, a single electronic device can execute this method, or multiple electronic devices can cooperate with each other to execute this method. When the above execution subject is software, this method can be implemented as multiple software or software modules, or can be implemented as a single software or software module. No specific limitation is made here.

[0094] As Figure 1 shown, the method specifically includes:

[0095] S101. Obtain a video to be processed.

[0096] The video to be processed refers to the original video data that needs to be proportionally corrected, including but not limited to video files in packaging formats such as MP4 and AVI.

[0097] In an embodiment of the present application, the original video stream is received through a video input interface, and then a decoder such as FFmpeg is used to decode the video stream into a sequence of frames that can be processed, that is, the video to be processed.

[0098] S102. Extract multiple video frames to be processed from the video to be processed.

[0099] The video frame to be processed is used to represent a static image sample extracted from the video stream and serves as the input data for key point detection.

[0100] In one embodiment, video frames can be extracted from different time points of the video to be processed through a random sampling strategy.

[0101] In another embodiment, first, the total number of video frames (F_total = T × FPS) is calculated according to the total video duration (T) and the original frame rate (FPS), and then the reference sampling interval (Δ = floor(F_total / N)) is determined in combination with the preset target sampling number (N). The uniformity of frame extraction is ensured through equal-interval sampling, and the first-frame and last-frame must-be-sampled rules are set to ensure content integrity. In this way, the risk of content omission in random sampling is avoided.

[0102] S103. For each video frame to be processed, detect the first human key point information in the video frame to be processed, and determine the first key part ratio according to the first human key point information.

[0103] The first human body key point information refers to the coordinates of human body key points. For example, there are a total of 17 human anatomical feature points such as the head, shoulders, elbows, waist, knees, feet, etc.

[0104] The first key part ratio refers to the pixel distance ratio between key points (such as the head-shoulder ratio, trunk-lower limb ratio, etc.).

[0105] In the embodiments of the present application, first, each video frame to be processed is respectively input into a deep learning model (such as OpenPose, HRNet, etc.), and the coordinates of human body key points in each video frame are detected by the model. Then, key part ratios such as the head-shoulder ratio are calculated according to the coordinates of human body key points.

[0106] S104. Determine the video ratio correction parameter according to the first key part ratio corresponding to each of the video frames to be processed and the standard human body ratio.

[0107] The standard human body ratio refers to a database of normal human body ratios established based on statistical learning.

[0108] The video ratio correction parameter refers to a geometric transformation parameter including width / height scaling factors. Specifically, it includes: width scaling factor (Sx): the adjustment ratio in the horizontal direction (for example, 0.9 means the width is compressed by 10%); height scaling factor (Sy): the adjustment ratio in the vertical direction (for example, 1.1 means the height is stretched by 10%).

[0109] In the embodiments of the present application, S104 may specifically include the following steps:

[0110] Step A1. Perform statistical analysis on the first key part ratios corresponding to each of the video frames to be processed to obtain a statistical result;

[0111] Step A2. Remove the outliers in the statistical result to obtain valid data;

[0112] Step A3. Calculate the reference human body ratio according to the valid data;

[0113] Step A4. Compare the reference human body ratio with the standard human body ratio to obtain the correction parameter.

[0114] In this solution, first, statistical analysis is performed on the ratios of human key body parts (such as head-shoulder ratio, torso-limb ratio, etc.) detected in multiple video frames to be processed to obtain statistical results, such as mean, median, and standard deviation analysis results. Then, statistical methods such as IQR (interquartile range) are used to eliminate abnormal detection results that deviate from the normal range (such as distorted data caused by occlusion or detection errors) to ensure the reliability of the subsequent calculation data. Furthermore, for the filtered valid data, the reference human body ratio is calculated through weighted average and robust regression algorithms, where the weights of each data point are dynamically adjusted according to the confidence of key points and the frame clarity (that is, the higher the confidence of key points, the greater the weight of the frame data; the higher the frame clarity, the greater the weight of the frame data), and finally, a reference human body ratio with statistical significance is obtained; finally, the reference human body ratio is compared with the standard human body ratio database statistically obtained from a large number of normal videos, and the optimal aspect ratio correction parameters are calculated through optimization algorithms such as the least squares method.

[0115] This technical solution realizes the determination of correction parameters through a four-step processing flow of "multi-frame statistics - anomaly filtering - reference calculation - standard comparison", improves the processing efficiency while ensuring the correction accuracy, and is particularly suitable for the batch processing requirements of a large amount of content on video platforms.

[0116] S105. Perform ratio correction processing on the video to be processed according to the video ratio correction parameters to obtain a target video.

[0117] Correction processing refers to the geometric transformation process performed on the video.

[0118] The target video refers to the corrected output video stream.

[0119] In the embodiments of the present application, S105 may specifically include the following steps:

[0120] Step B1. Calculate the target video size according to the video ratio correction parameters;

[0121] Step B2. Perform resolution resampling on the video to be processed according to the target video size to obtain a first video that meets the target video size;

[0122] Step B3. Perform edge processing on the first video to obtain a second video;

[0123] Step B4. Optimize the quality of the second video through a video enhancement algorithm to obtain a target video.

[0124] The target video size refers to the theoretical output resolution of the corrected video. Specifically, the target width = the original width × Sx; the target height = the original height × Sy.

[0125] In this solution, first, based on the aspect ratio adjustment coefficient in the calibration parameters, the target video size is accurately calculated through affine transformation to ensure that the output size strictly conforms to the standard human body ratio. Then, the "bicubic interpolation algorithm" is used to perform resolution resampling. While adjusting the video size, the edge details are retained by weighted calculation of 16-neighbor pixels, avoiding the problem of blurred images caused by traditional linear interpolation. Then, for the blank edge area generated by resampling, content-aware filling technology combined with Poisson fusion algorithm is applied to intelligently synthesize the background content, eliminating black edges or stretching distortion. Finally, through a combined optimization strategy of temporal noise reduction (eliminating inter-frame flicker) and adaptive sharpening (enhancing detail texture), the visual quality of the corrected video is comprehensively improved.

[0126] In addition, in another embodiment of the present application, the method may further include the following steps:

[0127] Step C1: Extract verification frames from the target video;

[0128] Step C2: Detect the second human body key point information in the verification frames, and determine the second key part ratio according to the second human body key point information;

[0129] Step C3: Calculate the matching degree between the second key part ratio and the standard human body ratio;

[0130] Step C4: When the matching degree is lower than the preset threshold, recalculate the video ratio calibration parameters.

[0131] Verification frames are detection sample frames extracted at equal intervals from the corrected video (such as taking 1 frame every 5 seconds), which are used to characterize the quality state of the final output video.

[0132] The second human body key point information refers to the key point data obtained by re-detecting the corrected video frames, which is used to verify the correction effect.

[0133] The second key part ratio refers to the human body ratio parameter newly calculated based on the key points of the verification frames, which should normally be highly consistent with the standard ratio.

[0134] In the embodiment of the present application, first, verification frames are extracted from the target video by the time-equal sampling method, and the sampling density is dynamically adjusted according to the video length (for example, 12 frames are taken for a 1-minute video) to ensure coverage of the entire video period. Then, the same HRNet model as the initial detection (ensuring the same standard) is used to detect the key point information, but the following optimizations are made to the model input: increasing the input resolution (higher precision); enabling the TTA (Test Time Augmentation) strategy; outputting the three-dimensional coordinates (x, y, confidence) of 17 key points. Then, a hierarchical evaluation strategy is adopted to calculate the matching degree: Basic matching: calculating the absolute error between the proportion of each part and the standard value; Advanced matching: evaluating the consistency of the spatial distribution of key points; Comprehensive scoring: weighted averaging each index (weight configuration: 60% for basic + 40% for advanced). Finally, when the matching degree < the preset threshold (such as 0.85), a three-level optimization mechanism is triggered: Parameter fine-tuning: iterating within the range of ±5% based on the original calibration parameters, Region optimization: recalculating only the local areas with low matching degree, Full process rollback: when there is a serious mismatch (such as the matching degree < 0.7), re-executing S101 - S105.

[0135] Through this solution, an independent verification mechanism can be used to ensure the calibration accuracy and improve the accuracy of the human body proportion in the output video.

[0136] In the embodiment of the present application, first, a video to be processed is obtained, and a plurality of video frames to be processed are extracted from the video to be processed. Then, for each video frame to be processed, the first human body key point information in the video frame to be processed is detected, and the first key part proportion is determined according to the first human body key point information. Furthermore, according to the first key part proportion corresponding to each video frame to be processed and the standard human body proportion, a video proportion correction parameter is determined. Finally, the video to be processed is subjected to proportion correction processing according to the video proportion correction parameter to obtain a target video. Through the present application, the video proportion can be automatically adjusted efficiently and accurately, which is applicable to the batch processing of a large amount of content on a video platform, greatly improving the processing efficiency while ensuring the calibration accuracy and significantly improving the video viewing experience.

[0137] See Figure 2 , which is a flowchart of an embodiment of another video processing method provided by the embodiment of the present application. The Figure 2 shown process is based on the process shown above Figure 1 and describes how to extract a plurality of video frames to be processed from the video to be processed. As Figure 2 shown, the process may include the following steps:

[0138] S201. Analyze the motion characteristics and scene change characteristics of the video content in the video to be processed, where the motion characteristics are used to characterize the intensity distribution of pixel displacement between video frames, and the scene change characteristics are used to characterize the spatio-temporal distribution characteristics of video shot switching.

[0139] The motion feature refers to the intensity of pixel displacement between frames calculated by the optical flow algorithm, which quantifies the severity of object motion in the video (e.g., fast actions will generate high-intensity optical flow values).

[0140] The scene change feature refers to the shot transition event detected based on the difference in HSV histograms, which identifies the time points when the video content changes abruptly (such as shot transitions and scene transitions).

[0141] The embodiments of this application include two parts: motion feature extraction and scene change detection. Among them, motion feature extraction: The Farneback dense optical flow algorithm is used to calculate the displacement vectors of all pixel points between consecutive frames, and the 90th percentile of its amplitude is statistically calculated as the feature value; Scene change detection: Construct a sequence of HSV color histograms, and identify shot transition points through the chi-square distance (threshold 0.35); Output feature matrix: It includes the time-domain motion intensity curve and the scene transition marker sequence.

[0142] S202. Determine a sampling strategy according to the motion feature and the scene change feature, and the sampling strategy is used to characterize the mapping relationship between the frame sampling frequency and the video content change.

[0143] The sampling strategy refers to establishing a dynamic mapping rule between the content complexity and the sampling density, which includes a basic sampling rate, a motion compensation coefficient, and a scene compensation parameter.

[0144] In the embodiments of this application, S202 may specifically include the following steps: When the intensity of pixel displacement between frames represented by the motion feature of the video content is greater than the motion intensity threshold, determine that the sampling frequency corresponding to the video content is the first sampling frequency, and the first sampling frequency is less than the reference sampling frequency; When the shot transition frequency represented by the scene change feature of the video content is greater than the scene transition threshold, determine that the sampling frequency corresponding to the video content is the second sampling frequency, and the second sampling frequency is greater than the reference sampling frequency; When the intensity of pixel displacement between frames represented by the motion feature of the video content is less than or equal to the motion intensity threshold, and the shot transition frequency represented by the scene change feature of the video content is less than or equal to the scene transition threshold, determine that the sampling frequency corresponding to the video content is the reference sampling frequency; Generate the sampling strategy according to the sampling frequencies corresponding to all video content.

[0145] This solution includes two parts: basic sampling and dynamic adjustment mechanism:

[0146] Basic sampling: Sample 1 frame every N frames (default N = 5); Motion compensation: When the optical flow intensity > threshold T1, the value of N decreases linearly (down to 1 at least); Scene compensation: Add M frames before and after the shot transition point (default M = 2).

[0147] Dynamic adjustment mechanism: Motion-sensitive area: Automatically increase the sampling rate to 30fps; Scene transition area: Ensure that the sampling density doubles within 0.5 seconds before and after the switching point; Static area: Maintain the basic sampling rate to save resources.

[0148] S203. Extract multiple video frames to be processed from the video to be processed according to the sampling strategy.

[0149] In the embodiments of the present application, a non-uniform sampling time series is generated according to the sampling strategy, and multiple video frames to be processed are extracted from the video to be processed according to the sampling time series.

[0150] In the application, forced sampling of the first and last frames can be adopted to ensure sampling integrity.

[0151] The solution provided by the embodiments of the present application significantly improves the integrity and accuracy of key frame capture based on the bimodal analysis of motion features and scene changes. Moreover, through the adaptive sampling strategy, while ensuring comprehensive coverage of multi-scene content in the video, the amount of redundant frame processing is significantly reduced, and the overall processing efficiency is improved.

[0152] See Figure 3 , which is a flowchart of an embodiment of another video processing method provided by the embodiments of the present application. The Figure 3 shown process is based on the process shown above Figure 1 and describes how to detect the first human key point information in the video frame to be processed. As Figure 3 shown, the process may include the following steps:

[0153] S301. Perform a preprocessing operation on the video frame to be processed.

[0154] The preprocessing operation refers to the image optimization processing performed on the original video frame.

[0155] In the embodiments of the present application, the preprocessing operation includes the following three aspects: Light compensation: Use the CLAHE algorithm to equalize the luminance channel to achieve light normalization (eliminate luminance differences); Noise processing: Use non-local means denoising (σ = 3) to reduce image noise; Image enhancement: Edge sharpening and contrast stretching to achieve image enhancement.

[0156] S302. Use a pre-trained key point detection model to perform key point recognition on the video frame to be processed after the preprocessing operation to obtain human key point coordinate data.

[0157] The key point detection model refers to a 17-point human pose estimation model trained based on datasets such as COCO and deep learning architectures such as HRNet.

[0158] In the embodiments of the present application, the video frame to be processed after preprocessing is input into the key point detection model, and the (x, y, confidence) three-dimensional data of 17 key points (i.e., human key point coordinate data) is output by the key point detection model.

[0159] S303. Verify the credibility of the human key point coordinate data according to the topological constraint relationship between the key points.

[0160] The topological constraint relationship is used to describe the spatial constraint rules of the physiological structure between human key points, including the range of bone length ratio and the limitation of joint movement angle (such as the knee joint cannot be located above the hip joint).

[0161] In the embodiments of the present application, verifying the credibility of the human key point coordinate data through the topological constraint relationship between the key points includes: checking whether the positions of the key points conform to the human structure to avoid errors such as knees growing on the shoulders; analyzing whether the human actions in consecutive frames are smooth to avoid unnatural phenomena such as instantaneous hand movement.

[0162] S304. Store the verified human key point coordinate data in time series to obtain the first human key point information.

[0163] In the embodiments of the present application, the verified human key point coordinate data is retained, and the unverified data is excluded. The verified human key point coordinate data and corresponding data such as timestamps are encoded and stored in a structured format (such as ProtocolBuffers binary format) to ensure data integrity and access efficiency. Secondly, by establishing a multi-level index system based on the video timeline, fast data retrieval with millisecond-level accuracy is supported. At the same time, a real-time compression algorithm can be used to reduce the storage space occupancy.

[0164] The solution provided by the embodiments of the present application, firstly, improves the image quality through preprocessing, providing a stable input basis for subsequent detection. Then, the key point recognition based on the advanced deep learning model achieves a high-precision positioning effect, greatly improving the detection accuracy. At the same time, abnormal data is effectively filtered through strict physiological constraints and kinematic analysis, improving the data accuracy and data reliability, and laying a solid data foundation for the subsequent ratio correction process.

[0165] Based on the same technical concept, the embodiments of the present application also provide a video processing device, as Figure 4 shown, the device includes:

[0166] An acquisition module 41, configured to acquire a video to be processed;

[0167] An extraction module 42, configured to extract a plurality of video frames to be processed from the video to be processed;

[0168] A detection module 43, configured to detect first human key-point information in each video frame to be processed, and determine a first key part ratio according to the first human key-point information;

[0169] A determination module 44, configured to determine a video ratio correction parameter according to the first key part ratio corresponding to each video frame to be processed and a standard human ratio;

[0170] A correction module 45, configured to perform a ratio correction process on the video to be processed according to the video ratio correction parameter to obtain a target video.

[0171] In a possible implementation manner, the extraction module is specifically configured to:

[0172] Analyze the motion features and scene change features of the video content in the video to be processed, where the motion features are used to characterize the intensity distribution of pixel displacement between video frames, and the scene change features are used to characterize the spatio-temporal distribution characteristics of video shot switching;

[0173] Determine a sampling strategy according to the motion features and scene change features, where the sampling strategy is used to characterize the mapping relationship between the frame sampling frequency and the video content change;

[0174] Extract a plurality of video frames to be processed from the video to be processed according to the sampling strategy.

[0175] In a possible implementation manner, the extraction module is further configured to:

[0176] When the intensity of pixel displacement between frames represented by the motion features of the video content is greater than a motion intensity threshold, determine that the sampling frequency corresponding to the video content is a first sampling frequency, and the first sampling frequency is less than a reference sampling frequency;

[0177] When the shot switching frequency represented by the scene change features of the video content is greater than a scene switching threshold, determine that the sampling frequency corresponding to the video content is a second sampling frequency, and the second sampling frequency is greater than the reference sampling frequency;

[0178] When the intensity of pixel displacement between frames represented by the motion features of the video content is less than or equal to the motion intensity threshold, and the shot switching frequency represented by the scene change features of the video content is less than or equal to the scene switching threshold, determine that the sampling frequency corresponding to the video content is the reference sampling frequency;

[0179] Generate the sampling strategy according to the sampling frequencies corresponding to all video contents.

[0180] In a possible implementation manner, the detection module is specifically configured to:

[0181] Perform preprocessing operations on the video frame to be processed;

[0182] Use a pre-trained key point detection model to identify key points in the preprocessed video frame to be processed, and obtain human key point coordinate data;

[0183] Verify the credibility of the human key point coordinate data according to the topological constraint relationship between key points;

[0184] Store the verified human key point coordinate data in time series to obtain the first human key point information.

[0185] In a possible implementation manner, the determining module is specifically configured to:

[0186] Perform statistical analysis on the first key part ratio corresponding to each video frame to be processed to obtain a statistical result;

[0187] Remove outliers from the statistical result to obtain valid data;

[0188] Calculate the reference human body ratio according to the valid data;

[0189] Compare the reference human body ratio with the standard human body ratio to obtain the correction parameter.

[0190] In a possible implementation manner, the correction module is specifically configured to:

[0191] Calculate the target video size according to the video ratio correction parameter;

[0192] Perform resolution resampling on the video to be processed according to the target video size to obtain a first video that meets the target video size;

[0193] Perform edge processing on the first video to obtain a second video;

[0194] Optimize the quality of the second video through a video enhancement algorithm to obtain the target video.

[0195] In a possible implementation manner, the device further includes a verification module for:

[0196] Extract verification frames from the target video;

[0197] Detect the second human key point information in the verification frame, and determine the second key part ratio according to the second human key point information;

[0198] Calculate the matching degree between the second key part ratio and the standard human body ratio;

[0199] In the case that the matching degree is lower than the preset threshold, recalculate the video ratio correction parameter.

[0200] Based on the same technical concept, an embodiment of the present application further provides an electronic device, as Figure 5 shown, including a processor 111, a communication interface 112, a memory 113, and a communication bus 114. Among them, the processor 111, the communication interface 112, and the memory 113 complete mutual communication through the communication bus 114.

[0201] The memory 113 is used to store a computer program.

[0202] When the processor 111 is used to execute the program stored on the memory 113, the following steps are implemented:

[0203] Obtain a video to be processed.

[0204] Extract a plurality of video frames to be processed from the video to be processed.

[0205] For each video frame to be processed, detect the first human body key point information in the video frame to be processed, and determine the first key part ratio according to the first human body key point information.

[0206] Determine the video ratio correction parameter according to the first key part ratio corresponding to each video frame to be processed and the standard human body ratio.

[0207] Perform ratio correction processing on the video to be processed according to the video ratio correction parameter to obtain a target video.

[0208] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0209] The communication interface is used for communication between the above electronic device and other devices.

[0210] The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0211] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0212] In another embodiment provided by the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above video processing methods are implemented.

[0213] In another embodiment provided by the present application, there is also provided a computer program product containing instructions, which when run on a computer, causes the computer to execute any of the video processing methods in the above embodiments.

[0214] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0215] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the related technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0216] It should be understood that the terms used herein are for the purpose of describing particular example embodiments only and are not intended to be limiting. Unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" as used herein may also include the plural forms. The terms "comprising", "including", "containing", and "having" are inclusive and thus specify the presence of stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be performed in the particular order described or illustrated, unless an execution order is explicitly stated. It should also be understood that additional or alternative steps may be used.

[0217] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.

Claims

1. A video processing method, characterized in that, The method includes: Obtaining a video to be processed; Extracting a plurality of video frames to be processed from the video to be processed; For each video frame to be processed, detecting first human key point information in the video frame to be processed, and determining a first key part ratio according to the first human key point information; Determining a video ratio correction parameter according to the first key part ratio corresponding to each video frame to be processed and a standard human ratio; Performing ratio correction processing on the video to be processed according to the video ratio correction parameter to obtain a target video.

2. The method according to claim 1, characterized in that The extracting a plurality of video frames to be processed from the video to be processed includes: Analyzing the motion characteristics and scene change characteristics of the video content in the video to be processed, where the motion characteristics are used to characterize the intensity distribution of pixel displacement between video frames, and the scene change characteristics are used to characterize the spatio-temporal distribution characteristics of video shot switching; Determining a sampling strategy according to the motion characteristics and scene change characteristics, where the sampling strategy is used to characterize the mapping relationship between the frame sampling frequency and the video content change; Extracting a plurality of video frames to be processed from the video to be processed according to the sampling strategy.

3. The method according to claim 2, wherein The determining a sampling strategy according to the motion characteristics and scene change characteristics includes: When the intensity of pixel displacement between frames represented by the motion characteristics of the video content is greater than a motion intensity threshold, determining that the sampling frequency corresponding to the video content is a first sampling frequency, and the first sampling frequency is less than a reference sampling frequency; When the shot switching frequency represented by the scene change characteristics of the video content is greater than a scene switching threshold, determining that the sampling frequency corresponding to the video content is a second sampling frequency, and the second sampling frequency is greater than the reference sampling frequency; When the intensity of pixel displacement between frames represented by the motion characteristics of the video content is less than or equal to the motion intensity threshold, and the shot switching frequency represented by the scene change characteristics of the video content is less than or equal to the scene switching threshold, determining that the sampling frequency corresponding to the video content is the reference sampling frequency; Generating the sampling strategy according to the sampling frequencies corresponding to all video content.

4. The method according to claim 1, wherein The detecting first human key point information in the video frame to be processed includes: Performing a preprocessing operation on the video frame to be processed; Using a pre-trained key point detection model to perform key point recognition on the preprocessed video frame to be processed to obtain human key point coordinate data; Performing credibility verification on the human key point coordinate data according to the topological constraint relationship between key points; Storing the verified human key point coordinate data in time series to obtain the first human key point information.

5. The method according to claim 1, wherein The determining a video ratio correction parameter according to the first key part ratio corresponding to each video frame to be processed and a standard human ratio includes: Performing statistical analysis on the first key part ratio corresponding to each video frame to be processed to obtain a statistical result; Removing outliers from the statistical result to obtain valid data; Calculating a reference human ratio according to the valid data; Comparing the reference human ratio with the standard human ratio to obtain the correction parameter.

6. The method according to claim 1, characterized in that, Performing proportional correction processing on the video to be processed according to the video proportional correction parameter to obtain a target video, including: Calculating the target video size according to the video proportional correction parameter; Performing resolution resampling on the video to be processed according to the target video size to obtain a first video that meets the target video size; Performing edge processing on the first video to obtain a second video; Optimizing the quality of the second video through a video enhancement algorithm to obtain a target video.

7. The method according to claim 1, wherein The method further includes: Extracting a verification frame from the target video; Detecting second human key point information in the verification frame and determining a second key part ratio according to the second human key point information; Calculating the matching degree between the second key part ratio and the standard human ratio; Recalculating the video proportional correction parameter when the matching degree is lower than a preset threshold.

8. A video processing device, characterized in that, The device includes: An acquisition module, configured to acquire a video to be processed; An extraction module, configured to extract a plurality of video frames to be processed from the video to be processed; A detection module, configured to detect first human key point information in each video frame to be processed and determine a first key part ratio according to the first human key point information; A determination module, configured to determine a video proportional correction parameter according to the first key part ratio corresponding to each video frame to be processed and the standard human ratio; A correction module, configured to perform proportional correction processing on the video to be processed according to the video proportional correction parameter to obtain a target video.

9. An electronic device, characterized in that, Including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used for storing a computer program; The processor is configured to implement the video processing method according to any one of claims 1-7 when executing the program stored on the memory.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the video processing method according to any one of claims 1-7 is implemented.