Video processing method and device, electronic equipment and storage medium

By automatically detecting and analyzing the flower-word area of the video frame and calculating the video proportion correction parameters, the problem of manual adjustment is solved, adaptive adjustment and lossless conversion of the flower-word area is realized, and the display quality of the video on different platforms is improved.

CN120264072APending Publication Date: 2025-07-04BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510477744.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the existing video processing technology, the method of manually adjusting the flower-word area is inefficient and cannot meet the batch processing needs, resulting in the flower-word area being easily deformed or lost when the resolution is adjusted, affecting the viewing experience and the integrity of information communication.

Method used

By automatically detecting the flower-like areas in four corners of the video frame, obtaining position information, analyzing and filtering the stable high-frequency flower-like areas, calculating the video proportion correction parameters, and performing geometric transformation processing to achieve adaptive adjustment of the video proportion.

Benefits of technology

The precise positioning and adaptive adjustment of the flower character area is realized, which avoids deformation or loss caused by resolution adjustment, eliminates the time bottleneck of manual frame-by-frame operation, and ensures the integrity and adaptability of flower character in different playback scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120264072A_ABST
    Figure CN120264072A_ABST
Patent Text Reader

Abstract

The invention provides a video processing method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring a to-be-processed video, and splitting the to-be-processed video into a plurality of video frames to obtain a video frame sequence; aiming at each video frame in the video frame sequence, detecting flower character areas at four corners of the video frame to obtain flower character area data containing flower character area position information; analyzing and processing the flower character region data of all the video frames to obtain a target flower character region set meeting a preset condition; performing comparison calculation on the basis of the size information of each target flower character region in the target flower character region set and a standard proportion to obtain a video proportion correction parameter; and performing geometric transformation processing on the video frame sequence according to the video proportion correction parameter to obtain a corrected target video. The deformation or loss of the patterns caused by resolution adjustment is avoided, and the time bottleneck of manual frame-by-frame operation is eliminated through a batch processing mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of video processing, and in particular, to a video processing method, apparatus, electronic device, and storage medium. Background Art

[0002] In the process of video production and dissemination, subtitles (such as logos, corner marks, specific text identifiers, etc.) are usually embedded in the four corners of the video frame for marking key information such as copyright information and channel identifiers. However, since the video may undergo resolution adjustment, scaling, or cropping when played on different platforms or devices, the subtitle area is prone to being squeezed, deformed, or even partially lost, affecting the viewing experience and the integrity of information transmission. Therefore, how to accurately detect the subtitle area and adaptively adjust the video ratio has become one of the key issues in video post-processing and adaptation. Currently, the commonly used video ratio adjustment method is the manual adjustment method, specifically: adjusting the frame ratio frame by frame through professional video editing software (such as Adobe Premiere, Final Cut Pro).

[0003] However, this manual adjustment method is inefficient and cannot meet the batch processing requirements. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a video processing method, apparatus, electronic device, and storage medium to solve the problem that the manual adjustment method is inefficient and cannot meet the batch processing requirements. The specific technical solutions are as follows:

[0005] In a first aspect, this application provides a video processing method, including:

[0006] Obtain a video to be processed, and split the video to be processed into a plurality of video frames to obtain a video frame sequence;

[0007] For each video frame in the video frame sequence, detect the subtitle area at the four corners of the video frame to obtain subtitle area data including the position information of the subtitle area;

[0008] Analyze and process the subtitle area data of all video frames to obtain a set of target subtitle areas that meet the preset conditions;

[0009] Based on the comparison calculation between the size information of each target subtitle area in the set of target subtitle areas and the standard ratio, obtain a video ratio correction parameter;

[0010] Perform geometric transformation processing on the video frame sequence according to the video ratio correction parameter to obtain a corrected target video.

[0011] In a possible implementation, detecting the captioned regions at the four corners of the video frame to obtain captioned region data including the position information of the captioned regions includes:

[0012] Performing image segmentation processing on the four preset corner regions of the video frame respectively to obtain four regions to be processed;

[0013] For each region to be processed, performing edge detection processing on the region to be processed to obtain edge feature data;

[0014] Performing contour extraction processing on the edge feature data to obtain a candidate contour set;

[0015] Performing feature screening processing on the candidate contour set to obtain valid captioned contours;

[0016] Performing position calculation processing based on the valid captioned contours to obtain captioned region data including the position information of the captioned regions.

[0017] In a possible implementation, the performing feature screening processing on the candidate contour set to obtain valid captioned contours includes:

[0018] For each contour in the candidate contour set, performing geometric feature extraction processing on the contour to obtain contour feature data, where the contour feature data includes at least one of the following features: the ratio of the contour area to a preset area threshold, the width-to-height ratio of the circumscribed rectangle of the contour, and the complexity eigenvalue of the contour edge;

[0019] Performing matching degree calculation processing on the contour feature data and a preset captioned feature model to obtain a matching degree score;

[0020] Determining the contours with corresponding matching degree scores greater than a preset threshold as valid captioned contours.

[0021] In a possible implementation, the analyzing and processing the captioned region data of all video frames to obtain a set of target captioned regions meeting preset conditions includes:

[0022] Calculating the average position of each captioned region within a sliding window and excluding the captioned regions with a position offset exceeding a first threshold to obtain a preliminary screening result;

[0023] Calculating the stability score of each captioned region in the preliminary screening result and excluding the captioned regions with a stability score lower than a second threshold to obtain a set of target captioned regions.

[0024] In a possible implementation, the calculating the stability score of each captioned region in the preliminary screening result includes:

[0025] Statistically analyze the occurrence frequency of each subtitle area in the preliminary screening results, where the occurrence frequency refers to the proportion of the subtitle area in the total number of frames;

[0026] Calculate the average value of the shape similarity scores of subtitle areas between adjacent video frames in the preliminary screening results to obtain a consistency score;

[0027] For each subtitle area in the preliminary screening results, perform a weighted summation calculation on the occurrence frequency and consistency score of the subtitle area to obtain the stability score of the subtitle area.

[0028] In a possible implementation manner, the method of calculating the video ratio correction parameter by comparing the size information of each target subtitle area in the target subtitle area set with a standard ratio includes:

[0029] Perform statistical analysis on the size information of each target subtitle area in the target subtitle area set to obtain statistical size data;

[0030] Perform a difference calculation on the statistical size data and the standard ratio to obtain a ratio difference parameter;

[0031] Perform an optimization and adjustment process on the ratio difference parameter to obtain a video ratio correction parameter.

[0032] In a possible implementation manner, the method of performing geometric transformation on the video frame sequence according to the video ratio correction parameter to obtain a corrected target video includes:

[0033] Perform frame-by-frame scaling on the video frame sequence according to the video ratio correction parameter to obtain a video frame sequence with adjusted ratio;

[0034] Perform edge compensation on the video frame sequence with adjusted ratio to obtain a video frame sequence with edge compensation;

[0035] Perform encoding and synthesis on the video frame sequence with edge compensation to obtain a corrected target video.

[0036] In a second aspect, the present application provides a video processing device, including:

[0037] An acquisition module, configured to acquire a video to be processed and split the video to be processed into a plurality of video frames to obtain a video frame sequence;

[0038] A detection module, configured to detect subtitle areas at four corners of each video frame in the video frame sequence to obtain subtitle area data including subtitle area position information;

[0039] An analysis module for analyzing and processing the caption area data of all video frames to obtain a set of target caption areas that meet preset conditions;

[0040] A calculation module for comparing and calculating the size information of each target caption area in the set of target caption areas with a standard ratio to obtain a video ratio correction parameter;

[0041] A correction module for performing geometric transformation processing on the video frame sequence according to the video ratio correction parameter to obtain a corrected target video.

[0042] In a possible implementation manner, the detection module is specifically configured to:

[0043] Perform image segmentation processing on four preset corner areas of the video frame to obtain four areas to be processed;

[0044] For each area to be processed, perform edge detection processing on the area to be processed to obtain edge feature data;

[0045] Perform contour extraction processing on the edge feature data to obtain a set of candidate contours;

[0046] Perform feature screening processing on the set of candidate contours to obtain effective caption contours;

[0047] Perform position calculation processing based on the effective caption contours to obtain caption area data including caption area position information.

[0048] In a possible implementation manner, the detection module is further configured to:

[0049] For each contour in the set of candidate contours, perform geometric feature extraction processing on the contour to obtain contour feature data, where the contour feature data includes at least one of the following features: the ratio of the contour area to a preset area threshold, the width-to-height ratio of the circumscribed rectangle of the contour, and the complexity eigenvalue of the contour edge;

[0050] Perform matching degree calculation processing on the contour feature data and a preset caption feature model to obtain a matching degree score;

[0051] Determine the contours with corresponding matching degree scores greater than a preset threshold as effective caption contours.

[0052] In a possible implementation manner, the analysis module is specifically configured to:

[0053] Calculate the average position of each caption area within a sliding window, and eliminate caption areas with a position offset exceeding a first threshold to obtain a preliminary screening result;

[0054] Calculate the stability scores of each caption area in the preliminary screening results, and remove the caption areas with stability scores lower than the second threshold to obtain a set of target caption areas.

[0055] In a possible implementation, the analysis module is further configured to:

[0056] Count the occurrence frequencies of each caption area in the preliminary screening results, where the occurrence frequency refers to the proportion of the caption area appearing in the total number of frames;

[0057] Calculate the average value of the shape similarity scores of caption areas between adjacent video frames in the preliminary screening results to obtain a consistency score;

[0058] For each caption area in the preliminary screening results, perform a weighted summation calculation on the occurrence frequency and the consistency score of the caption area to obtain the stability score of the caption area.

[0059] In a possible implementation, the calculation module is specifically configured to:

[0060] Perform statistical analysis processing on the size information of each target caption area in the set of target caption areas to obtain statistical size data;

[0061] Perform a difference calculation process on the statistical size data and a standard ratio to obtain a ratio difference parameter;

[0062] Perform an optimization and adjustment process on the ratio difference parameter to obtain a video ratio correction parameter.

[0063] In a possible implementation, the correction module is specifically configured to:

[0064] Perform frame-by-frame scaling processing on the video frame sequence according to the video ratio correction parameter to obtain a video frame sequence with adjusted ratio;

[0065] Perform edge compensation processing on the video frame sequence with adjusted ratio to obtain a video frame sequence with edge compensation;

[0066] Perform encoding and synthesis processing on the video frame sequence with edge compensation to obtain a corrected target video.

[0067] In a third aspect, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus, where the processor, the communication interface, and the memory communicate with each other through the communication bus;

[0068] The memory is used to store a computer program;

[0069] The processor is configured to, when executing the program stored in the memory, implement the method steps of any one of the first aspect.

[0070] In a fourth aspect, a computer-readable storage medium is provided, characterized in that a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method steps described in any one of the first aspects are implemented.

[0071] In a fifth aspect, a computer program product containing instructions is provided, which, when running on a computer, causes the computer to execute the video processing method described above.

[0072] Advantages of the embodiments of the present application:

[0073] The embodiments of the present application provide a video processing method, device, electronic device and storage medium. In the embodiments of the present application, first, a video to be processed is obtained, and the video to be processed is split into a plurality of video frames to obtain a video frame sequence. For each video frame in the video frame sequence, the caption areas at the four corners of the video frame are detected to obtain caption area data containing the position information of the caption areas. Then, the caption area data of all video frames are analyzed and processed to obtain a set of target caption areas that meet the preset conditions. Next, based on the comparison calculation of the size information of each target caption area in the set of target caption areas with the standard ratio, a video ratio correction parameter is obtained. Finally, geometric transformation processing is performed on the video frame sequence according to the video ratio correction parameter to obtain a corrected target video. The present application accurately locates the caption areas through automated frame sequence analysis, and uses mathematical modeling to derive the optimal cropping parameters, which not only avoids caption distortion or loss caused by resolution adjustment (such as the logo being cropped or the corner mark being stretched), but also eliminates the time bottleneck of manual frame-by-frame operation through a batch processing mechanism. Finally, while maintaining the integrity of the captions, a lossless conversion of the video ratio to different playback scenarios is achieved.

[0074] Of course, it is not necessary for any product or method implementing the present application to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention.

[0076] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0077] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings represent similar elements. Unless otherwise stated, the drawings in the figures do not constitute a scale limitation.

[0078] Figure 1 It is a flowchart of a video processing method provided by an embodiment of the present application;

[0079] Figure 2 It is a flowchart of another video processing method provided by an embodiment of the present application;

[0080] Figure 3 It is a flowchart of yet another video processing method provided by an embodiment of the present application;

[0081] Figure 4 It is a schematic structural diagram of a video processing device provided by an embodiment of the present application;

[0082] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0083] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.

[0084] The following disclosure provides many different embodiments or examples for implementing different structures of the present invention. To simplify the disclosure of the present invention, components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present invention. In addition, the present invention may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed.

[0085] Figure 1The flowchart of a video processing method provided by an embodiment of this application. This method can be applied to one or more electronic devices such as smartphones, laptops, desktop computers, portable computers, servers, etc. In addition, the execution subject of this method can be hardware or software. When the above execution subject is hardware, the execution subject can be one or more of the above electronic devices. For example, a single electronic device can execute this method, or multiple electronic devices can cooperate with each other to execute this method. When the above execution subject is software, this method can be implemented as multiple software or software modules, or can be implemented as a single software or software module. No specific limitation is made here.

[0086] As Figure 1 shown, this method specifically includes:

[0087] S101. Obtain the video to be processed, and split the video to be processed into several video frames to obtain a video frame sequence.

[0088] The video to be processed (V) refers to the original video file (such as MP4, MOV, etc.) that needs to be scaled, and may contain caption elements such as station logos and corner marks.

[0089] The video frame sequence refers to that a video is composed of consecutive static images (frames), and the set of frames obtained after disassembly, denoted as {F1, F2,..., Fn}, where N is the total number of video frames, that is, the video (V) is split into independent frames

[0090] In the embodiment of this application, FFmpeg or VideoCapture of OpenCV can be used for frame splitting, and a fixed interval (such as 1 frame per second) or dynamic key frame extraction (I-frame priority) can be set. The output is a video frame sequence sorted by time stamps, retaining the original resolution and color space (such as RGB or YUV).

[0091] S102. For each video frame in the video frame sequence, detect the caption areas at the four corners of the video frame to obtain caption area data including the position information of the caption areas.

[0092] The caption area refers to the fixed elements at the four corners (upper left, upper right, lower left, lower right) of the video screen, such as station logos, corner marks, copyright information, etc.

[0093] The position information is usually represented by the coordinates of a rectangular box, in the format of (x_min, y_min, x_max, y_max), or relative coordinates (normalized to the [0, 1] interval).

[0094] In the embodiment of this application, for each video frame F t, for the caption areas at its four corners (upper left, upper right, lower left, lower right), the function (D(F t , p)) is used for caption detection, where (p) is the detection position parameter.

[0095] Define the caption area detection function:

[0096]

[0097] Among them, (p ∈ LU, RU, LD, RD) represents the positions of the upper left, upper right, lower left, and lower right corners.

[0098] The execution process of this function is as follows:

[0099] Input parameters: Include the current video frame and the specified detection area (upper left, upper right, lower left, lower right).

[0100] Select the detection area: According to the input detection position parameter, select a specific corner of the image as the detection area.

[0101] Contour recognition: Extract the contours in the selected area, which may contain captions.

[0102] Feature screening: Screen the extracted contours according to predefined features (such as the area, shape, edge features, etc. of the contours) to determine whether they are caption areas protecting captions.

[0103] Output result: Return the position of the detected caption area in the image, usually represented by the coordinates of a rectangular box.

[0104] In this way, the caption area can be effectively detected and marked within a specific area of the video frame.

[0105] As for how to specifically screen the caption areas containing captions to obtain caption area data, it will be explained through the following embodiments and will not be elaborated here first.

[0106] S103. Analyze and process the caption area data of all video frames to obtain a set of target caption areas that meet the preset conditions.

[0107] The preset conditions refer to the thresholds for screening caption areas that appear stably and frequently, such as the minimum appearance frequency (e.g., 80% of the frames) and spatial consistency (overlap degree > 90%).

[0108] The set of target caption areas refers to the list of caption areas that are finally determined to need protection, excluding temporary elements (such as floating subtitles).

[0109] In the embodiments of this application, by analyzing and processing the caption area data of all video frames, caption areas that appear stably and frequently are screened out, and thus a set of target caption areas is obtained.

[0110] As for how to specifically analyze and process the caption area data of all video frames to obtain a set of target caption areas that meet the preset conditions, it will be explained in the following embodiments and will not be elaborated here first.

[0111] S104. Compare and calculate the size information of each target caption area in the set of target caption areas with the standard ratio to obtain a video ratio correction parameter.

[0112] The standard ratio refers to the aspect ratio of the video required by the target platform (such as 16:9, 9:16, 1:1, etc.).

[0113] The video ratio correction parameter refers to parameters such as the scaling factor, cropping area, or padding boundary required for geometric transformation.

[0114] In the embodiment of the present application, S104 may specifically include the following steps: statistically analyze and process the size information of each target caption area in the set of target caption areas to obtain statistical size data; perform difference calculation processing on the statistical size data and the standard ratio to obtain a ratio difference parameter; perform optimization and adjustment processing on the ratio difference parameter to obtain a video ratio correction parameter.

[0115] The statistical size data refers to the geometric features extracted from the stable set of caption areas. In applications, the average value of the sizes of each target caption area can be taken as the statistical size data.

[0116] In this embodiment, the calculation formula for the ratio difference parameter is as follows:

[0117]

[0118] Among them, Widthdetected is the width in the statistical size data; Heightdetected is the height in the statistical size data; Widthstandard is the width in the standard ratio, that is, the width of the target size; Heightstandard is the height in the standard ratio, that is, the height of the target size; ScaleFactor is the ratio difference parameter, which is a tuple containing two elements (the scale factor of the width scale_width, the scale factor of the height scale_height).

[0119] After obtaining the ratio difference parameters, the system will intelligently optimize and adjust them to generate the final video ratio correction parameters. First, the original ratio parameters are normalized to keep the width and height ratio factors coordinated. Then, different constraint rules are applied according to the video content type: for character videos, the height ratio is prioritized to prevent face deformation, while landscape videos allow more flexible width adjustment. At the same time, it will detect whether the adjusted text area exceeds the safe boundary of the picture, and automatically limit the zoom range if necessary to ensure that all key text content is fully visible.

[0120] The correction parameters output by this solution not only include the optimized scaling ratio, but also include whether the edges need to be cropped or filled, as well as the coordinate information of the flower character area that needs special protection. In this way, the image deformation caused by simple scaling is avoided, and the flower characters in different positions can be displayed correctly under various scaling adjustments.

[0121] S105 . Perform geometric transformation processing on the video frame sequence according to the video ratio correction parameter to obtain a corrected target video.

[0122] In the embodiment of the present application, the geometric transformation processing adopts a three-level processing architecture to ensure the accuracy and completeness of the video ratio adjustment.

[0123] Specifically, S105 may include the following steps: performing frame-by-frame scaling processing on the video frame sequence according to the video ratio correction parameter to obtain a video frame sequence after ratio adjustment; performing edge compensation processing on the video frame sequence after ratio adjustment to obtain an edge compensated video frame sequence; performing encoding and synthesis processing on the video frame sequence after edge compensation to obtain a corrected target video.

[0124] This step realizes the adaptive adjustment of the video ratio through a hierarchical processing mechanism: First, it is a frame-by-frame scaling processing stage: according to the calculated correction parameters, each frame of the image is geometrically transformed (such as scaling, cropping or rotation). Mathematically expressed as: let the original image width (W) and height (H), and the correction parameter be α, then the width after scaling is (W'): W'=W×α; during the scaling process, the detected flower character area is protectively scaled to keep its original width-to-height ratio unchanged.

[0125] Then, it is the edge compensation processing stage: detect the invalid edge areas (black edges) that appear in the scaled video frame, and use the content-aware edge filling algorithm: a) perform Gaussian blur expansion processing on the static background area; b) implement motion compensation prediction generation on the dynamic scene area; c) apply intelligent repair technology to the edge area containing important visual content. Smooth transition processing is performed on the compensated edge area, and the edge-compensated video frame sequence is output.

[0126] Finally, it is the encoding and synthesis processing stage: encoding the processed video frame sequence according to the format requirements of the target platform: a) The video encoding adopts the H.264 / AVC or H.265 / HEVC standard; b) The audio stream maintains the original sampling rate and bit rate; c) The encapsulation format adapts to the requirements of the target platform (such as MP4, MOV, etc.). Furthermore, reconstruct the video file header information and metadata to ensure that the frame rate and timestamp of the output video are consistent with the original video, and generate the final corrected target video file.

[0127] In this solution, first, perform protective frame-by-frame scaling based on accurately calculated correction parameters to complete the core proportional transformation while ensuring the integrity of the caption area; second, adopt intelligent edge compensation technology to dynamically eliminate the black edges generated by scaling and maintain the visual coherence of the picture through content-aware filling; finally, ensure the compatibility and quality stability of the output video through standardized encoding processing. This processing method effectively solves the problems existing in traditional video proportion adjustment, such as caption distortion, content cropping loss, and edge defects. On the premise of maintaining the main visual information of the original video, it realizes efficient and accurate adaptive proportion correction, significantly improving the display quality and user experience when the video is transmitted across platforms.

[0128] In another embodiment of this application, the method may further include the following steps: evenly distribute the video frame sequence to multiple computing units through multi-threaded task scheduling; use GPU parallel processing for the following core computing tasks: accelerating the caption area detection algorithm based on CUDA; optimizing geometric transformation calculations through texture memory; using stream processors to parallelize the edge compensation operation; establishing an inter-thread synchronization mechanism to ensure the temporal consistency of the video segments output by each processing unit.

[0129] In this step, an efficient processing is achieved through a parallel computing architecture that combines multi-threading and GPU: First, use the dynamic load balancing algorithm to intelligently distribute the video frame sequence to multiple CPU threads, and each thread independently manages the corresponding frame processing process; at the same time, utilize the parallel computing power of the GPU to accelerate key links, including: 1) Parallelly execute caption area detection based on CUDA kernels and optimize the feature matching efficiency through shared memory; 2) Map geometric transformation calculations to texture sampling operations and accelerate scaling calculations using the GPU hardware interpolation unit; 3) Deploy multiple stream processors to concurrently execute edge compensation operations in different regions. Inter-thread synchronization is achieved through a double-buffering mechanism and atomic operations to ensure that the processed video frames strictly maintain the original temporal relationship.

[0130] In this solution, through the collaborative processing of the heterogeneous parallel computing architecture, end-to-end acceleration of the video proportion correction process is achieved, obtaining a linear-level performance improvement while ensuring processing accuracy, and meeting the requirements of real-time large-scale video processing.

[0131] In another embodiment of the present application, the method may further include the following steps: establishing a video frame buffer queue; performing real-time processing on the video frames in the video frame buffer queue to obtain real-time processing results; and performing output processing on the real-time processing results to form a real-time video stream.

[0132] In this step, first, a multi-level buffer queue is created to receive input video frames, and the producer-consumer model is adopted for frame management; then, an independent processing thread is started to extract video frames from the queue, and operations such as captions detection, ratio correction, and edge compensation are performed in real time; finally, the processing results are pushed to the output streaming media server through an asynchronous encoding module.

[0133] This solution decouples the video acquisition and processing links through a buffer queue, avoids I / O blocking, meets the requirements of low-latency scenarios such as real-time live broadcast, and effectively solves the technical problem that the traditional offline processing mode cannot meet the requirements of real-time video stream processing.

[0134] In the embodiment of the present application, first, the video to be processed is obtained, and the video to be processed is split into a plurality of video frames to obtain a video frame sequence. For each video frame in the video frame sequence, the caption regions at the four corners of the video frame are detected to obtain caption region data including the position information of the caption regions. Then, the caption region data of all video frames are analyzed and processed to obtain a set of target caption regions that meet the preset conditions. Next, based on the comparison calculation of the size information of each target caption region in the set of target caption regions and the standard ratio, a video ratio correction parameter is obtained. Finally, geometric transformation processing is performed on the video frame sequence according to the video ratio correction parameter to obtain the corrected target video. The present application accurately locates the caption regions through automated frame sequence analysis and derives the optimal cropping parameters using mathematical modeling, which not only avoids caption distortion or loss caused by resolution adjustment (such as the logo being cropped or the corner mark being stretched), but also eliminates the time bottleneck of manual frame-by-frame operation through a batch processing mechanism. Finally, while maintaining the integrity of the captions, lossless conversion of the video ratio to different playback scenarios is achieved.

[0135] See Figure 2 , which is a flowchart of an embodiment of another video processing method provided by the embodiment of the present application. The Figure 2 shown process is based on the process shown above Figure 1 and describes how to detect the caption regions at the four corners of the video frame to obtain caption region data including the position information of the caption regions. As Figure 2 shown, the process may include the following steps:

[0136] S201. Perform image segmentation processing on the four preset corner regions of the video frame respectively to obtain four regions to be processed.

[0137] Preset corner regions refer to the four fixed regions in the upper left, upper right, lower left, and lower right of the video frame. In applications, each region generally accounts for 5%-15% of the total screen area.

[0138] Image segmentation processing refers to extracting pixel data of a specified rectangular region from a complete frame to generate an independent sub-image.

[0139] In the embodiments of the present application, first, according to the preset four corner coordinates (upper left, upper right, lower left, lower right), four sub-images, that is, regions to be processed, are accurately cropped from the video frame using cv::Rect (OpenCV) or equivalent methods. In applications, the size of each region is generally controlled within the range of 5%-15% of the total screen area. In this way, it is possible to avoid copying the complete frame data to save memory, and at the same time, improve the subsequent processing efficiency.

[0140] S202. For each region to be processed, perform edge detection processing on the region to be processed to obtain edge feature data.

[0141] Edge detection processing refers to identifying the pixel positions where the brightness / color changes abruptly in an image to generate a binary edge map.

[0142] Edge feature data refers to structured data containing edge pixel coordinates and gradient intensity.

[0143] In the embodiments of the present application, apply an edge detection algorithm (such as Canny edge detection) to each segmented region to generate a binary edge map. During the detection process, balance noise suppression and edge continuity through double thresholds (threshold1 = 50, threshold2 = 150), and output feature data containing edge pixel coordinates and gradient intensity to provide a basis for contour extraction.

[0144] S203. Perform contour extraction processing on the edge feature data to obtain a candidate contour set.

[0145] Contour extraction processing refers to clustering edge pixel points into closed or polygonal contours.

[0146] Candidate contour set refers to a list of all initially extracted contours, which may contain floating words and noise (such as subtitle afterimages).

[0147] In the embodiments of the present application, based on the edge feature data, use cv::findContours to extract all potential contours and filter and retain the outermost closed contours. Compress the number of contour points through polygon approximation (error threshold 1.5 pixels) to improve the calculation efficiency and obtain a candidate contour set.

[0148] S204. Perform feature screening processing on the candidate contour set to obtain valid floating word contours.

[0149] Valid caption outline: The outline selected through screening, including targets such as logos and corner marks.

[0150] In the embodiment of the present application, S204 may specifically include the following steps: For each outline in the candidate outline set, perform geometric feature extraction processing on the outline to obtain outline feature data, where the outline feature data includes at least one of the following features: the ratio of the outline area to a preset area threshold, the width-to-height ratio of the circumscribed rectangle of the outline, and the complexity eigenvalue of the outline edge; perform a matching degree calculation process on the outline feature data and a preset caption feature model to obtain a matching degree score; determine the outline with a corresponding matching degree score greater than the preset threshold as a valid caption outline.

[0151] Outline feature data refers to a set of geometric features stored in a structured manner for subsequent matching. The set includes at least one of the ratio of the outline area to a preset area threshold, the width-to-height ratio of the circumscribed rectangle of the outline, and the complexity eigenvalue of the outline edge.

[0152] The preset caption feature model refers to a pre-trained caption outline feature template library (such as the typical width-to-height ratio range of logos and corner marks), which is obtained through statistical learning.

[0153] In this solution, first, calculate the following features for each candidate outline: Area ratio: the ratio of the actual area of the outline to a preset threshold (such as 100 pixels) to filter out too small / too large noise; Width-to-height ratio: the ratio of the width to the height of the minimum circumscribed rectangle of the outline to exclude non-caption shapes (such as thin lines or circular elements); Edge complexity: distinguish simple icons from complex background noise by the ratio of the outline perimeter to the area (compactness) or Hu moments to quantify the degree of edge tortuosity.

[0154] Then, calculate the matching degree score through weighted matching degree:

[0155] Score = w1·AreaMatch + w2·AspectMatch + w3·ComplexityMatch, where the weights w1 + w2 + w3 = 1 (recommended values: 0.4, 0.3, 0.3); AreaMatch: the degree of matching between the actual area of the contour and the preset standard area of the subtitle. If it is a perfect match (measured area = standard area), the full score is 1 point. For every 10% increase in the area deviation, the score decreases linearly (e.g., a 20% deviation results in a score of 0.8); AspectMatch: the similarity between the aspect ratio of the bounding rectangle of the contour and the standard aspect ratio of the subtitle. When the aspect ratio is exactly the same as the standard value, the score is 1 point. The greater the difference between the actual ratio and the standard value, the lower the score (e.g., for a standard ratio of 1.5 and a measured ratio of 1.4, the score is 0.93); ComplexityMatch: the matching score based on the complexity of the contour edge, quantifying the similarity through a mathematical model (such as Hu moments or compactness). A perfect match gets 1 point, and the difference decays exponentially (e.g., when the difference is 0.3, the score is 0.74).

[0156] Dynamic threshold screening: The matching degree threshold is adaptively adjusted according to the video resolution (e.g., for a 1080P video, the threshold = 0.7; for a 4K video, the threshold = 0.75).

[0157] Finally, perform effective contour determination: Retain the contours with Score > threshold and mark them as effective subtitle contours.

[0158] Through the feature fusion detection method in this solution, it can adapt to different styles of subtitles (such as transparent logos, shadowed icons, etc.), achieve stable and reliable subtitle area positioning in complex video scenarios, and provide accurate input data for subsequent video aspect ratio correction.

[0159] S205. Perform position calculation and processing based on the effective subtitle contours to obtain subtitle area data containing the position information of the subtitle area.

[0160] In the embodiments of this application, after obtaining the effective subtitle contours, the following processing is performed to generate standardized subtitle area position data:

[0161] Bounding box generation: Calculate the minimum bounding rectangle for each effective contour (e.g., using cv::boundingRect) to obtain its pixel-level coordinates (x_min, y_min, x_max, y_max), accurately bounding the range of the subtitle area.

[0162] Coordinate normalization: Convert the absolute coordinates to relative coordinates (e.g., the position of the top-left subtitle is represented as (0.05, 0.03, 0.15, 0.10)), thereby eliminating the influence of video resolution differences and adapting to the requirements of multiple platforms.

[0163] Confidence assignment: According to the proportion of the contour area and the edge intensity, calculate the confidence (0-1) of the existence of the fancy words, which is used for subsequent temporal filtering or weight assignment.

[0164] In addition, in another embodiment of the present application, after performing position calculation processing based on the effective fancy word contour, the following steps may further be included: successively perform dilation operation and erosion operation on the binary mask generated by the effective fancy word contour, using a rectangular structuring element with a preset size; perform refined adjustment on the boundary of the fancy word area through a combination of opening operation and closing operation, where the opening operation uses an elliptical structuring element and the closing operation uses a rectangular structuring element; output the fancy word area mask after morphological optimization.

[0165] In this solution, first, perform dilation-erosion operation on the binary mask generated by the effective fancy word contour, use the rectangular structuring element to bridge the edge breaks and eliminate small noise points at the same time; then remove isolated pixel points through the opening operation (elliptical structuring element), and fill the internal voids of the contour through the closing operation (rectangular structuring element) to form a smooth and continuous boundary; the finally output optimized mask not only retains the original morphological features of the fancy words, but also eliminates the edge burrs and discrete noises generated during the detection process, enabling the subsequent ratio correction processing to be performed based on more accurate regional positioning data, effectively avoiding correction errors caused by incomplete contours.

[0166] This solution effectively eliminates the noise interference at the edge of the fancy word area through morphological operations, improves the continuity and smoothness of the contour boundary, and at the same time maintains the original structural features of the fancy words, providing more accurate regional positioning data for subsequent video ratio correction.

[0167] Figure 2 In the shown process, first, narrow the processing range through preset area segmentation to reduce the calculation amount; then, based on the collaborative processing of edge detection and contour extraction, effectively retain the structural features of the fancy words; and exclude interfering contours through multi-dimensional geometric feature screening to ensure the specificity of fancy word recognition; the finally output standardized position data not only contains accurate boundary coordinates, but also provides confidence evaluation, providing a reliable basis for subsequent video ratio correction. While ensuring the detection accuracy, the processing speed is improved.

[0168] See Figure 3 , which is a flowchart of an embodiment of another video processing method provided by an embodiment of the present application. This Figure 3 shown process is based on the above Figure 1 shown process, and describes how to analyze and process the fancy word area data of all video frames to obtain a set of target fancy word areas that meet the preset conditions. As Figure 3 shown, this process may include the following steps:

[0169] S301. Calculate the average position of each subtitle region within the sliding window, and remove subtitle regions with a position deviation exceeding the first threshold to obtain a preliminary screening result.

[0170] In the embodiments of the present application, a moving average calculation is performed on the detected subtitle regions through Formula 1. Then, a first threshold (\(\epsilon\)) is set through Formula 2 to filter out subtitle regions with large fluctuations:

[0171]

[0172] Where The moving average value at time \(t\) and position \(p\). Here, \(t\) represents the time frame (or frame number), and \(p\) represents the position of the pixel. \(R\) i,p : The original pixel value at time frame \(i\) and position \(p\). \(w\): The size of the moving window, representing the number of time frames used for calculating the moving average. \(e\): The first threshold, used to filter out subtitle regions with large fluctuations. This threshold determines the maximum allowed deviation. \(|R_{t,p} - R_{t,p}|\): The absolute difference between the original pixel value and the moving average value.

[0173] Formula 2 is used to calculate the absolute difference between the original pixel value and the moving average value of each pixel \(p\) at time \(t\), and determine whether to retain the pixel by comparing this difference with the threshold, that is, if this difference is less than the set threshold \(e\), the pixel is retained; otherwise, it is considered that the pixel is an outlier with large fluctuations and is filtered out. In this way, pixels with large fluctuations caused by noise or other factors can be removed, thereby obtaining a more stable subtitle region.

[0174] S302. Calculate the stability score of each subtitle region in the preliminary screening result, and remove subtitle regions with a stability score lower than the second threshold to obtain a set of target subtitle regions.

[0175] In the embodiments of the present application, the stability score is calculated through the appearance frequency and consistency score of each subtitle region, specifically including the following steps: count the appearance frequency of each subtitle region in the preliminary screening result, where the appearance frequency refers to the proportion of the subtitle region in the total number of frames; calculate the average value of the shape similarity scores between subtitle regions in adjacent video frames in the preliminary screening result to obtain the consistency score; for each subtitle region in the preliminary screening result, perform a weighted summation calculation on the appearance frequency and consistency score of the subtitle region to obtain the stability score of the subtitle region.

[0176] In this solution, first, calculate the shape similarity ShapeSim(R t,p ,R t+1,p ) through Formula 3:

[0177]

[0178] where R t,p and R t+1,p are two regions, At is the area of R t,p , At+1 is the area of R t+1,p , and Aintersection is the intersection area of the two.

[0179] Then, calculate the consistency score Consistency(R p ) through Equation 4:

[0180]

[0181] Consistency(Rp): The finally calculated consistency score, representing the average of the shape similarities of a specific region Rp between adjacent time frames over all time frames. N: Represents the total number of video frames. For example, if the video has 100 frames, then N = 100. Rt,p represents the region (or shape) at time frame t and position p. Rt+1,p represents the region (or shape) at the next time frame t + 1 and the same position p.

[0182] The consistency calculation process of Equation 4 is as follows: For each pair of adjacent time frames t and t + 1, calculate the shape similarity between regions Rt.p and Rt+1,p. After summing up the shape similarities of all adjacent time frames, divide by the total number of frames N to obtain the consistency score of this region.

[0183] Next, calculate the frequency of occurrence Frequency(Rp) through Equation 5:

[0184]

[0185] Frequency(Rp) is the finally calculated frequency value, representing the frequency of occurrence of a specific region Rp over all time frames. δ(Rt,p≠0) is an indicator function used to determine whether Rt,p is empty. If Rt,p≠0, that is, the region exists in time frame t, then δ(Rt,p≠0) = 1; if Rt,p = 0, that is, the region does not exist in time frame t, then δ(Rt,p≠0) = 0.

[0186] Finally, select the regions that meet the frequency and consistency thresholds through the following formula:

[0187] [Select R p where Frequency(R p )>ψ1 and Consistency(R p )>

[0188] ψ2].

[0189] In this solution, first, temporary interference elements (such as floating advertisements or transient lens noise) are effectively filtered by counting the occurrence frequency (such as accounting for ≥ 80% in the total number of frames); second, the morphological consistency of subtitles on the time axis is ensured based on the average value of the shape similarity of adjacent frames (such as the Hu moment difference ≤ 0.1); finally, weighted summation (such as frequency weight 0.6 + consistency weight 0.4) is used to comprehensively quantify the spatio-temporal stability, which not only avoids misjudgment by a single index (such as high-frequency but deformed subtitles), but also adapts to different types of subtitles (static logos have a weight bias towards frequency, and dynamic corner marks have a weight bias towards consistency). This effectively improves the continuous tracking accuracy of key subtitles, while reducing the false detection rate of flickering, providing high-stability regional positioning data for video scale correction.

[0190] Figure 3 In the shown process, first, based on the average position of the sliding window, abnormal offset regions caused by video jitter or instantaneous occlusion are effectively filtered (such as misdetected targets with a position mutation exceeding ±5 pixels), ensuring the spatial continuity of subtitles; second, by comprehensively evaluating the occurrence frequency and shape consistency of subtitles in the time sequence through stability scoring, temporary interference elements (such as floating advertisements or lens glare) are further removed, so that the final output set of target subtitle regions simultaneously meets the dual criteria of spatial position stability and time continuity. This hierarchical screening mechanism reduces the false detection rate and is particularly suitable for scenarios with dynamic noise such as live streaming media, providing high-confidence input data for subsequent video scale correction.

[0191] Based on the same technical concept, an embodiment of the present application also provides a video processing device, as Figure 4 shown, the device includes:

[0192] An acquisition module 41, configured to acquire a video to be processed and split the video to be processed into a plurality of video frames to obtain a video frame sequence;

[0193] A detection module 42, configured to detect subtitle regions at four corners of each video frame in the video frame sequence to obtain subtitle region data including position information of the subtitle regions;

[0194] An analysis module 43, configured to analyze and process the subtitle region data of all video frames to obtain a set of target subtitle regions meeting preset conditions;

[0195] A calculation module 44, configured to perform a comparison calculation based on the size information of each target subtitle region in the set of target subtitle regions and a standard ratio to obtain a video scale correction parameter;

[0196] A correction module 45, configured to perform geometric transformation processing on the video frame sequence according to the video scale correction parameter to obtain a corrected target video.

[0197] In a possible implementation, the detection module is specifically configured to:

[0198] Perform image segmentation processing on four preset corner regions of the video frame to obtain four regions to be processed;

[0199] For each region to be processed, perform edge detection processing on the region to be processed to obtain edge feature data;

[0200] Perform contour extraction processing on the edge feature data to obtain a candidate contour set;

[0201] Perform feature screening processing on the candidate contour set to obtain effective caption contours;

[0202] Perform position calculation processing based on the effective caption contours to obtain caption region data including caption region position information.

[0203] In a possible implementation, the detection module is further configured to:

[0204] For each contour in the candidate contour set, perform geometric feature extraction processing on the contour to obtain contour feature data, where the contour feature data includes at least one of the following features: the ratio of the contour area to a preset area threshold, the width-to-height ratio of the circumscribed rectangle of the contour, and the complexity eigenvalue of the contour edge;

[0205] Perform matching degree calculation processing on the contour feature data and a preset caption feature model to obtain a matching degree score;

[0206] Determine the contours with corresponding matching degree scores greater than a preset threshold as effective caption contours.

[0207] In a possible implementation, the analysis module is specifically configured to:

[0208] Calculate the average position of each caption region in the sliding window, and eliminate the caption regions with position offsets exceeding a first threshold to obtain a preliminary screening result;

[0209] Calculate the stability scores of the caption regions in the preliminary screening result, and eliminate the caption regions with stability scores lower than a second threshold to obtain a target caption region set.

[0210] In a possible implementation, the analysis module is further configured to:

[0211] Count the occurrence frequencies of the caption regions in the preliminary screening result, where the occurrence frequency refers to the proportion of the caption regions in the total number of frames;

[0212] Calculate the average value of the shape similarity scores of the caption regions between adjacent video frames in the preliminary screening results to obtain a consistency score;

[0213] For each caption region in the preliminary screening results, perform a weighted summation calculation on the occurrence frequency and the consistency score of the caption region to obtain the stability score of the caption region.

[0214] In a possible implementation manner, the calculation module is specifically configured to:

[0215] Perform statistical analysis processing on the size information of each target caption region in the target caption region set to obtain statistical size data;

[0216] Perform a difference calculation on the statistical size data and a standard ratio to obtain a ratio difference parameter;

[0217] Perform optimization and adjustment processing on the ratio difference parameter to obtain a video ratio correction parameter.

[0218] In a possible implementation manner, the correction module is specifically configured to:

[0219] Perform frame-by-frame scaling processing on the video frame sequence according to the video ratio correction parameter to obtain a video frame sequence after ratio adjustment;

[0220] Perform edge compensation processing on the video frame sequence after ratio adjustment to obtain a video frame sequence after edge compensation;

[0221] Perform encoding and synthesis processing on the video frame sequence after edge compensation to obtain a corrected target video.

[0222] Based on the same technical concept, an embodiment of the present application further provides an electronic device, as Figure 5 shown, including a processor 111, a communication interface 112, a memory 113, and a communication bus 114, where the processor 111, the communication interface 112, and the memory 113 complete communication with each other through the communication bus 114,

[0223] The memory 113 is used to store a computer program;

[0224] When the processor 111 is used to execute the program stored in the memory 113, the following steps are implemented:

[0225] Obtain a video to be processed, and split the video to be processed into a plurality of video frames to obtain a video frame sequence;

[0226] For each video frame in the video frame sequence, detect the caption regions at the four corners of the video frame to obtain caption region data including the position information of the caption regions;

[0227] Analyze and process the caption area data of all video frames to obtain a set of target caption areas that meet the preset conditions;

[0228] Based on the comparison calculation of the size information of each target caption area in the set of target caption areas with the standard ratio, obtain a video ratio correction parameter;

[0229] Perform geometric transformation processing on the video frame sequence according to the video ratio correction parameter to obtain a corrected target video.

[0230] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used in the figure to represent it, but it does not mean that there is only one bus or one type of bus.

[0231] The communication interface is used for communication between the above electronic device and other devices.

[0232] The memory can include a Random Access Memory (RAM), and can also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory can also be at least one storage device located far from the aforementioned processor.

[0233] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0234] In another embodiment provided by the present application, a computer-readable storage medium is also provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above video processing methods are implemented.

[0235] In another embodiment provided by the present application, a computer program product including instructions is further provided. When it runs on a computer, it causes the computer to execute any one of the video processing methods in the above embodiments.

[0236] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0237] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution or the part that contributes to the related technology can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0238] It should be understood that the terms used herein are only for the purpose of describing specific example embodiments and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" as used herein may also include the plural forms. The terms "include", "comprise", "contain", and "have" are inclusive and thus specify the presence of the stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be executed in the specific order described or illustrated, unless the execution order is clearly indicated. It should also be understood that additional or alternative steps can be used.

[0239] The above description is only the specific implementation manners of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A video processing method, characterized in that, The method comprises: Obtain a video to be processed, and split the video to be processed into a plurality of video frames to obtain a video frame sequence; For each video frame in the video frame sequence, detecting the flower character regions at four corners of the video frame to obtain flower character region data including the position information of the flower character region; Analyze and process the flower character area data of all video frames to obtain a target flower character area set that meets preset conditions; Comparing and calculating the size information of each target flower character area in the target flower character area set with the standard ratio, a video ratio correction parameter is obtained; The video frame sequence is geometrically transformed according to the video ratio correction parameter to obtain a corrected target video.

2. The method according to claim 1, wherein The detecting of the flower character regions at the four corners of the video frame to obtain flower character region data containing the position information of the flower character regions includes: Performing image segmentation processing on four preset corner areas of the video frame respectively to obtain four areas to be processed; For each area to be processed, edge detection processing is performed on the area to be processed to obtain edge feature data; Performing contour extraction processing on the edge feature data to obtain a candidate contour set; Performing feature screening on the candidate contour set to obtain a valid flower character contour; Position calculation processing is performed based on the effective flower character outline to obtain flower character area data containing flower character area position information.

3. The method according to claim 2, characterized in that, The step of performing feature screening on the candidate contour set to obtain a valid flower character contour comprises: For each contour in the candidate contour set, a geometric feature extraction process is performed on the contour to obtain contour feature data, wherein the contour feature data includes at least one of the following features: a ratio of a contour area to a preset area threshold, a width-to-height ratio of a contour circumscribed rectangle, and a complexity feature value of a contour edge; Calculate the matching degree of the outline feature data and the preset flower character feature model to obtain a matching degree score; The contours whose corresponding matching scores are greater than a preset threshold are determined as valid flower character contours.

4. The method according to claim 1, characterized in that The step of analyzing and processing the flower character region data of all video frames to obtain a target flower character region set that meets preset conditions includes: Calculate the average position of each flower character region in the sliding window, and remove the flower character regions whose position deviation exceeds the first threshold to obtain a preliminary screening result; The stability score of each flower character region in the preliminary screening result is calculated, and the flower character regions with stability scores lower than the second threshold are eliminated to obtain a target flower character region set.

5. The method according to claim 4, characterized in that, The calculating of the stability score of each flower character region in the preliminary screening result comprises: Counting the occurrence frequency of each flower character region in the preliminary screening result, wherein the occurrence frequency refers to the appearance ratio of the flower character region in the total number of frames; Calculating an average of shape similarity scores of the squiggle regions between adjacent video frames in the preliminary screening results to obtain a consistency score; For each flower character region in the preliminary screening result, a weighted sum calculation process is performed on the occurrence frequency and consistency score of the flower character region to obtain a stability score of the flower character region.

6. The method according to claim 1, wherein The step of comparing and calculating the size information of each target flower character area in the target flower character area set with the standard ratio to obtain the video ratio correction parameter includes: Statistically analyze the size information of each target subtitle area in the target subtitle area set to obtain statistical size data; Perform a difference calculation process on the statistical size data and a standard ratio to obtain a ratio difference parameter; Perform an optimization and adjustment process on the ratio difference parameter to obtain a video ratio correction parameter.

7. The method according to claim 1, characterized in that The geometric transformation process of the video frame sequence according to the video ratio correction parameter to obtain a corrected target video includes: Perform a frame-by-frame scaling process on the video frame sequence according to the video ratio correction parameter to obtain a video frame sequence with adjusted ratio; Perform an edge compensation process on the video frame sequence with adjusted ratio to obtain a video frame sequence with edge compensation; Perform an encoding and synthesis process on the video frame sequence with edge compensation to obtain a corrected target video.

8. A video processing device, characterized in that, The device includes: An acquisition module, configured to acquire a video to be processed and split the video to be processed into a plurality of video frames to obtain a video frame sequence; A detection module, configured to detect subtitle areas at four corners of each video frame in the video frame sequence to obtain subtitle area data including subtitle area position information; An analysis module, configured to analyze and process subtitle area data of all video frames to obtain a set of target subtitle areas meeting preset conditions; A calculation module, configured to perform a comparison calculation based on the size information of each target subtitle area in the set of target subtitle areas and a standard ratio to obtain a video ratio correction parameter; A correction module, configured to perform a geometric transformation process on the video frame sequence according to the video ratio correction parameter to obtain a corrected target video.

9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used to store a computer program; The processor, when executing the program stored on the memory, implements the video processing method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the video processing method according to any one of claims 1-7.

Citation Information

Cited By

  • Video generation method and electronic equipment

    CN121000921A

  • Video content area detection method and device, electronic equipment and storage medium

    CN121661570A