Video correction method and device based on subtitles, storage medium and electronic equipment
By disassembling the video frames and text area detection, combining the target clustering algorithm and font width and height calculation, the video proportions are automatically corrected, solving the problems of low processing efficiency and insufficient accuracy in the prior art, and improving the efficiency and adaptability of video processing.
Patent Information
- Application Number
- CN202510235778.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art has low processing efficiency and insufficient accuracy in video proportion adjustment, especially in large-scale processing, which is difficult to achieve intelligence and efficiency.
By disassembling the original video, detecting text area and trajectory feature analysis, the target clustering algorithm is used to determine the subtitle area, and the corrected proportion of the video is calculated based on the font width and height ratio of the subtitle area, geometric transformation and video encoding are performed to generate the target video.
It realizes automated video proportion correction, improves processing efficiency and accuracy, is highly adaptable, can be applied to videos from different sources, and optimizes the viewing experience.
Smart Images

Figure CN120050462A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a video correction method, device, storage medium, and electronic device based on subtitles. Background Art
[0002] In scenarios such as online video platforms, social media, video production, surveillance analysis, and online education, the problem of inaccurate video aspect ratios is common and affects the viewing experience. Currently, video aspect ratio adjustment mainly relies on manual editing, preset templates, or simple cropping and filling. However, these methods have problems such as low efficiency and insufficient accuracy when dealing with large-scale processing. The difficulty of automated aspect ratio adjustment lies in intelligently identifying video content and determining the appropriate aspect ratio, which involves technologies such as computer vision, image processing, and machine learning. At the same time, different scenarios have different requirements for video aspect ratios, and there is a lack of a unified standard, making it difficult to achieve a general solution. In addition, the computational performance and efficiency of algorithms pose high requirements for real-time processing and large-scale batch processing. Due to the diverse sources of videos, the shooting devices and parameters are different, and automated methods need to adapt to complex and changing data environments. Therefore, to achieve an intelligent, efficient, and general video aspect ratio adjustment technology, it is necessary to combine deep learning, image recognition, and adaptive algorithms to improve processing accuracy and optimize computing resources to meet the needs of various application scenarios. Summary of the Invention
[0003] This application provides a video correction method, device, storage medium, and electronic device based on subtitles to solve the technical problem of low processing efficiency of traditional video aspect ratio adjustment methods.
[0004] In a first aspect, this application provides a video correction method based on subtitles, including: performing frame splitting operations on the original video according to a preset frequency to obtain an image sequence; performing text region detection on each frame image of the image sequence to obtain the corresponding text regions and their trajectory features, where the trajectory features include duration, position coordinates, and region area; based on a target clustering algorithm and the trajectory features, determining a subtitle region from all the text regions, where the subtitle region is a text region with the longest duration and the smallest change in position among all the text regions; calculating the correction ratio of the original video according to the font width-to-height ratio in the subtitle region and the standard font width-to-height ratio; performing geometric transformation on each frame image of the image sequence according to the correction ratio to obtain a processed image sequence, and performing video encoding on the processed image sequence to obtain a target video.
[0005] In a second aspect, the present application provides a video correction device based on subtitles, including: a first processing module, configured to perform frame splitting operations on an original video according to a preset frequency to obtain an image sequence; a second processing module, configured to perform text region detection on each frame image of the above image sequence to obtain corresponding text regions and their trajectory features, where the above trajectory features include duration, position coordinates, and region area; a first determination module, configured to determine a subtitle region from all the above text regions based on a target clustering algorithm and the above trajectory features, where the subtitle region is a text region with the longest duration and the smallest position change amplitude among all the above text regions; a calculation module, configured to calculate a correction ratio of the above original video according to a font width-to-height ratio in the subtitle region and a standard font width-to-height ratio; a correction module, configured to perform geometric transformation on each frame image of the above image sequence according to the above correction ratio to obtain a processed image sequence, and perform video encoding on the above processed image sequence to obtain a target video.
[0006] As an optional example, the above device further includes: a second determination module, configured to determine a duration threshold and a position change amplitude threshold before determining the subtitle region from all the above text regions based on the target clustering algorithm and the above trajectory features; a third processing module, configured to determine each text region among all the above text regions as the current text region, and perform the following operations on the current text region: obtain the current duration and the current position change amplitude of the current text region according to the trajectory features of the current text region; filter out the current text region when the current duration is less than the duration threshold or the current position change amplitude is greater than the position change amplitude threshold.
[0007] As an optional example, the above first determination module includes: a clustering unit, configured to cluster all the above text regions based on the target clustering algorithm to group text regions with similar trajectory features into one cluster; a first determination unit, configured to determine the cluster with the longest average duration and the smallest average position change amplitude among all the above clusters as the target cluster; a second determination unit, configured to determine any one text region in the target cluster as the subtitle region.
[0008] As an optional example, the above calculation module includes: an acquisition unit, configured to acquire the font width-to-height ratio in the subtitle region and the standard font width-to-height ratio; a calculation unit, configured to calculate a ratio of the font width-to-height ratio in the subtitle region to the standard font width-to-height ratio to obtain the correction ratio of the above original video.
[0009] As an alternative example, the above correction module includes: a correction unit configured to perform proportional scaling on the height and width of each frame of the above image sequence in parallel according to the above correction ratio to obtain a processed image sequence, and perform edge padding processing on each frame of the processed image sequence.
[0010] In a third aspect, the present application provides a storage medium storing a computer program, wherein the computer program, when run by a processor, executes the above-mentioned subtitle-based video correction method.
[0011] In a fourth aspect, the present application further provides an electronic device including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the above-mentioned subtitle-based video correction method through the computer program.
[0012] The above technical solutions provided by the embodiments of the present application have the following advantages compared with the prior art:
[0013] The present application adopts the method of splitting frames of the original video according to a preset frequency to obtain an image sequence; detecting text regions for each frame of the image sequence to obtain corresponding text regions and their trajectory features, where the trajectory features include duration, position coordinates, and region area; based on a target clustering algorithm and the above trajectory features, determining a subtitle region from all the above text regions, where the subtitle region is a text region with the longest duration and the smallest position change amplitude among all the above text regions; calculating a correction ratio of the original video according to the font width-to-height ratio in the subtitle region and a standard font width-to-height ratio; performing geometric transformation on each frame of the image sequence according to the correction ratio to obtain a processed image sequence, and performing video encoding on the processed image sequence to obtain a target video. Since in the above method, an image sequence is extracted by frame splitting, text regions are recognized, and a subtitle region is selected by combining trajectory features. The correction ratio of the video is calculated based on the font width-to-height ratio of the subtitle, and geometric transformation is performed on the image sequence, finally generating a target video with the correct ratio. Thus, automatic video ratio correction is achieved, reducing the workload of manual adjustment, improving accuracy, processing efficiency, and adaptability, ensuring that videos from different sources can be correctly displayed, optimizing the viewing experience, and further solving the technical problem of low processing efficiency of traditional video ratio adjustment methods. Description of the Drawings
[0014] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.
[0015] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0016] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the drawings represent similar elements, unless otherwise stated, the drawings in the figures do not constitute a proportional limitation.
[0017] Figure 1 is a flowchart of an optional subtitle-based video correction method according to an embodiment of the present application;
[0018] Figure 2 is an overall implementation flowchart of an optional subtitle-based video correction method according to an embodiment of the present application;
[0019] Figure 3 is a schematic structural diagram of an optional subtitle-based video correction device according to an embodiment of the present application;
[0020] Figure 4 is a schematic diagram of an optional electronic device according to an embodiment of the present application. Detailed implementation manners
[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.
[0022] The following disclosure provides many different embodiments or examples for implementing different structures of the present application. To simplify the disclosure of the present application, the components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present application. In addition, the present application may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed.
[0023] According to the first aspect of the embodiments of the present application, a subtitle-based video correction method is provided. Optionally, as Figure 1 shown, the above method includes:
[0024] S102. Perform frame splitting on the original video according to a preset frequency to obtain an image sequence;
[0025] S104. Detect the text regions for each frame of the image sequence to obtain the corresponding text regions and their trajectory features, where the trajectory features include duration, position coordinates, and region area;
[0026] S106. Based on the target clustering algorithm and the trajectory features, determine the subtitle region from all the text regions, where the subtitle region is a text region with the longest duration and the smallest position change amplitude among all the text regions;
[0027] S108. Calculate the correction ratio of the original video according to the font width-to-height ratio in the subtitle region and the standard font width-to-height ratio;
[0028] S110. Perform geometric transformation on each frame of the image sequence according to the correction ratio to obtain a processed image sequence, and perform video encoding on the processed image sequence to obtain the target video.
[0029] Optionally, in this embodiment, the stable characteristics of the subtitle area are utilized to automatically correct the video ratio, covering key steps such as frame splitting, text detection, trajectory analysis, ratio calculation, geometric transformation, and video reconstruction. This method can be applied to scenarios such as online videos, UGC (user-generated content), surveillance videos, and online education, and can be extended to requirements such as multi-language subtitles, different shooting devices, and real-time streaming media processing. Specifically, through common video processing tools, the sampling frequency (such as 10 frames per second) is set, and image frames are extracted at equal intervals from the original video to generate a series of static frame images for subsequent analysis. OCR (Optical Character Recognition) or deep learning models are used to detect the text areas in the images, and the relevant information of each text area, that is, the trajectory features, is recorded, including: duration (the appearance time of the area in multiple frames), position coordinates (the specific position of the subtitle area), and area (the pixel area occupied by the text area). Since the video may contain multiple text areas (such as subtitles, watermarks, annotations, advertisement texts, etc.), the subtitle area needs to be screened out based on the target clustering algorithm and trajectory features. By analyzing all text areas through the clustering algorithm, the text area with the longest duration and the smallest position change can be finally determined and identified as the subtitle area. This is because subtitles usually appear stably, while other texts (such as advertisement slogans) may be transient, and subtitles are usually fixed at the bottom or top of the video, while the positions of non-fixed texts (such as moving watermarks) change greatly. By analyzing the font aspect ratio of the subtitle area, the aspect ratio of the actual video is calculated, compared with the standard font aspect ratio, and the possible distortion or stretching of the original video is deduced, and the correction ratio is calculated, that is, how to adjust the aspect ratio of the video to restore normal display. According to the correction ratio, geometric transformations (such as scaling, cropping, stretching) are performed on each frame image to adjust to the correct ratio, and the processed image sequence is re-encoded to generate the final target video.
[0030] Optionally, in this embodiment, by automatically detecting the subtitle area and calculating the correction ratio, the accurate correction of the video ratio is achieved, avoiding the cumbersome and error of manual adjustment. Using OCR recognition and clustering algorithms ensures accurate positioning of the subtitle area, improving applicability and processing efficiency. Adapting to different video formats from different sources, ensuring that the picture is not distorted, optimizing the viewing experience, and being applicable to various application scenarios such as online video platforms, video editing, and surveillance analysis at the same time.
[0031] As an optional example, before determining the subtitle area from all text areas based on the target clustering algorithm and trajectory features, the above method further includes:
[0032] Determining a duration threshold and a position change amplitude threshold;
[0033] Identify each text region in all text regions as the current text region, and perform the following operations on the current text region:
[0034] According to the trajectory characteristics of the current text region, obtain the current duration and the current position change amplitude of the current text region;
[0035] In the case where the current duration is less than the duration threshold or the current position change amplitude is greater than the position change amplitude threshold, filter out the current text region.
[0036] Optionally, in this embodiment, before determining the subtitle region based on the target clustering algorithm and trajectory characteristics, pre-screening based on thresholds is added to improve the recognition accuracy and calculation efficiency of the subtitle region. Specifically, set the duration threshold and the position change amplitude threshold. The duration threshold is used to measure the stability of the subtitle region to avoid interference texts that appear for a short time (such as advertising slogans, bullet screens, etc.), and the position change amplitude threshold is used to filter out texts that move with the screen (such as watermarks, scrolling subtitles) to ensure the stability of the subtitle region. Process each text region in parallel. Taking the current text region as an example, extract the trajectory characteristics and duration of the current text region. If the duration of the current text region is less than the set duration threshold, it means that the text region may appear temporarily, such as short prompt messages, advertising texts, etc., and it is directly screened out. If the current position change amplitude is greater than the set position change amplitude threshold, it means that the displacement of the text region is large, such as watermarks, moving annotations, scrolling bullet screens, etc. in the video, which does not conform to the stable characteristics of the subtitle, and it is also screened out. Only text regions that meet the conditions of "long enough duration" and "small position change" can enter the subsequent target clustering algorithm. After threshold screening, the remaining text regions basically conform to the characteristics of the subtitle. These text regions are further analyzed by the target clustering algorithm to finally determine the subtitle region, and then the video correction ratio is calculated based on the font width-to-height ratio of the subtitle region and the ratio adjustment is performed.
[0037] Optionally, in this embodiment, the misrecognition interference is effectively reduced, the subsequent subtitle region detection is optimized, and the final video ratio correction is made more accurate.
[0038] As an optional example, determining the subtitle region from all text regions based on the target clustering algorithm and trajectory characteristics includes:
[0039] Cluster all text regions based on the target clustering algorithm to group text regions with similar trajectory characteristics into a cluster;
[0040] Determine the target cluster as the cluster with the longest average duration and the smallest average position change amplitude among all clusters;
[0041] Determine any text region in the target cluster as the subtitle region.
[0042] Optionally, in this embodiment, first, based on the previously filtered text regions, apply a target clustering algorithm. This algorithm classifies the text regions according to trajectory features (such as duration, position change amplitude, etc.) to ensure that the text regions in each cluster have similar trajectory features. During the clustering process, text regions with similar time features and position change features will be grouped into the same cluster. For example, stable and long-lasting text regions will be placed together with other similar regions. Among all the clusters, select the cluster with the longest average duration and the smallest average position change amplitude as the target cluster. The text regions in this cluster exist for a longer time and have less regional change, and are more likely to be the subtitles in the video rather than short-lived advertisements or prompt messages. Finally, select any text region in the target cluster as the subtitle region. By using the target clustering algorithm and trajectory features to screen the subtitle region, the accuracy and stability of subtitle region recognition are effectively improved, the possibility of misidentifying interference regions is reduced, and the quality of subsequent video processing (such as scale adjustment, picture optimization, etc.) is enhanced.
[0043] As an optional example, calculating the correction ratio of the original video according to the font width-to-height ratio in the subtitle region and the standard font width-to-height ratio includes:
[0044] Obtain the font width-to-height ratio in the subtitle region and the standard font width-to-height ratio;
[0045] Calculate the ratio of the font width-to-height ratio in the subtitle region to the standard font width-to-height ratio to obtain the correction ratio of the original video.
[0046] Optionally, in this embodiment, the subtitle area has been determined. For this subtitle area, the width and height information of the font therein is extracted, and the font width-to-height ratio is calculated, which refers to the ratio between the width and height of the font in this area and is usually used to describe the shape of the font. For example, if the font width in the subtitle area is 100 pixels and the font height is 50 pixels, then the font width-to-height ratio of this subtitle area is 2:1. The standard font width-to-height ratio is a predefined standard value used to compare with the font in the actual subtitle area. This standard value is determined according to common video production specifications or subtitle styles and represents an ideal or standardized font ratio. By comparing the font width-to-height ratio of the subtitle area with the preset standard font width-to-height ratio and calculating the ratio between them, the correction ratio is obtained. Suppose the font width-to-height ratio of the subtitle area is 2:1, while the standard font width-to-height ratio is 1.5:1, then the correction ratio is (2:1) / (1.5:1) = 1.33. This calculation result provides a scaling factor, that is, a corresponding geometric transformation needs to be performed on the entire original video so that the font width-to-height ratio of the subtitle area in the video is close to the standard font width-to-height ratio. Through this ratio, the overall ratio of the video can be appropriately adjusted to meet the standardization requirements. By correcting the font width-to-height ratio, the standardization, readability, and aesthetics of the video subtitles are ensured, effectively improving the quality of the video content. Based on the automatic analysis and calculation of the video subtitle area, the workload of manual adjustment is reduced, and the processing efficiency of video correction is improved.
[0047] As an optional example, according to the correction ratio, a geometric transformation is performed on each frame of the image sequence, and the obtained processed image sequence includes:
[0048] According to the correction ratio, the height and width of each frame of the image sequence are scaled proportionally in parallel to obtain a processed image sequence, and edge padding processing is performed on each frame of the processed image sequence.
[0049] Optionally, in this embodiment, according to the correction ratio calculated previously, scaling operations for the width and height of each frame of the image are performed. The correction ratio can be understood as a scaling factor that determines the overall size adjustment of the video. This ratio affects the width and height of the image to ensure that the font ratio in the subtitle area meets the standard. For each frame of the image, parallel scaling is performed according to the correction ratio, that is, the width and height of the image are scaled proportionally at the same time, maintaining the original aspect ratio of the image. For example, if the correction ratio is 1.33, the original width of the image is 1000 pixels, and the height is 500 pixels, then the width of the processed image is 1330 pixels, and the height is 665 pixels. After proportional scaling, the size of the image may change (width and height change), but there may be requirements for the original size of the video or the video display area needs to maintain a fixed size. At this time, if the size of the scaled image does not exactly match the size of the target display area, edge padding processing is required. Edge padding can add additional pixel areas around the image, usually filled with the background color or transparent color. This can ensure that the image adapts to the target size without distortion and avoid distortion or incompleteness of the picture content due to stretching or cropping. For example, if the target size is 1280x720 pixels, and the size of the image after proportional scaling is 1330x665 pixels, then padding pixels need to be added to the right and bottom of the image to fill to the target size of 1280x720 pixels. By performing proportional scaling and edge padding processing on the image sequence according to the correction ratio, the adaptability, visual effect, and playback consistency of the video content can be significantly improved, ensuring consistent video playback effects on different devices and platforms. At the same time, manual adjustment work is reduced, and the processing efficiency is improved.
[0050] Illustrated with an example, this application relates to a video correction method based on subtitles. The overall implementation process is as Figure 2 shown, including video frame splitting, text detection, subtitle area discrimination, time-domain trajectory clustering, subtitle area extraction, ratio calculation, and video correction. Specifically:
[0051] (1) Video frame splitting and image sequence generation: The video file is split into frames according to a preset frequency, that is, the original video is decomposed into a series of individual image frames, which will become the basic units for subsequent processing.
[0052] (2) Text area detection and trajectory feature extraction: For each frame of the image, text area detection is performed to identify all text areas in the image. For each text area, its trajectory features are extracted, including: duration (the time length that the text area is displayed in the video), position coordinates (the position of the text area in the image), and area (the area of the text area).
[0053] (3) Subtitle area screening: Set the duration threshold and the position change amplitude threshold to filter out text areas that do not meet the conditions. For each text area, if its duration is less than the duration threshold or the position change amplitude is greater than the position change amplitude threshold, it is excluded.
[0054] (4) Subtitle area determination: Use the target clustering algorithm to cluster all text areas, group text areas with similar trajectory characteristics into one cluster, and from all clusters, select the cluster with the longest duration and the smallest position change amplitude as the target cluster. The text areas included in this cluster are considered subtitle areas. Finally, select any one text area from the target cluster as the subtitle area for subsequent video correction and processing.
[0055] (5) Font ratio and correction ratio calculation: Calculate the correction ratio according to the font width-to-height ratio in the subtitle area and the standard font width-to-height ratio. The specific steps are as follows: Obtain the font width-to-height ratio and the standard font width-to-height ratio of the subtitle area; calculate the ratio of the two to obtain the correction ratio.
[0056] (6) Geometric transformation and processing of the image sequence: According to the correction ratio, perform parallel scaling on the width and height of each frame of the image to ensure that the size of the video content meets the standard. Fill in the four sides of the image to ensure that the scaled image fits the target display area. The filling usually uses the background color or the transparent color to maintain the integrity of the video.
[0057] (7) Video encoding and output: Perform video encoding on the processed image sequence, recombine it into the target video format, and generate the final processed video.
[0058] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0059] According to another aspect of the embodiments of the present application, there is also provided a subtitle-based video correction device, as Figure 3 shown, including:
[0060] The first processing module 302 is used to perform frame splitting on the original video according to a preset frequency to obtain an image sequence;
[0061] The second processing module 304 is configured to perform text region detection on each frame of the image sequence to obtain the corresponding text region and its trajectory features, where the trajectory features include duration, position coordinates, and region area;
[0062] The first determination module 306 is configured to determine the subtitle region from all text regions based on the target clustering algorithm and the trajectory features, where the subtitle region is a text region with the longest duration and the smallest position change amplitude among all text regions;
[0063] The calculation module 308 is configured to calculate the correction ratio of the original video according to the font width-to-height ratio and the standard font width-to-height ratio in the subtitle region;
[0064] The correction module 310 is configured to perform geometric transformation on each frame of the image sequence according to the correction ratio to obtain the processed image sequence, and perform video encoding on the processed image sequence to obtain the target video.
[0065] It should be noted that the first processing module 302 in this embodiment may be configured to execute step S102 in the embodiment of the present application, the second processing module 304 in this embodiment may be configured to execute step S104 in the embodiment of the present application, the first determination module 306 in this embodiment may be configured to execute step S106 in the embodiment of the present application, the calculation module 308 in this embodiment may be configured to execute step S108 in the embodiment of the present application, and the correction module 310 in this embodiment may be configured to execute step S110 in the embodiment of the present application.
[0066] As an optional example, the above device further includes:
[0067] The second determination module is configured to determine the duration threshold and the position change amplitude threshold before determining the subtitle region from all text regions based on the target clustering algorithm and the trajectory features;
[0068] The third processing module is configured to determine each text region in all text regions as the current text region, and perform the following operations on the current text region:
[0069] Obtain the current duration and the current position change amplitude of the current text region according to the trajectory features of the current text region;
[0070] Filter out the current text region when the current duration is less than the duration threshold or the current position change amplitude is greater than the position change amplitude threshold.
[0071] As an optional example, the first determination module includes:
[0072] A clustering unit, configured to cluster all text regions based on a target clustering algorithm, so as to classify text regions with similar trajectory features into one cluster;
[0073] A first determination unit, configured to determine the cluster with the longest average duration and the smallest average position change amplitude among all clusters as the target cluster;
[0074] A second determination unit, configured to determine any one text region in the target cluster as the subtitle region.
[0075] As an optional example, the calculation module includes:
[0076] An acquisition unit, configured to acquire the font width-to-height ratio and the standard font width-to-height ratio in the subtitle region;
[0077] A calculation unit, configured to calculate the ratio of the font width-to-height ratio in the subtitle region to the standard font width-to-height ratio, so as to obtain the correction ratio of the original video.
[0078] As an optional example, the correction module includes:
[0079] A correction unit, configured to perform proportional scaling on the height and width of each frame of the image sequence in parallel according to the correction ratio, so as to obtain a processed image sequence, and perform edge padding processing on each frame of the processed image sequence.
[0080] For other examples of this embodiment, please refer to the above examples, which will not be elaborated here.
[0081] Figure 4 It is a schematic diagram of an optional electronic device according to an embodiment of the present application. As shown in Figure 4 the figure, it includes a processor 402, a communication interface 404, a memory 406, and a communication bus 408. Among them, the processor 402, the communication interface 404, and the memory 406 complete mutual communication through the communication bus 408. Among them,
[0082] The memory 406 is configured to store a computer program;
[0083] The processor 402, when executing the computer program stored on the memory 406, implements the following steps:
[0084] Perform frame splitting on the original video according to a preset frequency to obtain an image sequence;
[0085] Perform text region detection on each frame of the image sequence to obtain the corresponding text regions and their trajectory features, where the trajectory features include duration, position coordinates, and region area;
[0086] Based on the target clustering algorithm and trajectory features, a subtitle area is determined from all text areas, where the subtitle area is a text area with the longest duration and the smallest position change amplitude among all text areas;
[0087] According to the font width-to-height ratio and the standard font width-to-height ratio in the subtitle area, the correction ratio of the original video is calculated;
[0088] According to the correction ratio, geometric transformation is performed on each frame of the image sequence to obtain a processed image sequence, and video encoding is performed on the processed image sequence to obtain the target video.
[0089] Optionally, in this embodiment, the above communication bus may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 4 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used for communication between the above electronic device and other devices.
[0090] The memory may include RAM, and may also include non-volatile memory, for example, at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0091] As an example, the above memory 406 may but is not limited to include the first processing module 302, the second processing module 304, the first determination module 306, the calculation module 308, and the correction module 310 in the above subtitle-based video correction device. In addition, it may also include but is not limited to other module units in the above subtitle-based video correction device, which will not be elaborated in this example.
[0092] The above-mentioned processor may be a general-purpose processor, including but not limited to: CPU (Central Processing Unit, central processing unit), NP (Network Processor, network processor), etc.; it may also be a DSP (Digital Signal Processing, digital signal processor), ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), FPGA (Field-Programmable Gate Array, field-programmable gate array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0093] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiment, and will not be elaborated here.
[0094] Those of ordinary skill in the art can understand that Figure 4 The structure shown is only schematic. The device for implementing the above-mentioned subtitle-based video correction method may be a terminal device, which may be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, and a mobile Internet device (Mobile Internet Devices, MID), a PAD and other terminal devices. Figure 4 It does not limit the structure of the above-mentioned electronic device. For example, the electronic device may also include more or fewer components (such as a network interface, a display device, etc.) than those shown in Figure 4 or have a different configuration from that shown in Figure 4 Those shown.
[0095] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a ROM, a RAM, a magnetic disk or an optical disc, etc.
[0096] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored. When the computer program is run by a processor, it executes the steps in the above-mentioned subtitle-based video correction method.
[0097] Optionally, in this embodiment, those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program. This program can be stored in a computer-readable storage medium, and the storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, etc.
[0098] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.
[0099] If the integrated unit in the above embodiments is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in the above computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing one or more computer devices (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application.
[0100] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0101] In the several embodiments provided by the present application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0102] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0103] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0104] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A video correction method based on subtitles, characterized in that: include: Perform frame splitting operation on the original video according to a preset frequency to obtain an image sequence; Performing text region detection on each frame of the image sequence to obtain a corresponding text region and its trajectory features, wherein the trajectory features include duration, position coordinates, and region area; Based on the target clustering algorithm and the trajectory features, a subtitle area is determined from all the text areas, wherein the subtitle area is a text area with the longest duration and the smallest position change amplitude among all the text areas; Calculating a correction ratio of the original video according to the font width-to-height ratio in the subtitle area and the standard font width-to-height ratio; According to the correction ratio, each frame of the image sequence is geometrically transformed to obtain a processed image sequence, and the processed image sequence is video encoded to obtain a target video.
2. The method according to claim 1, characterized in that Before determining the subtitle area from all the text areas based on the target clustering algorithm and the trajectory features, the method further includes: Determine a duration threshold and a position change magnitude threshold; Determine each of the text regions as a current text region, and perform the following operations on the current text region: According to the trajectory characteristics of the current text area, obtaining the current duration and the current position change amplitude of the current text area; When the current duration is less than the duration threshold, or the current position change amplitude is greater than the position change amplitude threshold, the current text area is filtered out.
3. The method according to claim 1, characterized in that The step of determining the subtitle area from all the text areas based on the target clustering algorithm and the trajectory features includes: Clustering all the text regions based on the target clustering algorithm to classify the text regions with similar trajectory features into one cluster; Determine the cluster with the longest average duration and the smallest average position change amplitude among all the clusters as the target cluster; Any text area in the target cluster is determined as the subtitle area.
4. The method according to claim 1, characterized in that: The step of calculating the correction ratio of the original video according to the font width-to-height ratio and the standard font width-to-height ratio in the subtitle area includes: Obtaining the aspect ratio of the font in the subtitle area and the standard font aspect ratio; The ratio of the font width-to-height ratio in the subtitle area to the standard font width-to-height ratio is calculated to obtain a correction ratio of the original video.
5. The method according to claim 1, characterized in that The step of performing geometric transformation on each frame of the image sequence according to the correction ratio to obtain a processed image sequence comprises: According to the correction ratio, the height and width of each frame image in the image sequence are scaled in parallel to obtain a processed image sequence, and edge filling processing is performed on each frame image in the processed image sequence.
6. A video correction device based on subtitles, characterized in that: include: The first processing module is used to perform a frame splitting operation on the original video according to a preset frequency to obtain an image sequence; A second processing module is used to perform text region detection on each frame of the image sequence to obtain a corresponding text region and its trajectory features, wherein the trajectory features include duration, position coordinates and region area; A first determination module is used to determine a subtitle area from all the text areas based on a target clustering algorithm and the trajectory features, wherein the subtitle area is a text area with the longest duration and the smallest position change amplitude among all the text areas; A calculation module, used for calculating the correction ratio of the original video according to the width-to-height ratio of the font in the subtitle area and the standard width-to-height ratio of the font; The correction module is used to perform geometric transformation on each frame of the image sequence according to the correction ratio to obtain a processed image sequence, and perform video encoding on the processed image sequence to obtain a target video.
7. The device according to claim 1, characterized in that The device also includes: A second determination module is used to determine a duration threshold and a position change amplitude threshold before determining a subtitle area from all the text areas based on a target clustering algorithm and the trajectory feature; The third processing module is used to determine each of the text regions as a current text region, and perform the following operations on the current text region: According to the trajectory characteristics of the current text area, obtaining the current duration and the current position change amplitude of the current text area; When the current duration is less than the duration threshold, or the current position change amplitude is greater than the position change amplitude threshold, the current text area is filtered out.
8. The device according to claim 1, characterized in that The first determining module comprises: A clustering unit, configured to cluster all the text regions based on the target clustering algorithm, so as to classify the text regions with similar trajectory features into one cluster; A first determining unit is used to determine the cluster with the longest average duration and the smallest average position change amplitude among all the clusters as the target cluster; The second determining unit is used to determine any text area in the target cluster as the subtitle area.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is executed.
10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 5 through the computer program.