A method and system for natural scene video text detection based on contour modeling

Through the natural scene video text detection method based on contour modeling, using Fourier inter-frame fusion and GPU acceleration, the problem of difficult detection speed and accuracy in the prior art is solved, and real-time text detection in high-definition video is realized.

CN116092069BActive Publication Date: 2025-07-25SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310058072.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-17
Publication Date
2025-07-25
Estimated Expiration
2043-01-17

AI Technical Summary

Technical Problem

The existing natural scene video text detection methods are difficult to take into account the detection speed and accuracy. There is a big gap in real-time and detection effects of traditional methods, especially in high-definition videos, which are difficult to achieve a detection rate of 30fps or higher.

Method used

The contour modeling method is adopted, and the text contour is modeled using Fourier inter-frame fusion, combined with the matching algorithm to track text targets, and accelerated inference through GPU, including video frame reading, image frame information extraction, text area prediction, inter-frame information fusion and video frame tracking. Multi-scale feature extraction is used using ResNet network and feature pyramid network with a depth of 50, set thresholds for information screening and non-maximum suppression, and match tracking is used using the IOU matrix and Hungarian algorithm.

Benefits of technology

Real-time detection of text in high-definition videos is achieved, which improves detection speed and accuracy, reduces the amount of calculation, and can meet the requirements of real-time detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116092069B_ABST
    Figure CN116092069B_ABST
Patent Text Reader

Abstract

The present invention discloses a natural scene video text detection method and system based on contour modeling, including video frame reading and initialization, extracting image frame information, predicting text region information, fusing inter-frame text information, GPU-accelerated post-processing, and video frame tracking. The inter-frame text information fusion is to set two thresholds with different sizes to fuse and screen the text information predicted by two adjacent frames to obtain enhanced text information. This method uses Fourier inter-frame fusion to model the text contour, supplements with a matching algorithm to track the text target, and at the same time uses GPU-accelerated inference to achieve real-time detection of video text while ensuring a high level of detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision, and particularly to a method and system for natural scene video text detection based on contour modeling. Background Art

[0002] With the rapid development of the Internet and the wide application of digital image capturing devices such as smart phones, digital cameras, and digital TVs, content-based image processing methods have received extensive attention. One of the most demanding applications is the accurate detection of text in natural scene videos. This technology plays an indispensable role in the fields of computer vision, machine learning, autonomous driving, real-time translation, etc. However, text detection in natural scene videos often faces problems such as slow detection speed and poor detection effect.

[0003] Traditional algorithms for natural scene video text detection are all carried out in the spatial domain and mainly fall into two categories. One category is achieved through bounding box regression. In this method, the size setting of the bounding box is usually fixed, which makes it difficult for the bounding box to fit the fine text contour. The other category is achieved through pixel segmentation. This method not only makes it difficult to aggregate complete text, but also the per-pixel operation will increase the huge computational amount, resulting in an extremely slow inference speed and making it difficult to achieve the video real-time detection effect.

[0004] At the same time, previous methods usually have difficulty in achieving both detection and speed performance. In real-life video detection applications, real-time performance is a very important requirement. That is, at least a detection rate of 30fps is required, and even 60fps or 75fps is required in high-frame-rate videos. Although existing methods have achieved good performance in terms of accuracy, there is still a large gap from the real-time goal in terms of the detection inference speed. Summary of the Invention

[0005] In order to overcome the above-mentioned shortcomings and deficiencies of the prior art, the purpose of the present invention is to provide a method and system for natural scene video text detection based on contour modeling. This method uses Fourier inter-frame fusion to model the text contour, supplemented by a matching algorithm to track the text target, and at the same time uses GPU to accelerate the inference, and can achieve real-time detection of video text while ensuring a relatively high level of detection accuracy.

[0006] The purpose of the present invention is achieved through the following technical solutions:

[0007] A method for natural scene video text detection based on contour modeling, including:

[0008] Video frame reading and initialization: Specifically, perform scale transformation on the read video frame and perform normalization operation to obtain the input image frame;

[0009] Extract image frame information: Use a ResNet network with a depth of 50 to extract the image frame information of the input image frame, and use a Feature Pyramid Network to obtain the multi-scale information of the image frame;

[0010] Text region information prediction: According to the multi-scale information, predict the text contour confidence map of the corresponding scale and the Fourier series of the text corresponding to each pixel point;

[0011] Inter-frame text information fusion: Set two thresholds with different sizes to fuse and filter the text information predicted by two adjacent frames to obtain enhanced text information. Specifically:

[0012] Set two thresholds β1 and β2 with different sizes;

[0013] First, compare the text contour confidence map cls t-1 of the previous frame with the threshold β1, and filter out the part greater than β1 to obtain the useful supplementary information cls t-1 ′ of the previous frame. Subsequently, fuse the filtered cls t-1 ′ and the text contour confidence map cls t of the current frame to enhance the prediction effect of the current frame. The obtained fused text information map is then compared with the threshold β2 to obtain the effective part greater than the threshold β2 as the final result;

[0014] GPU-accelerated post-processing: Accelerate on the GPU, model the text contour through inverse Fourier transform, and use non-maximum suppression to filter out redundant text to obtain the final text detection result;

[0015] Video frame tracking: For the text detection results of adjacent frames, construct an IOU matrix through the IOU value, and perform matching tracking through the KM algorithm and the Hungarian algorithm.

[0016] Furthermore, the text region information prediction is specifically:

[0017] The image frame information uses a Feature Pyramid Network to obtain the multi-scale features of the image frame, and the multi-scale features are respectively passed through a classification prediction head and a regression prediction head to obtain the text region information. Among them, the classification prediction head predicts the text region TR and the text center region TCR at the corresponding scale, and the regression prediction head predicts the Fourier series of the text contour at the corresponding scale.

[0018] Furthermore, the video frame tracking is specifically:

[0019] For the predicted text contours in adjacent image frames, use a matching algorithm to track them. For the contours in the image frame at the previous moment t-1 and the contours in the image frame at the current moment t, calculate the IOU values pairwise to construct an IOU matrix. Through the IOU matrix, use the KM algorithm for matching. If the matching is successful, update the tracking status of the text contour; if the matching fails, check the tracking status. If the maximum tracking duration is reached, delete the text contour; if the maximum tracking duration is not reached, retain the text contour and update its tracking duration.

[0020] Furthermore, the Fourier series of the text corresponding to each pixel specifically abstracts the text contour point sequence into a Fourier series, including:

[0021] Use a complex-valued function f: R→C of a real variable t∈[0,1] to represent any text closed contour as follows:

[0022] f(t) = x(t) + iy(t)

[0023] i represents the imaginary unit, and (x(t), (t)) are the spatial coordinates at a specific time t. Since f is a closed contour, f(t) = f(t + 1), and f(t) is re-expressed by the inverse Fourier transform (IFT) as:

[0024]

[0025] k∈Z represents the frequency, and c k is the complex-valued Fourier coefficient used to characterize the initial state of the frequency k.

[0026] Furthermore, the text contour confidence map is obtained by multiplying the text region confidence and the text center region confidence, where the center region of the text is obtained by indenting the text region inward by a distance of 0.3 times the average height of the text.

[0027] Furthermore, before network training, divide the text instances into three categories: small, medium, and large according to the size ratio of the text instance samples, where the size ratio r is determined by the ratio of the larger value of the maximum horizontal difference dx and the maximum vertical difference dy of the text instance to the image height h:

[0028] r = max(dx, dy) / h

[0029] The small, medium, and large three types of targets respectively correspond to the multi-scale feature outputs in the feature pyramid.

[0030] A system for a natural scene video text detection method, including

[0031] Video reading and initialization module: used to perform scale transformation on the read video frames and perform normalization operations to obtain input image frames;

[0032] Image frame extraction module: It is used to extract the image frame information of the input image frame using a ResNet network with a depth of 50, and obtain the multi-scale features of the image frame using a feature pyramid network;

[0033] Text area information prediction module: It is used to predict the binary map of the text contour confidence of the corresponding scale and the Fourier series of the text corresponding to each pixel according to the multi-scale information;

[0034] Inter-frame text information fusion module: It is used to set two thresholds with different sizes to fuse and screen the text information predicted by two adjacent frames to obtain enhanced text information, and select the pixel points where the binary map of the confidence is greater than the threshold for operation with the predicted regression Fourier series;

[0035] GPU-accelerated post-processing module: It is used to perform acceleration on the GPU, model the text contour through inverse Fourier transform, use non-maximum suppression to filter out redundant text, and obtain the final text detection result;

[0036] Video frame tracking module: It is used to construct an IOU matrix for the text detection results of adjacent frames, and perform matching and tracking through the KM algorithm and the Hungarian algorithm.

[0037] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0038] This method can predict the Fourier coefficients corresponding to the text contour at multiple scales, then reconstruct the point sequence of the text contour through the inverse Fourier transform method, and perform tracking through text matching of adjacent frames. It is more accurate than the traditional methods of per-pixel segmentation prediction and bounding box regression prediction, and at the same time greatly reduces the calculation amount, improves the inference speed, and realizes real-time detection on high-definition videos. Description of the Drawings

[0039] Figure 1 is the workflow diagram of the present invention. Detailed Embodiments

[0040] The following combines embodiments to further elaborate on the present invention in detail, but the implementation manners of the present invention are not limited thereto.

[0041] As Figure 1 shown, this embodiment provides a method for detecting natural scene video text based on contour modeling. This method can predict the Fourier coefficients corresponding to the text contour at multiple scales, then reconstruct the point sequence of the text contour through the inverse Fourier transform method, and perform tracking through text matching of adjacent frames. It is more accurate than the traditional methods of per-pixel segmentation prediction and bounding box regression prediction, and at the same time greatly reduces the calculation amount, improves the inference speed, and realizes real-time detection on high-definition videos. The method specifically includes the following steps:

[0042] S1. Read the video frames, perform scale transformation on the read video frames, and perform normalization operations to obtain the final input image frames.

[0043] Specifically:

[0044] S11. Before network training, it is first necessary to process the natural scene video text dataset. The videos in the dataset usually have a large number of low-quality image frames and blurred text instances. It is necessary to mark the low-quality image frames and text instances in advance and exclude these samples during training to prevent interference with the training of normal samples. At the same time, the video dataset is too large, and it is necessary to select images by frame extraction from the video instead of directly inputting all frame images into the training. This operation can not only greatly reduce the training time consumption but also improve the robustness of the network.

[0045] S12. Before the video image frames are input into network training, they need to undergo certain data augmentation and normalization. Each image frame is enhanced through random size adjustment, random cropping and flipping, rotation and padding, and random brightness, contrast, and mirror adjustments, and finally, the final input image is obtained through normalization.

[0046] S2. Use a ResNet network with a depth of 50 to extract the image frame information, and use a Feature Pyramid Network to obtain the multi-scale information of the image frames.

[0047] S3. According to the multi-scale information extracted by the Feature Pyramid Network, predict the text confidence map of the corresponding scale and the Fourier series of the text corresponding to each pixel point.

[0048] Prediction of text region information: After the image frame F passes through the backbone network, multi-scale features C3, C4, and C5 are obtained. The multi-scale features are respectively passed through a classification prediction head and a regression prediction head to obtain the predicted text region information. Among them, the classification prediction head predicts the text region TR and the text center region TCR at the corresponding scale, and the regression prediction head predicts the Fourier series of the text contour at the corresponding scale.

[0049] Specifically:

[0050] S31. During network training, each text contour is represented by a complex-valued function f: R → C of a real variable t ∈ [0,1]:

[0051] f(t) = x(t) + iy(t)

[0052] i represents the imaginary unit, (x(t), (t)) are the spatial coordinates at a specific time t, and it is these spatial coordinate sequences that are fed into the network. Since f is a closed contour, f(t) = f(t + 1). f(t) can be reformulated through the Inverse Fourier Transform (IFT) as:

[0053]

[0054] k ∈ Z represents the frequency, and c k is a complex-valued Fourier coefficient used to characterize the initial state of the frequency k. In this method, k only takes five frequency values of -2, -1, 0, 1, and 2, and the Fourier series predicted by the network corresponds to c in the corresponding formula k .

[0055] S32. During network training, the text instances are divided into three categories: small, medium, and large according to the size ratio of the text instance samples, and are assigned to different output levels of the feature pyramid for supervised learning of the network according to the classification results. Among them, the size ratio r is determined by the ratio of the larger value between the maximum horizontal difference dx and the maximum vertical difference dy of the text instance to the image height h

[0056] r = max(dx, dy) / h

[0057] S33. During network training, the final confidence C of the predicted text contour is obtained by multiplying the text region confidence C tr and the text center region confidence C tcr where the center region of the text is obtained by indenting the text region inward by a distance of 0.3 times the average height of the text

[0058]

[0059] S4. Use a video frame-interval enhanced fusion module proposed by us to fuse and screen the predicted text information of two adjacent frames through two different thresholds to obtain enhanced text information. Specifically, for the text contour confidence maps cls t-1 and cls t of the predicted outputs of two adjacent frames, we set two thresholds β1 and β2. First, the predicted value cls t-1 of the previous frame is screened through the threshold β1 to obtain the useful supplementary information cls t-1 ′ of the previous frame. Subsequently, the screened cls t-1 ′ and the predicted value cls t of the current frame are fused to enhance the prediction effect of the current frame. The obtained fused text information map is then screened again using the threshold β2 to obtain the effective part as the final result. For the results of each scale, the confidence of the text contour is calculated by weighted summation of the text region TR and the text center region TCR. Finally, the pixel points with confidence greater than the threshold are selected, and the result obtained by operating on the predicted regression Fourier series is sent to the subsequent post-processing

[0060] Furthermore, the β1 and β2 thresholds need to be set according to the moving speed and complexity in the video. For example​

[0061] When the train passes at high speed, the two threshold values are as follows: 0.95, 1.7.

[0062] When a person holds a camera and moves forward, the two threshold values are as follows: 0.8, 1.2.

[0063] S5. Model the text contour through Fourier transform, integrate all predicted text contours, use non-maximum suppression to filter out redundant text, and obtain the final text detection result. The entire process is accelerated using a GPU.

[0064] Post-processing of text information: For all obtained text contours, filter out overlapping text contours through non-maximum suppression.

[0065] Text contour reconstruction model: During inference, use the inverse Fourier transform (IFT) to reconstruct the predicted Fourier series model to obtain the point sequence of the text contour:

[0066]

[0067] S6. Perform video frame tracking on the text detection results of adjacent frames, construct an IOU matrix through their IOU values, and then perform matching tracking through the KM algorithm and the Hungarian algorithm.

[0068] Video frame tracking: For the text contours predicted in adjacent image frames, use a matching algorithm to track them. Calculate the IOU values pairwise for the contours in the previous frame at time t - 1 and the contours in the current frame at time t to construct an IOU matrix. Through the IOU matrix, use the KM algorithm for matching. If the matching is successful, update the tracking status of the text contour; if the matching fails, check the tracking status. If the maximum tracking duration is reached, delete the text contour; if the maximum tracking duration is not reached, retain the text contour and update the tracking duration of the text.

[0069] The benefits and advantages of the present invention include the following points:

[0070] The present invention can use Fourier inter-frame fusion modeling to model text instances in videos and detect complete and accurate text instances.

[0071] According to the designed GPU acceleration method, the present invention has an extremely fast inference speed and can achieve real-time detection in high-definition natural scene videos.

[0072] This embodiment also provides a natural scene video text detection system based on contour modeling, including:

[0073] Video reading and initialization module: used to perform scale transformation on the read video frames and perform normalization operations to obtain input image frames;

[0074] Image frame extraction module: It is used to extract the image frame information of the input image frame using a ResNet network with a depth of 50, and obtain the multi-scale features of the image frame using a feature pyramid network;

[0075] Text region information prediction module: It is used to predict the binary map of the text contour confidence at the corresponding scale and the Fourier series of the text corresponding to each pixel according to the multi-scale information;

[0076] Inter-frame text information fusion module: It is used to set two thresholds with different sizes to fuse and screen the text information predicted by two adjacent frames to obtain enhanced text information, and select the pixel points with the binary map of the confidence greater than the threshold for operation with the predicted regression Fourier series;

[0077] GPU-accelerated post-processing module: It is used to perform acceleration on the GPU, model the text contour through inverse Fourier transform, and use non-maximum suppression to filter out redundant text to obtain the final text detection result;

[0078] Video frame tracking module: It is used to construct an IOU matrix for the text detection results of adjacent frames, and perform matching and tracking through the KM algorithm and the Hungarian algorithm.

[0079] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited by the described embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A method for natural scene video text detection based on contour modeling, characterized in that, Including: Video frame reading and initialization: Specifically, scale transformation is performed on the read video frame, and normalization operation is carried out to obtain the input image frame; Extracting image frame information: Using a ResNet network with a depth of 50 to extract the image frame information of the input image frame, and using a feature pyramid network to obtain the multi-scale information of the image frame; Predicting text region information: According to the multi-scale information, predict the text contour confidence map of the corresponding scale and the Fourier series of the text corresponding to each pixel point; Inter-frame text information fusion: Set two thresholds with different sizes to fuse and screen the text information predicted by two adjacent frames to obtain enhanced text information. Specifically: Set two thresholds β1 and β2 with different sizes; First, compare the text contour confidence map cls of the previous frame t-1 with the threshold β1, and select the part greater than β1 to obtain the useful supplementary information cls of the previous frame t-1 ′. Subsequently, fuse the filtered cls t-1 ′ and the text contour confidence map cls of the current frame t to enhance the prediction effect of the current frame. Then, compare the obtained fused text information map with the threshold β2, and take the valid part greater than the threshold β2 as the final result; GPU-accelerated post-processing: Accelerate on the GPU, model the text contour through inverse Fourier transform, and use non-maximum suppression to screen out redundant text to obtain the final text detection result; Video frame tracking: For the text detection results of adjacent frames, construct an IOU matrix through the IOU value, and perform matching tracking through the KM algorithm and the Hungarian algorithm; The video frame tracking is specifically: For the text contours predicted in adjacent image frames, use a matching algorithm to track them. For the contours in the image frame at the previous moment t-1 and the contours in the image frame at the current moment t, calculate the IOU value pairwise to construct an IOU matrix. Through the IOU matrix, use the KM algorithm for matching. If the matching is successful, the tracking status of the text contour is updated; If the matching fails, check the tracking status. If the maximum tracking duration is reached, delete the text contour. If the maximum tracking duration is not reached, retain the text contour and update the tracking duration of the text.

2. The natural scene video text detection method according to claim 1, characterized in that The prediction of the text region information is specifically: The image frame information uses a feature pyramid network to obtain the multi-scale features of the image frame. The multi-scale features are respectively passed through a classification prediction head and a regression prediction head to obtain the text region information. Among them, the classification prediction head predicts the text region TR and the text center region TCR at the corresponding scale, and the regression prediction head predicts the Fourier series of the text contour at the corresponding scale.

3. The natural scene video text detection method according to claim 1, characterized in that, The Fourier series corresponding to each pixel of the text is specifically to abstract the text contour point sequence into a Fourier series, including: Use a complex-valued function f: R→C of a real variable t∈[0,1] to represent any text closed contour as follows: f(t) = x(t) + iy(t) i represents the imaginary unit, and (x(t), y(t)) is the spatial coordinate at a specific time t. Since f is a closed contour, f(t) = f(t + 1), and f(t) is re-expressed through inverse Fourier transform as: $k\in\mathbb{Z}$ represents the frequency, and $c$ k is a complex Fourier coefficient used to characterize the initial state of the frequency $k$.

4. The natural scene video text detection method according to claim 1, wherein The text contour confidence map is obtained by multiplying the text region confidence and the text center region confidence, where the center region of the text is obtained by indenting the text region inward by a distance of 0.3 times the average character height of the text.

5. The natural scene video text detection method according to claim 1, wherein Before network training, text instances are divided into three categories: small, medium, and large according to the size ratio of text instance samples, where the size ratio r is determined by the ratio of the larger value between the maximum horizontal difference dx and the maximum vertical difference dy of the text instance to the image height h: r = max(dx, dy) / h Small, medium, and large targets respectively correspond to the multi-scale feature outputs in the feature pyramid.

6. A system for implementing the natural scene video text detection method according to any one of claims 1-5, characterized in that, Including Video reading and initialization module: used to perform scale transformation on the read video frames and perform normalization operations to obtain the input image frames; Image frame extraction module: used to extract the image frame information of the input image frames using a ResNet network with a depth of 50 and obtain the multi-scale features of the image frames using the feature pyramid network; Text area information prediction module: used to predict the binary map of the text contour confidence at the corresponding scale and the Fourier series of each pixel corresponding to the text according to the multi-scale information; Inter-frame text information fusion module: used to set two thresholds of different sizes to fuse and screen the text information predicted by adjacent two frames to obtain enhanced text information, and select the pixel points with the binary map of confidence greater than the threshold to perform operations with the predicted regression Fourier series; GPU-accelerated post-processing module: used to perform acceleration on the GPU, model the text contour through inverse Fourier transform, and use non-maximum suppression to screen out redundant text to obtain the final text detection result; Video frame tracking module: used to construct an IOU matrix for the text detection results of adjacent frames and perform matching tracking through the KM algorithm and the Hungarian algorithm.

Citation Information

Patent Citations

  • Text detection method and device, electronic equipment and storage medium

    CN112800954A

  • Natural scene character detection method and system for slender text and medium

    CN114882485A