A method, apparatus, device, and storage medium for video text detection.
By acquiring three single-frame images and aligning feature maps using optical flow information, and combining text probability maps and weights to determine text regions, the problem of low accuracy caused by motion in video text detection is solved, achieving higher detection accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-17
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies for text detection in videos suffer from low accuracy because they ignore the temporal changes in the video, especially when the text is moved, the lighting changes, or blurred.
By acquiring three single-frame images, an initial feature map is obtained using a feature extraction network. The feature map is then aligned using optical flow information. The text region is determined by combining the text probability map and weights, compensating for feature shifts caused by motion between frames and improving detection accuracy.
It effectively improves the accuracy of video text detection. By pre-calculating the optical flow values of the previous and next frames and the current frame, the extracted features from the previous and next frames are mapped to the current frame, compensating for feature offset caused by motion and improving the accuracy of text detection.
Smart Images

Figure CN116824115B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information processing, and includes, but is not limited to, a video text detection method, apparatus, device, and storage medium. Background Technology
[0002] Existing technical solutions for video text detection rely on detecting text in a single frame. This involves extracting video frames at intervals and then performing text detection on the extracted images; alternatively, it involves extracting video images at time intervals and then performing text detection. Both methods detect text within a single frame of the video.
[0003] The above methods have limitations for text detection in videos. Text in videos undergoes changes in shape, lighting, speed, and blurring due to video movement. Not all text information in a single frame is suitable for static image detection, and single-frame image detection ignores temporal changes in video features. Furthermore, they fail to distinguish between video text and image text; static image detection methods often deteriorate significantly when applied to video, resulting in lower text detection accuracy. Summary of the Invention
[0004] In view of this, embodiments of this application provide a video text detection method, apparatus, device, and storage medium.
[0005] The technical solution of this application embodiment is implemented as follows:
[0006] In a first aspect, embodiments of this application provide a video text detection method, the method comprising: acquiring three single-frame images based on a preset sampling interval, wherein the three single-frame images include image one, image two, and image three, and image two is located between image one and image three; and acquiring initial features of image one using a feature extraction network. Figure 1 Features of Image Two Figure 2 and the initial features of image three Figure 3 ; Determine the first optical flow information of image one and image two, and the second optical flow information of image three and image two; Set the initial features Figure 1 Alignment to the feature based on the first optical flow information Figure 2 , to obtain features Figure 1 ; the initial features Figure 3 Alignment to the feature based on the second optical flow information Figure 2 , to obtain features Figure 3 Based on the aforementioned features Figure 1 Text probability Figure 1 And weight 1, the features mentioned Figure 2 Text probability Figure 2And weight 2, the features mentioned Figure 3 Text probability Figure 3 The text region of the second image is determined by weight three.
[0007] Secondly, this application provides a video text detection device, the device comprising: an acquisition module, configured to acquire three single-frame images based on a preset sampling interval, wherein the three single-frame images include image one, image two and image three, and image two is located between image one and image three;
[0008] The extraction module is used to obtain the initial features of image one using a feature extraction network. Figure 1 Features of Image Two Figure 2 and the initial features of image three Figure 3 The first determining module is used to determine the first optical flow information of image one and image two, and the second optical flow information of image three and image two; the first alignment module is used to align the initial features Figure 1 Alignment to the feature based on the first optical flow information Figure 2 , to obtain features Figure 1 The second alignment module is used to align the initial features. Figure 3 Alignment to the feature based on the second optical flow information Figure 2 , to obtain features Figure 3 The second determining module is used to determine based on the features. Figure 1 Text probability Figure 1 And weight 1, the features mentioned Figure 2 Text probability Figure 2 And weight 2, the features mentioned Figure 3 Text probability Figure 3 The text region of the second image is determined by weight three.
[0009] Thirdly, embodiments of this application provide an electronic device, including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the program to implement the above-described method.
[0010] Fourthly, embodiments of this application provide a storage medium storing executable instructions for inducing a processor to execute the above-described method.
[0011] In this embodiment, three single-frame images are first acquired based on a preset sampling interval. These three single-frame images include image one, image two, and image three, with image two located between image one and image three. Then, a feature extraction network is used to obtain the initial features of image one. Figure 1 Features of Image Two Figure 2 and the initial features of image three Figure 3; Determine the first optical flow information of image one and image two, and the second optical flow information of image three and image two; Set the initial features Figure 1 Alignment to the feature based on the first optical flow information Figure 2 , to obtain features Figure 1 ; the initial features Figure 3 Alignment to the feature based on the second optical flow information Figure 2 , to obtain features Figure 3 Finally, based on the aforementioned features Figure 1 Text probability Figure 1 And weight 1, the features mentioned Figure 2 Text probability Figure 2 And weight 2, the features mentioned Figure 3 Text probability Figure 3 The text region in image two is determined by weight three. In this way, the optical flow values of the preceding and following frames and the current frame are pre-calculated, and the extracted features from the preceding and following frames are mapped to the features of the current frame; this compensates for the image feature shift caused by motion in the preceding and following frames, supplements the feature representation of the current frame, and aggregates the features using weights, effectively improving the text detection accuracy of the current frame. Attached Figure Description
[0012] Figure 1 A flowchart illustrating a video text detection method provided in an embodiment of this application;
[0013] Figure 2 A schematic diagram of a model architecture for detecting text in video provided in an embodiment of this application;
[0014] Figure 3 A flowchart illustrating a model for training and detecting video text, provided in an embodiment of this application;
[0015] Figure 4 This application provides a schematic diagram of a specific model architecture for detecting text in video.
[0016] Figure 5 This is a schematic diagram of the composition structure of a video text detection device provided in an embodiment of this application;
[0017] Figure 6 This is a schematic diagram of a hardware entity of an electronic device provided in an embodiment of this application. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the specific technical solutions of the embodiments will be further described in detail below with reference to the accompanying drawings. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application.
[0019] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0020] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0022] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0023] Optical flow (or optical flow) is the instantaneous velocity of pixels moving on the imaging plane of a spatially moving object. It's a method that uses the temporal changes of pixels in an image sequence and the correlation between adjacent frames to find the correspondence between the previous and current frames, thereby calculating the motion information of objects between adjacent frames. It describes the motion of the observed target, surface, or edge caused by the motion relative to the observer.
[0024] The VGG network is an improvement on the classic AlexNet Convolutional Neural Network (CNN) model. It is a classic CNN model. The VGG network uses very small 3x3 convolutional templates and five 2x2 pooling layers, and increases the depth of the convolutional layers to 16 to 19 layers.
[0025] The goal of Feature Pyramid Network (FPN) is to construct feature pyramids using the hierarchical semantic features inherent in convolutional networks. It can serve as a general feature extractor, delivering significant performance improvements across multiple tasks.
[0026] This application provides a video text detection method, such as... Figure 1 As shown, the method includes:
[0027] Step S110: Acquire three single-frame images based on a preset sampling interval, wherein the three single-frame images include image one, image two and image three, and image two is located between image one and image three;
[0028] Here, the preset sampling interval can be determined based on the video sampling rate. For example, if the sampling rate is determined to be 60 frames per second, three single-frame images can be obtained at 3-frame intervals; if the sampling rate is determined to be 20 frames per second, three single-frame images can be obtained at 1-frame intervals.
[0029] In some embodiments, images 1, 2, and 3 can be acquired sequentially, then image 2 can be determined as the current frame image of the text to be detected, and images 1 and 3 can be determined as adjacent frame images of the current frame image.
[0030] In some embodiments, image 2 can be acquired first and determined as the current frame image of the text to be detected. Then, images 1 and 3 are acquired before and after image 2, respectively, with a preset sampling interval. That is, images 1 and 3 are adjacent frame images of the current frame image.
[0031] Step S120: Obtain the initial features of image one using a feature extraction network. Figure 1 Features of Image Two Figure 2 and the initial features of image three Figure 3 ;
[0032] Figure 2 This application provides a schematic diagram of a model architecture for detecting text in video, as shown in the embodiments of this application. Figure 2 As shown, the model architecture includes: a feature extraction network 21, a warp component 22, a classification sub-network 23, a weight sub-network 24, a dot product component 25, and an aggregation weight component 26.
[0033] In some embodiments, the feature extraction network 21 can be used to obtain the initial features of image one. Figure 1 Features of Image Two Figure 2 and the initial features of image three Figure 3 The feature extraction network 21 can be composed of a VGG subnetwork and an FPN subnetwork.
[0034] Step S130: Determine the first optical flow information of the first image and the second image, and the second optical flow information of the third image and the second image;
[0035] Here, optical flow information can be used to represent the motion speed and direction of each pixel in two adjacent frames. The first optical flow information between image one and image two, and the second optical flow information between image three and image two, can be calculated using an optical flow information calculation method.
[0036] Step S140: The initial features Figure 1 Alignment to the feature based on the first optical flow information Figure 2 , to obtain features Figure 1 ;
[0037] In some embodiments, it is possible to utilize Figure 2 The Warp component 22 shown implements the initial features Figure 1 Alignment to features based on the first optical flow information Figure 2 , to obtain features Figure 1 Here, Warp can be implemented using a bilinear interpolation algorithm, which means finding the corresponding feature points of the previous frame based on the feature points of the current frame using optical flow information, obtaining those feature points, and combining the information from the current frame and the previous frame to obtain the final feature representation.
[0038] Step S150: The initial features Figure 3 Alignment to the feature based on the second optical flow information Figure 2 , to obtain features Figure 3 ;
[0039] In some embodiments, it is possible to utilize Figure 2 The Warp component 22 shown implements the initial features Figure 3 Alignment to features based on second optical flow information Figure 2 , to obtain features Figure 3 .
[0040] Step S160: Based on the features Figure 1 Text probability Figure 1 And weight 1, the features mentioned Figure 2 Text probability Figure 2 And weight 2, the features mentioned Figure 3 Text probability Figure 3 The text region of the second image is determined by weight three.
[0041] During implementation, it is possible to utilize Figure 2 The classification subnetwork 23 shown obtains features Figure 1 Text probability Figure 1 ,feature Figure 2 Text probability Figure 2 and characteristics Figure 3 Text probability Figure 3 Features are obtained using weighted subnetwork 24. Figure 1 Weights and features Figure 2 Weighted binary features Figure 3 Weight 3.
[0042] Based on weight one, weight two, and weight three, the text probability is fused. Figure 1 Probability of words Figure 2 and the probability of words Figure 3 The text region in image 2 is determined, which means the text region in the current frame is determined, thus realizing the detection of text in the video.
[0043] In this embodiment, three single-frame images are first acquired based on a preset sampling interval. These three single-frame images include image one, image two, and image three, with image two located between image one and image three. Then, a feature extraction network is used to obtain the initial features of image one. Figure 1 Features of Image Two Figure 2 and the initial features of image three Figure 3 ; Determine the first optical flow information of image one and image two, and the second optical flow information of image three and image two; Set the initial features Figure 1 Alignment to the feature based on the first optical flow information Figure 2 , to obtain features Figure 1 ; the initial features Figure 3 Alignment to the feature based on the second optical flow information Figure 2 , to obtain features Figure 3 Finally, based on the aforementioned features Figure 1 Text probability Figure 1 And weight 1, the features mentioned Figure 2 Text probability Figure 2 And weight 2, the features mentioned Figure 3 Text probability Figure 3 The text region in image two is determined by weight three. In this way, the optical flow values of the preceding and following frames and the current frame are pre-calculated, and the extracted features from the preceding and following frames are mapped to the features of the current frame; this compensates for the image feature shift caused by motion in the preceding and following frames, supplements the feature representation of the current frame, and aggregates the features using weights, effectively improving the text detection accuracy of the current frame.
[0044] In some embodiments, the above step S130, "determining the first optical flow information of the first image and the second image, and the second optical flow information of the third image and the second image," can be achieved through the following steps:
[0045] Step 131: Determine the first optical flow information of Image 1 and Image 2 using the TV-L1 algorithm;
[0046] The calculation of TVL1 optical flow information employs a numerical analysis mechanism based on bidirectional solution using image denoising to minimize the total variational optical flow energy function. The optical flow field is used as input to capture motion information. The optical flow information can be calculated by the rate of change of pixel grayscale values between two adjacent frames.
[0047] Step 132: Use the TV-L1 algorithm to determine the second optical flow information of image 3 and image 2.
[0048] During implementation, the TV-L1 algorithm can be used to predetermine the first optical flow information of image 1 and image 2, and the second optical flow information of image 3 and image 2.
[0049] Here, there is no restriction on the execution order of steps 131 and 132; they can be executed sequentially or simultaneously.
[0050] In this embodiment of the application, the TV-L1 algorithm can be used to effectively obtain the first optical flow information of image one and image two, and the second optical flow information of image three and image two in advance.
[0051] In some embodiments, step S140 above "the initial feature" Figure 1 Alignment to the feature based on the first optical flow information Figure 2 , to obtain features Figure 1 This can be achieved through the following steps:
[0052] Based on the first optical flow information, the bilinear interpolation algorithm is used in conjunction with the initial features. Figure 1 and the features Figure 2 , to obtain features Figure 1 ;
[0053] During implementation, the initial features can be... Figure 1 Based on the first optical flow information, warp to the feature Figure 2 , to obtain features Figure 1 .
[0054] Here, warp can infer information about the image at the next moment using optical flow information and current image information. It can be based on initial features. Figure 1 Linear sampling is performed with the first optical flow information, and then combined with the features Figure 2 Initial features of adjacent frames Figure 1 Based on the first optical flow information, warp to the feature Figure 2 , to obtain features Figure 1 .
[0055] Correspondingly, in some embodiments, step S150 above "transfers the initial features" Figure 3 Alignment to the feature based on the second optical flow information Figure 2, to obtain features Figure 3 This can be achieved through the following steps:
[0056] Based on the second optical flow information, the bilinear interpolation algorithm is combined with the initial features. Figure 3 and the features Figure 2 , to obtain features Figure 3 .
[0057] During implementation, the initial features can be... Figure 3 Based on the second optical flow information, warp to the feature Figure 2 , to obtain features Figure 3 .
[0058] In this embodiment, the optical flow values of the preceding and following frames and the current frame are pre-calculated, and the extracted features of the preceding and following frames are mapped to the features of the current frame. This compensates for the image feature shift caused by motion in the preceding and following frames and supplements the feature representation of the current frame.
[0059] In some embodiments, step S160 above "based on the features" Figure 1 Text probability Figure 1 And weight 1, the features mentioned Figure 2 Text probability Figure 2 And weight 2, the features mentioned Figure 3 Text probability Figure 3 "Determining the text region of image two using weight three" can be achieved through the following steps:
[0060] Step 161: Detect the features using a classification sub-network. Figure 1 The features Figure 2 and the features Figure 3 The features are obtained. Figure 1 Text probability Figure 1 The features Figure 2 Text probability Figure 2 and the features Figure 3 Text probability Figure 3 ;
[0061] Here, C can be used. t-1 Representing the probability of a character Figure 1 C t Representing the probability of a character Figure 2 C t+1 Representing the probability of a character Figure 3 .
[0062] Step 162: Obtain the features using weighted subnetworks. Figure 1 Weight 1, the aforementioned features Figure 2 The weights of the two features Figure 3 Weight three;
[0063] Here, an adaptive weighted subnetwork can be built to integrate the features. Figure 1 ,feature Figure 2 and characteristics Figure 3 The features are obtained by inputting the weighted subnetwork into the components. Figure 1 weight one W t-1 ,feature Figure 2 Weight 2W t and characteristics Figure 3 Weights of three Ws t+1 , where t represents the current frame, and t-1 and t+1 represent the frames before and after the current frame.
[0064] Step 163: Multiply the weight 1 by the weight 2, multiply the weight 2 by the weight 2, and multiply the weight 3 by the weight 2 to obtain the feature. Figure 1 First dot product similarity, the features Figure 2 The second dot product similarity and the features Figure 3 The third dot product similarity;
[0065] During implementation, features Figure 1 The first dot product similarity can be obtained using the following formula (1):
[0066] S t-1 =W t-1 ·W t (1);
[0067] feature Figure 2 The second dot product similarity can be obtained using the following formula (2):
[0068] S t =W t ·W t (2);
[0069] feature Figure 3 The third dot product similarity can be obtained using the following formula (3):
[0070] S t+1 =W t+1 ·W t (3);
[0071] In this way, the dot product similarity of the three images can be calculated.
[0072] Step 163, based on the aforementioned features Figure 1 Text probability Figure 1 Similarity to the first dot product, the features Figure 2 Text probability Figure 2 Second dot product similarity, the features Figure 3 Text probability Figure 3The target probability map of image two is determined by the third dot product similarity.
[0073] Step 164: Determine the text region based on the target probability map.
[0074] In this embodiment, a classification subnetwork is used to calculate the probability maps of text in the feature maps of different frames after warping; a weight subnetwork is used to calculate the weights of the current frame and the frames before and after warping; finally, the three feature probability maps are aggregated according to the weights and probability maps, and different aggregation weights are assigned to the three frame feature probability maps, which effectively enhances the formation of text detection features.
[0075] In some embodiments, step 163 above, "based on the features" Figure 1 Text probability Figure 1 Similarity to the first dot product, the features Figure 2 Text probability Figure 2 Second dot product similarity, the features Figure 3 Text probability Figure 3 "Determining the target probability map of image two by combining the third dot product similarity" can be achieved through the following steps:
[0076] Step 1631, based on the aforementioned features Figure 1 Text probability Figure 1 Similarity to the first dot product, the features Figure 2 Text probability Figure 2 Second dot product similarity, the features Figure 3 Text probability Figure 3 The features are determined accordingly by the third dot product similarity. Figure 1 The first aggregation weight, the feature Figure 2 The second aggregation weight and the feature Figure 3 The third aggregation weight;
[0077] During implementation, features Figure 1 The first aggregation weight can be obtained using the following formula (4):
[0078]
[0079] feature Figure 2 The second aggregation weight can be obtained using the following formula (5):
[0080]
[0081] feature Figure 3 The third aggregation weight can be obtained using the following formula (6):
[0082]
[0083] Among them, Agt-1 Indicates the first aggregation weight, Ag t Indicates the second aggregation weight, Ag t+1 C represents the third aggregation weight. t-1 Indicates the probability of the text Figure 1 C t Indicates the probability of the text Figure 2 C t+1 Indicates the probability of the text Figure 3 S t-1 S represents the first dot product similarity. t S represents the second dot product similarity. t+1 The third dot product is similar; exp is an exponential function with the natural constant e as its base. exp(x) represents e raised to the power of x, where x can be a function.
[0084] Step 1632: Based on the first aggregation weight, the second aggregation weight, and the third aggregation weight, calculate the text probability. Figure 1 The probability of the text Figure 2 and the probability of the text Figure 3 The target probability map is obtained by weighted summation.
[0085] During implementation, the target probability map can be obtained using the following formula (7):
[0086] ALL = Ag t-1 *C t-1 +Ag t *C t +Ag t+1 *C t+1 (7).
[0087] In this embodiment, based on the first aggregated weight, the second aggregated weight, and the third aggregated weight obtained from the text probability graph and weights, it is possible to achieve text probability... Figure 1 Probability of words Figure 2 and the probability of words Figure 3 Weighted summation yields the target probability map.
[0088] In some embodiments, step 164 above, "determining the text region based on the target probability map," can be achieved through the following steps:
[0089] Step 1641: Binarize the target probability map;
[0090] Here, the binarization of the target probability map can be achieved by first setting a probability threshold, then obtaining the probability of each pixel in the target probability map being text, and finally binarizing the target probability map based on the probability threshold and the probability of each pixel being text, so as to mark pixels with a text probability greater than or equal to the probability threshold as 1, and pixels with a text probability less than the probability threshold as 0.
[0091] Step 1642: Based on the binarized target probability map, determine the text pixels in the target probability map;
[0092] Step 1643: Determine the text region based on the text pixels.
[0093] During implementation, the smallest rectangle covering the text area can be generated based on the text pixel marked as 1.
[0094] In this embodiment of the application, the probability map is binarized. When a pixel is determined to be a text region, the smallest rectangle covering the text region can be generated, thus determining the text region.
[0095] In some embodiments, the video text detection model includes at least the feature extraction network, the classification subnetwork, and the weight subnetwork. The feature extraction network includes a VGG subnetwork and an FPN subnetwork, such as... Figure 3 As shown, training the model includes the following steps:
[0096] Step S310: Train the VGG subnetwork, the FPN subnetwork, the classification subnetwork, and the weight subnetwork using the training dataset to update the parameters of the VGG subnetwork, the FPN subnetwork, the classification subnetwork, and the weight subnetwork.
[0097] Step S320: During the training process, the DiceCoefficient loss function is used to measure the error between the prediction result and the target result until the error converges.
[0098] Here, since the non-text region in a normal video frame is much larger than the text region, using binary classification cross-entropy loss will bias the results towards the non-text region. Therefore, DiceCoefficient can be used as the loss function, defined as follows:
[0099]
[0100] Among them, S x,y To predict the value of pixel (x, y) in an instance, G x,y This represents the value of the pixel (x, y) in the label.
[0101] In this embodiment, the VGG sub-network training parameters are pre-loaded, the training dataset is input into the network, and the network parameters are updated. During training, DiceCoefficient is used as the loss function, and the parameters are optimized using the DiceCoefficient loss function, resulting in parameters that better meet actual needs.
[0102] Figure 4 This application provides a schematic diagram of a specific model architecture for detecting text in video, as shown in the embodiments of this application. Figure 4 As shown, the model includes: a feature extraction network 41, a warp component 42, a classification sub-network 43, a weight sub-network 44, a dot product component 45, and an aggregation weight component 46, wherein,
[0103] The feature extraction network 41 can be composed of a VGG subnetwork and an FPN subnetwork, used to extract feature maps from the input image. For example, it can extract features from three input images (image 1, image 2, and image 3) to obtain initial features. Figure 1 ,feature Figure 2 and initial features Figure 3 ;
[0104] Warp component 42 can acquire optical flow information calculated using an optical flow information calculation method, separately calculating the optical flow information of image 2 and the other two frames (image 1 and image 3). Then, it performs linear sampling based on the feature map and optical flow information, and combines the feature maps of adjacent frames (initial feature maps) with the optical flow information. Figure 1 and initial features Figure 3 Based on optical flow information, warp to the feature map of the current frame (feature map). Figure 2 This yields a new feature map (feature map). Figure 1 and characteristics Figure 3 Among them, optical flow information calculation methods include the Lucas-Kanade algorithm and the TV-L1 algorithm; the implementation of Warp can be carried out by using bilinear interpolation, that is, based on the feature points of the current frame, the corresponding feature points of the previous frame are found from the optical flow information, and then the feature points are taken and combined with the information of the current frame and the previous frame to obtain the final feature representation.
[0105] The classification subnetwork 43 can be composed of multiple convolutional layers and activation layers. Through the classification subnetwork 43, the probability maps of text in the feature maps of different frames after warping can be calculated; that is, it can be used to obtain features. Figure 1 ,feature Figure 2 and characteristics Figure 3 The corresponding text probabilities Figure 1 Probability of words Figure 2 and the probability of words Figure 3 .
[0106] Weighted subnetwork 44 is used to obtain features. Figure 1 ,feature Figure 2 and characteristics Figure 3 The corresponding weights are weight one, weight two, and weight three, respectively.
[0107] The dot product component 45 is used to calculate the dot product similarity between weight 1 and weight 2, weight 2 and weight 2, and weight 3 and weight 2, respectively.
[0108] The aggregation weight component 46 is used to calculate the aggregation weight of each frame image (image 1, image 2 and image 3), and finally aggregate the frames in time to obtain the aggregated probability map.
[0109] In the implementation process, the video can first be extracted into single-frame images, and the current frame and adjacent frames can be established as network inputs. The optical flow information of adjacent frames and the current frame can be pre-calculated using the Lucas-Kanade algorithm. Then, three consecutive frames and the optical flow information of the current frame and the frames before and after it can be used as inputs to the model to form the input dataset.
[0110] In some embodiments, training the above model can involve pre-loading the VGG subnetwork to train the network parameters, then creating a dataset of video inputs, inputting the data into the network, updating the network parameters, and continuing until the loss converges. Here, since the non-text regions in normal video frames are much larger than the text regions, using binary classification cross-entropy loss will bias the results towards non-text regions; therefore, DiceCoefficient can be used as the loss function.
[0111] In some embodiments, during the prediction stage, the video can first be divided into frames, the current frame image and adjacent frame images can be obtained, the optical flow information between the current frame image and adjacent frame images can be calculated, the trained model can be input, the probability map of the current frame image can be obtained, the probability map can be binarized, it can be determined whether the pixel is a text region, and the smallest rectangle covering the text region can be generated, which is the text region.
[0112] In this embodiment, the current frame image and the previous and next frame images are respectively extracted using a feature extraction network. The extracted features of the previous and next frame images are mapped to the features of the current frame image by using the pre-calculated optical flow information of the previous and next frame images and the current frame image. This compensates for the image feature shift caused by motion in the previous and next frame images and supplements the feature expression method of the current frame image.
[0113] In this embodiment, a classification subnetwork and a weight subnetwork are designed to calculate the features of the preceding and following frames after they have been wrapped with optical flow information. The classification subnetwork calculates the probability maps of text in the feature images of different frames after the wrapping. The weight subnetwork calculates the weights of the current frame image and the preceding and following frames after the wrapping. Finally, the weights and probability maps are aggregated, assigning different aggregation weights to the three frame feature probability maps, effectively enhancing the formation of text detection features.
[0114] Text detection in videos differs from text detection in single images. Text in videos often suffers from motion blur, defocus, and brightness variations, and not every frame in a video is suitable for still image detection. Effective spatiotemporal modeling for videos presents a significant challenge. This application addresses both the temporal and spatial aspects of video by extracting features from the current frame and its preceding and following frames. The extracted features are then warped based on the optical flow information of the preceding and following frames, resulting in a warped feature map. A classification subnetwork is used to calculate the text probability map corresponding to this feature map. Weights are then calculated based on the warped feature map, and a weighted summation is performed to form the final probabilistic representation of the detected text. This process enhances the feature information of the current frame. Compared to existing technologies, this approach differs from single-frame detection methods by utilizing both classification and weighted subnetworks, and by strengthening the optical flow information of preceding and following frames, thus enriching the feature representation of the current frame and significantly improving the accuracy of text detection in videos.
[0115] Based on the foregoing embodiments, this application provides a video text detection device, which includes various modules, each module including sub-modules, which can be implemented by a processor in an electronic device; of course, it can also be implemented by specific logic circuits; in the implementation process, the processor can be a central processing unit (CPU), microprocessor (MPU), digital signal processor (DSP) or field programmable gate array (FPGA), etc.
[0116] Figure 5 This is a schematic diagram of the composition structure of the video text detection device provided in the embodiments of this application, as shown below. Figure 5 As shown, the device 500 includes:
[0117] The acquisition module 510 is used to acquire three single-frame images based on a preset sampling interval, wherein the three single-frame images include image one, image two and image three, and image two is located between image one and image three;
[0118] Extraction module 520 is used to obtain initial features of image one using a feature extraction network. Figure 1 Features of Image Two Figure 2 and the initial features of image three Figure 3 ;
[0119] The first determining module 530 is used to determine the first optical flow information of the first image and the second image, and the second optical flow information of the third image and the second image;
[0120] The first alignment module 540 is used to align the initial features Figure 1Alignment to the feature based on the first optical flow information Figure 2 , to obtain features Figure 1 ;
[0121] The second alignment module 550 is used to align the initial features Figure 3 Alignment to the feature based on the second optical flow information Figure 2 , to obtain features Figure 3 ;
[0122] The second determining module 560 is configured to determine based on the features Figure 1 Text probability Figure 1 And weight 1, the features mentioned Figure 2 Text probability Figure 2 And weight 2, the features mentioned Figure 3 Text probability Figure 3 The text region of the second image is determined by weight three.
[0123] In some embodiments, the first alignment module 540 is further configured to, based on the first optical flow information, utilize a bilinear interpolation algorithm combined with the initial features. Figure 1 and the features Figure 2 , to obtain features Figure 1 Correspondingly, the second alignment module 550 is further configured to, based on the second optical flow information, utilize the bilinear interpolation algorithm combined with the initial features. Figure 3 and the features Figure 2 , to obtain features Figure 3 .
[0124] In some embodiments, the second determining module 560 includes a detection submodule, an acquisition submodule, a dot product submodule, a first determining submodule, and a second determining submodule, wherein the detection submodule is used to detect the features respectively using a classification subnetwork. Figure 1 The features Figure 2 and the features Figure 3 The features are obtained. Figure 1 Text probability Figure 1 The features Figure 2 Text probability Figure 2 and the features Figure 3 Text probability Figure 3 The acquisition submodule is used to acquire the features respectively using the weighted subnetwork. Figure 1 Weight 1, the aforementioned features Figure 2 The weights of the two features Figure 3 The weight three; the dot product submodule is used to dot product the weight one and the weight two, the weight two and the weight two, and the weight three and the weight two, respectively, to obtain the feature. Figure 1 First dot product similarity, the features Figure 2 The second dot product similarity and the features Figure 3 The third dot product similarity; the first determining submodule, used to determine the similarity based on the features. Figure 1 Text probability Figure 1 Similarity to the first dot product, the features Figure 2 Text probability Figure 2 Second dot product similarity, the features Figure 3 Text probability Figure 3 The second determination submodule is used to determine the target probability map of the image based on the third dot product similarity.
[0125] In some embodiments, the first determining submodule includes a first determining unit and a weighted summation unit, wherein the first determining unit is configured to, based on the features... Figure 1 Text probability Figure 1 Similarity to the first dot product, the features Figure 2 Text probability Figure 2 Second dot product similarity, the features Figure 3 Text probability Figure 3 The features are determined accordingly by the third dot product similarity. Figure 1 The first aggregation weight, the feature Figure 2 The second aggregation weight and the feature Figure 3 The third aggregation weight; the weighted summation unit is used to calculate the text probability based on the first aggregation weight, the second aggregation weight, and the third aggregation weight. Figure 1 The probability of the text Figure 2 and the probability of the text Figure 3 The target probability map is obtained by weighted summation.
[0126] In some embodiments, the second determining submodule includes a binarization unit, a second determining unit, and a third determining unit, wherein the binarization unit is used to binarize the target probability map; the second determining unit is used to determine text pixels in the target probability map based on the binarized target probability map; and the third determining unit is used to determine the text region based on the text pixels.
[0127] In some embodiments, the video text detection model includes at least the feature extraction network, the classification sub-network, and the weight sub-network. The feature extraction network includes a VGG sub-network and an FPN sub-network. The device further includes a training module and a measurement module. The training module is used to train the VGG sub-network, the FPN sub-network, the classification sub-network, and the weight sub-network using a training dataset to update the parameters of the VGG sub-network, the FPN sub-network, the classification sub-network, and the weight sub-network. The measurement module is used to measure the error between the text region prediction result and the text region target result using the DiceCoefficient loss function during training until the error converges.
[0128] The description of the above device embodiments is similar to that of the above method embodiments, and has similar beneficial effects. For technical details not disclosed in the device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.
[0129] It should be noted that, in the embodiments of this application, if the above methods are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of software products. These computer software products are stored in a storage medium and include several instructions to cause electronic devices (such as mobile phones, tablets, laptops, desktop computers, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.
[0130] Correspondingly, embodiments of this application provide a storage medium storing a computer program thereon, which, when executed by a processor, implements the steps in the video text detection method provided in the above embodiments.
[0131] Correspondingly, embodiments of this application provide an electronic device, Figure 6 A schematic diagram of a hardware entity of an electronic device provided in an embodiment of this application, such as... Figure 6 As shown, the hardware entity of the device 600 includes a memory 601 and a processor 602. The memory 601 stores a computer program that can run on the processor 602. When the processor 602 executes the program, it implements the steps in the video text detection method provided in the above embodiments.
[0132] The memory 601 is configured to store instructions and applications executable by the processor 602, and can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data and video communication data) of the processor 602 and various modules in the electronic device 600, and can be implemented by flash memory or random access memory (RAM).
[0133] It should be noted that the descriptions of the storage medium and device embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0134] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely descriptive and do not represent the superiority or inferiority of the embodiments.
[0135] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0136] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.
[0137] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.
[0138] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0139] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.
[0140] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device (which may be a mobile phone, tablet computer, laptop computer, desktop computer, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, magnetic disks, or optical disks.
[0141] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0142] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0143] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.
[0144] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A video text detection method, characterized in that, The method includes: Three single-frame images are acquired based on a preset sampling interval. The three single-frame images include image 1, image 2, and image 3, with image 2 located between image 1 and image 3. The initial feature map 1 of image 1, the feature map 2 of image 2, and the initial feature map 3 of image 3 are obtained using a feature extraction network; Determine the first optical flow information of image one and image two, and the second optical flow information of image three and image two; The initial feature map 1 is aligned to the feature map 2 based on the first optical flow information to obtain feature map 1; The initial feature map 3 is aligned to the feature map 2 based on the second optical flow information to obtain feature map 3; The feature map 1, feature map 2, and feature map 3 are detected by a classification sub-network respectively to obtain the text probability map 1 of feature map 1, the text probability map 2 of feature map 2, and the text probability map 3 of feature map 3; The weights of feature map 1, feature map 2, and feature map 3 are obtained by using a weighted subnetwork. By dot-producting the weights 1 and 2, 2 and 3 respectively, we obtain the first dot-product similarity of feature map 1, the second dot-product similarity of feature map 2, and the third dot-product similarity of feature map 3. The target probability map of image 2 is determined based on the text probability map 1 and the first dot product similarity of feature map 1, the text probability map 2 and the second dot product similarity of feature map 2, and the text probability map 3 and the third dot product similarity of feature map 3. The text region of image two is determined based on the target probability map.
2. The method as described in claim 1, characterized in that, The step of aligning the initial feature map one to the feature map two based on the first optical flow information to obtain feature map one includes: Based on the first optical flow information, the first feature map is obtained by combining the first initial feature map and the second feature map using the bilinear interpolation algorithm; Correspondingly, aligning the initial feature map three to the feature map two based on the second optical flow information to obtain feature map three includes: Based on the second optical flow information, the feature map 3 is obtained by combining the initial feature map 3 and the feature map 2 using the bilinear interpolation algorithm.
3. The method as described in claim 1, characterized in that, The process of determining the target probability map of image two based on the text probability map one and the first dot product similarity of feature map one, the text probability map two and the second dot product similarity of feature map two, and the text probability map three and the third dot product similarity of feature map three includes: Based on the text probability map 1 and the first dot product similarity of feature map 1, the text probability map 2 and the second dot product similarity of feature map 2, and the text probability map 3 and the third dot product similarity of feature map 3, the first aggregation weight of feature map 1, the second aggregation weight of feature map 2 and the third aggregation weight of feature map 3 are determined accordingly. The target probability map is obtained by weighted summation of the first aggregation weight, the second aggregation weight, and the third aggregation weight on the text probability map 1, the text probability map 2, and the text probability map 3.
4. The method as described in claim 3, characterized in that, The method of determining the first aggregation weight of feature map 1, the second aggregation weight of feature map 2, and the third aggregation weight of feature map 3 based on the text probability map 1 and the first dot product similarity, the text probability map 2 and the second dot product similarity of feature map 2, and the text probability map 3 and the third dot product similarity of feature map 3, respectively, includes: based on Calculate the first aggregation weight of the feature map 1; based on Calculate the second aggregation weight of the second feature map; based on Calculate the third aggregation weight of the feature map 3; in, This represents the first aggregation weight. This represents the second aggregation weight. This represents the third aggregation weight. This represents the probability graph of the text. This represents the probability graph of the aforementioned text. This represents the probability diagram of the text in Figure 3; This represents the first dot product similarity. This represents the second dot product similarity. This represents the third dot product similarity.
5. The method as described in claim 3, characterized in that, The step of determining the text region of image two based on the target probability map includes: Binarize the target probability map; Based on the binarized target probability map, determine the text pixels in the target probability map; The text region is determined based on the text pixels.
6. The method as described in claim 1, characterized in that, The video text detection model includes at least the feature extraction network, the classification subnetwork, and the weight subnetwork. The feature extraction network includes a VGG subnetwork and an FPN subnetwork. Training the model includes: The VGG subnetwork, FPN subnetwork, classification subnetwork, and weight subnetwork are trained using the training dataset to update the parameters of the VGG subnetwork, the FPN subnetwork, the classification subnetwork, and the weight subnetwork. During training, the DiceCoefficient loss function is used to measure the error between the predicted text region and the target text region until the error converges.
7. A video text detection device, characterized in that, The device includes: The acquisition module is used to acquire three single-frame images based on a preset sampling interval, wherein the three single-frame images include image one, image two and image three, and image two is located between image one and image three; The extraction module is used to obtain the initial feature map 1 of image 1, the feature map 2 of image 2, and the initial feature map 3 of image 3 using a feature extraction network; The first determining module is used to determine the first optical flow information of the first image and the second image, and the second optical flow information of the third image and the second image; The first alignment module is used to align the initial feature map one to the feature map two based on the first optical flow information to obtain feature map one. The second alignment module is used to align the initial feature map 3 to the feature map 2 based on the second optical flow information to obtain feature map 3; The second determining module is used to determine the text region of the second image based on the text probability map 1 and weight 1 of the first feature map, the text probability map 2 and weight 2 of the second feature map, and the text probability map 3 and weight 3 of the third feature map; wherein, the second determining module includes a detection submodule, an acquisition submodule, a dot product submodule, a first determining submodule, and a second determining submodule; The detection submodule is used to detect the feature map 1, the feature map 2, and the feature map 3 respectively using the classification subnetwork to obtain the text probability map 1 of the feature map 1, the text probability map 2 of the feature map 2, and the text probability map 3 of the feature map 3; The acquisition submodule is used to acquire the weight one of feature map one, the weight two of feature map two, and the weight three of feature map three using the weight subnetwork. The dot product submodule is used to dot product the weight one and the weight two, the weight two and the weight two, and the weight three and the weight two, respectively, to obtain the first dot product similarity of the feature map one, the second dot product similarity of the feature map two, and the third dot product similarity of the feature map three. The first determining submodule is used to determine the target probability map of image two based on the text probability map one and the first dot product similarity of feature map one, the text probability map two and the second dot product similarity of feature map two, and the text probability map three and the third dot product similarity of feature map three. The second determining submodule is used to determine the text region of the second image based on the target probability map.
8. An electronic device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 6.
9. A storage medium, characterized in that, It stores executable instructions for causing a processor to execute, thereby implementing the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Frame level feature aggregation method for video target detection
CN109993095A
Biometric recognition for uncontrolled acquisition environments
US20190080068A1