Video data processing method, device and equipment

Through the neural network model trained by using attention mechanism and shape-aware loss function in video text detection, the problem of low text detection efficiency in the prior art is solved, and efficient and accurate text detection is achieved.

CN114067237BActive Publication Date: 2025-05-13TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111264126.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-28
Publication Date
2025-05-13
Estimated Expiration
2041-10-28

AI Technical Summary

Technical Problem

The existing video text detection methods have poor results when detecting curved text, and the segmentation-based method requires post-processing, resulting in slow detection speed, resulting in low text detection efficiency.

Method used

A neural network model trained based on attention mechanism and shape-aware loss function is used for video text detection. The amount of parameters and calculations are reduced through attention mechanism, and the shape-aware loss function improves the accuracy of detection.

Benefits of technology

It improves the speed and accuracy of text detection, solves the problems of poor detection effect and slow speed, and realizes efficient text detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114067237B_ABST
    Figure CN114067237B_ABST
Patent Text Reader

Abstract

The present application provides a video data processing method, device and equipment, which relates to computer technology. The method includes: obtaining a video to be detected, wherein the video to be detected includes multiple texts; detecting the text in the video to be detected according to a preset text detection model, wherein the text detection model is obtained by training a neural network model according to an attention mechanism and a preset shape-aware loss function; and outputting a video containing a text detection frame according to the detected text, wherein the text detection frame is used to indicate the position of the text in the video. The method of the present application can solve the problem that accuracy and speed cannot be taken into account at the same time in text detection. While achieving high-accuracy text detection, it greatly improves the speed of text detection, is more suitable for practical applications, and solves the technical problem of low efficiency in text detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to technology, and in particular to a method, device and equipment for processing video data. Background Art

[0002] At present, with the continuous development of computer vision technology, scene text detection technology is also constantly improving. Scene text detection refers to the task of annotating text in pictures or videos using visual text boxes. It is a basic and key task in the field of computer vision and a key step in introducing post-processing methods. Post-processing methods include text recognition, text retrieval, license plate recognition, text visualization question and answer, etc. Therefore, it is necessary to recognize the text in the video.

[0003] In the prior art, when recognizing text in a video, the text in the video is usually detected based on a deep learning neural network model. The text detection methods based on deep learning can be divided into two categories: regression-based text detection methods and segmentation-based text detection methods. Among them, the regression-based text detection method regards the text as the target to be detected, and obtains the text detection frame through direct regression; the segmentation-based text detection method classifies the pixels of the image, identifies whether the pixels of the image belong to text, and then combines the post-processing method to obtain the final text detection frame.

[0004] However, in the prior art, in the regression-based text detection method, due to the limitation of the shape of the text detection box, the detection effect of curved text is very poor. In the segmentation-based text detection method, it is necessary to combine the post-processing method to obtain the final text detection box, which will reduce the speed of detecting text. Therefore, the existing text detection method will lead to poor detection effect or slow detection speed, which in turn leads to low efficiency in detecting text. Summary of the invention

[0005] The present application provides a video data processing method, device and equipment to solve the technical problem of low efficiency in text detection.

[0006] In a first aspect, the present application provides a video data processing method, comprising:

[0007] Acquire a video to be detected, wherein the video to be detected includes multiple texts;

[0008] Detecting text in the video to be detected according to a preset text detection model, wherein the text detection model is obtained by training a neural network model according to an attention mechanism and a preset shape-aware loss function;

[0009] According to the detected text, a video including a text detection frame is output, wherein the text detection frame is used to mark the position of the text in the video.

[0010] Further, detecting the text in the video to be detected according to a preset text detection model includes:

[0011] Detecting the text in the video to be detected using the attention mechanism of a preset text detection model;

[0012] The area of ​​the pixel block where the text is located is determined using a preset shape-aware loss function.

[0013] Furthermore, according to the detected text, a video containing a text detection frame is output, including:

[0014] According to the area of ​​the pixel block where the text is located, generating a text detection frame equal to the area;

[0015] According to the text detection frame, a video including the text detection frame is output.

[0016] Furthermore, the video to be detected is a real-time video or an offline video.

[0017] Furthermore, the video to be detected is a real-time video; outputting a video containing a text detection frame according to the detected text includes:

[0018] As the real-time video plays each frame, a real-time image including a text detection frame corresponding to each frame is output to obtain a video including a text detection frame; wherein the text detection frame is specifically used to mark the position of the text in each frame.

[0019] Furthermore, the video to be detected is an offline video; outputting a video containing a text detection frame according to the detected text includes:

[0020] According to the text detection frame contained in each frame of the offline video, a video containing the text detection frame is output; wherein the text detection frame is specifically used to mark the position of the text in each frame.

[0021] Furthermore, the method further comprises:

[0022] Acquire multiple image data sets, wherein the image data sets include multiple texts;

[0023] Setting a shape-aware loss function; wherein the shape-aware loss function includes a text part loss function, a text core part loss function and a pixel vector part loss function;

[0024] Based on the shape-aware loss function, the neural network model is trained using the image data set until the shape-aware loss function reaches a minimum value, thereby obtaining a trained text detection model.

[0025] Furthermore, the image data set includes horizontal text, inclined text and text of arbitrary shape.

[0026] Furthermore, a shape-aware loss function is set, including:

[0027] Acquire a first actual image including text label values ​​of pixels where text is located, background label values ​​of pixels where background other than the text is located in the image dataset, and a first predicted image including the text label values ​​and the background label values ​​detected by the text detection model;

[0028] Compare the first actual image with the first predicted image for similarity, obtain a first similarity value between the text label value and the background label value in the first actual image and the text label value and the background label value in the first predicted image, and set the first similarity value as the value of the text part loss function;

[0029] According to the first similarity value and the text label value of the pixel where the text is located, adjusting the text label value of the pixel where the text is located corresponding to the attention mechanism;

[0030] Acquire a second actual image including a text label value of a pixel where a text core is located, a background label value of a pixel where a background other than the text core is located in the image data set, and a second predicted image including the text label value and the background label value detected by the text detection model;

[0031] Compare the second actual image with the second predicted image for similarity, obtain a second similarity value between the text label value and the background label value in the second actual image and the text label value and the background label value in the second predicted image, and set the second similarity value as the value of the text core part loss function;

[0032] According to the second similarity value and the text label value of the pixel where the text core is located, adjusting the text label value of the pixel where the text core is located corresponding to the attention mechanism;

[0033] Determine a third similarity value according to the number of the texts in the image data set, the number of pixels in the pixel block where each of the texts is located, and the average value of the feature vectors of the pixels of each of the texts, and set the third similarity value as the value of the pixel vector partial loss function;

[0034] According to the third similarity value, the area of ​​each of the texts corresponding to the attention mechanism is adjusted.

[0035] Further, based on the shape-aware loss function, the neural network model is trained using the image data set until the shape-aware loss function reaches a minimum value, thereby obtaining a trained text detection model, including:

[0036] Based on the shape-aware loss function, training a neural network model using the image dataset;

[0037] Determine the first similarity value, the second similarity value, and the third similarity value included in the shape-aware loss function, and determine a weighted sum of the first similarity value, the second similarity value, and the third similarity value according to a hyperparameter corresponding to the second similarity value and a hyperparameter corresponding to the third similarity value;

[0038] When the weighted sum reaches a minimum value, a trained text detection model is obtained.

[0039] In a second aspect, the present application provides a video data processing device, comprising:

[0040] A first acquisition unit is used to acquire a video to be detected, wherein the video to be detected includes a plurality of texts;

[0041] A detection unit, used to detect text in the video to be detected according to a preset text detection model, wherein the text detection model is obtained by training a neural network model according to an attention mechanism and a preset shape-aware loss function;

[0042] The output unit is used to output a video containing a text detection frame based on the detected text, and the text detection frame is used to mark the position of the text in the video.

[0043] Furthermore, the detection unit comprises:

[0044] A detection module, used to detect the text in the video to be detected by using the attention mechanism of a preset text detection model;

[0045] The first determination module is used to determine the area of ​​the pixel block where the text is located by using a preset shape-aware loss function.

[0046] Furthermore, the output unit comprises:

[0047] A generating module, used for generating a text detection frame having an area equal to the area of ​​the pixel block where the text is located;

[0048] An output module is used to output a video containing a text detection frame according to the text detection frame.

[0049] Furthermore, the video to be detected is a real-time video or an offline video.

[0050] Furthermore, the video to be detected is a real-time video; the output unit is specifically used to:

[0051] As the real-time video plays each frame, a real-time image including a text detection frame corresponding to each frame is output to obtain a video including a text detection frame; wherein the text detection frame is specifically used to mark the position of the text in each frame.

[0052] Furthermore, the video to be detected is an offline video; and the output unit is specifically used to:

[0053] According to the text detection frame contained in each frame of the offline video, a video containing the text detection frame is output; wherein the text detection frame is specifically used to mark the position of the text in each frame.

[0054] Furthermore, the device also includes:

[0055] A second acquisition unit, used to acquire a plurality of image data sets, wherein the image data sets include a plurality of texts;

[0056] A setting unit, used to set a shape-aware loss function; wherein the shape-aware loss function includes a text part loss function, a text core part loss function and a pixel vector part loss function;

[0057] A training unit is used to train a neural network model based on the shape-aware loss function using the image data set until the shape-aware loss function reaches a minimum value, thereby obtaining a trained text detection model.

[0058] Furthermore, the image data set includes horizontal text, inclined text and text of arbitrary shape.

[0059] Furthermore, the setting unit includes:

[0060] A first acquisition module is used to acquire a first actual image including text label values ​​of pixels where text is located, background label values ​​of pixels where background other than the text is located in the image data set, and a first predicted image including the text label values ​​and the background label values ​​detected by the text detection model;

[0061] A first setting module is used to compare the first actual image with the first predicted image for similarity, obtain a first similarity value between the text label value and the background label value in the first actual image and the text label value and the background label value in the first predicted image, and set the first similarity value as the value of the text part loss function;

[0062] A first adjustment module, configured to adjust the text label value of the pixel where the text is located corresponding to the attention mechanism according to the first similarity value and the text label value of the pixel where the text is located;

[0063] A second acquisition module is used to acquire a second actual image including a text label value of a pixel where a text core is located, a background label value of a pixel where background other than the text core is located in the image data set, and a second predicted image including the text label value and the background label value detected by the text detection model;

[0064] A second setting module is used to compare the second actual image with the second predicted image for similarity, obtain a second similarity value between the text label value and the background label value in the second actual image and the text label value and the background label value in the second predicted image, and set the second similarity value as the value of the text core part loss function;

[0065] A second adjustment module, used for adjusting the text label value of the pixel where the text core is located corresponding to the attention mechanism according to the second similarity value and the text label value of the pixel where the text core is located;

[0066] A third setting module is used to determine a third similarity value according to the number of the texts in the image data set, the number of pixels in the pixel block where each of the texts is located, and the average value of the number of pixels of each of the texts, and set the third similarity value as the value of the pixel vector partial loss function;

[0067] The third adjustment module is used to adjust the area of ​​each of the texts corresponding to the attention mechanism according to the third similarity value.

[0068] Furthermore, the training unit comprises:

[0069] A training module, configured to train a neural network model using the image dataset based on the shape-aware loss function;

[0070] a second determination module, configured to determine the first similarity value, the second similarity value, and the third similarity value included in the shape-aware loss function, and determine a weighted sum of the first similarity value, the second similarity value, and the third similarity value according to a hyperparameter corresponding to the second similarity value and a hyperparameter corresponding to the third similarity value;

[0071] The third determination module is used to obtain the trained text detection model when the weighted sum obtains a minimum value.

[0072] In a third aspect, the present application provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, and when the processor executes the computer program, the method described in the first aspect is implemented.

[0073] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the method described in the first aspect.

[0074] In a fifth aspect, the present application provides a computer program product, including a computer program, which implements the method described in the first aspect when executed by a processor.

[0075] The video data processing method, device and equipment provided by the present application obtain a video to be detected, wherein the video to be detected includes multiple texts; the text in the video to be detected is detected according to a preset text detection model, wherein the text detection model is obtained by training a neural network model according to an attention mechanism and a preset shape-aware loss function; according to the detected text, a video containing a text detection frame is output, and the text detection frame is used to mark the position of the text in the video. In this scheme, since the text detection model is obtained by training a neural network model according to an attention mechanism and a preset shape-aware loss function, the text in the video to be detected can be detected according to the preset text detection model, the position of the text in the video to be detected is marked using the text detection frame, and the corresponding video containing the text detection frame is output. Therefore, introducing the attention mechanism when training the neural network model can reduce the number of parameters and calculations, and improve the overall speed of the text detection method. Introducing a preset shape-aware loss function can more accurately distinguish pixels of different texts and pixels of the same text, and then train an optimized text detection model. Using the optimized text detection model to detect the video to be detected can solve the problem that accuracy and speed cannot be taken into account at the same time in text detection. While achieving high-accuracy text detection, it greatly improves the speed of text detection, is more suitable for practical applications, and solves the technical problem of low efficiency in text detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0077] Figure 1A schematic diagram of a flow chart of a video data processing method provided in an embodiment of the present application;

[0078] Figure 2 A flowchart of another video data processing method provided in an embodiment of the present application;

[0079] Figure 3 A schematic diagram of the structure of an attention module provided in this application;

[0080] Figure 4 A structural diagram of a feature deepening module and a feature fusion module provided in this application;

[0081] Figure 5 A schematic diagram of the structure of a video data processing device provided in an embodiment of the present application;

[0082] Figure 6 A schematic diagram of the structure of another video data processing device provided in an embodiment of the present application;

[0083] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0084] Figure 8 A block diagram of an electronic device provided in an embodiment of the present application.

[0085] The above drawings show clear embodiments of the present disclosure, which will be described in more detail below. These drawings and text descriptions are not intended to limit the scope of the present disclosure in any way, but to illustrate the concepts of the present disclosure to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0086] Here, exemplary embodiments will be described in detail, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present disclosure.

[0087] In one example, with the continuous development of computer vision technology, scene text detection technology is also constantly improving. Scene text detection refers to the task of annotating text in a picture or video using a visual text box. It is a basic and key task in the field of computer vision and a key step in introducing post-processing methods. Post-processing methods include text recognition, text retrieval, license plate recognition, text visualization question and answer, etc. Therefore, it is necessary to recognize the text in the video. In the prior art, when recognizing text in a video, the text in the video is usually detected based on a deep learning neural network model. The text detection method based on deep learning can be divided into two categories: regression-based text detection method and segmentation-based text detection method. Among them, the regression-based text detection method regards the text as the target to be detected and obtains the text detection box through direct regression; the segmentation-based text detection method classifies the pixels of the image, identifies whether the pixels of the image belong to the text, and then combines the post-processing method to obtain the final text detection box. However, in the prior art, in the regression-based text detection method, due to the limitation of the shape of the text detection box, the detection effect of curved text is very poor. In the segmentation-based text detection method, it is necessary to combine the post-processing method to obtain the final text detection box, which will reduce the speed of detecting text. Therefore, the existing text detection method will lead to poor detection effect or slow detection speed, which in turn leads to low efficiency in detecting text.

[0088] The video data processing method, device and equipment provided in this application are intended to solve the above technical problems in the prior art.

[0089] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0090] Figure 1 A flowchart of a video data processing method provided in an embodiment of the present application is shown as follows: Figure 1 As shown, this includes:

[0091] 101. Obtain a video to be detected, where the video to be detected includes multiple texts.

[0092] For example, the execution subject of this embodiment may be an electronic device, or a terminal device, or a video data processing device or device, or other devices or equipment that can execute this embodiment, and there is no limitation on this. In this embodiment, the execution subject is described as an electronic device.

[0093] First, it is necessary to obtain the video to be detected. The video to be detected can be captured, or obtained from a storage device; or obtained from a web page, or received from other devices. The video to be detected includes multiple texts.

[0094] 102. Detect text in the video to be detected according to a preset text detection model, wherein the text detection model is obtained by training a neural network model according to an attention mechanism and a preset shape-aware loss function.

[0095] Exemplarily, the electronic device may train a neural network model according to an attention mechanism and a preset shape-aware loss function to obtain a text detection model, and then input the video to be detected into the text detection model, and detect the video to be detected through the text detection model.

[0096] 103. Output a video containing a text detection frame according to the detected text, where the text detection frame is used to mark the position of the text in the video.

[0097] Exemplarily, the electronic device may output a video containing a text detection frame based on the detected text. The text detection frame is used to mark the position of the text in the video. The text detection frame may be a polygonal frame such as a rectangle or a square.

[0098] In an embodiment of the present application, a video to be detected is obtained, and the video to be detected includes multiple texts. The text in the video to be detected is detected according to a preset text detection model, wherein the text detection model is obtained by training a neural network model according to an attention mechanism and a preset shape-aware loss function. According to the detected text, a video containing a text detection frame is output, and the text detection frame is used to mark the position of the text in the video. In this solution, since the text detection model is obtained by training a neural network model according to an attention mechanism and a preset shape-aware loss function, the text in the video to be detected can be detected according to the preset text detection model, and the position of the text in the video to be detected is marked using the text detection frame, and the corresponding video containing the text detection frame is output. Therefore, introducing the attention mechanism when training the neural network model can reduce the number of parameters and calculations, and improve the overall speed of the text detection method. Introducing a preset shape-aware loss function can more accurately distinguish pixels of different texts and pixels of the same text, and then train an optimized text detection model. Using the optimized text detection model to detect the video to be detected can solve the problem that accuracy and speed cannot be taken into account at the same time in text detection. While achieving high-accuracy text detection, it greatly improves the speed of text detection, is more suitable for practical applications, and solves the technical problem of low efficiency in text detection.

[0099] Figure 2A flowchart of another video data processing method provided in an embodiment of the present application is shown as follows: Figure 2 As shown, the method includes:

[0100] 201. Obtain multiple image data sets, where the image data sets include multiple texts.

[0101] In one example, the image dataset includes horizontal text, tilted text, and text of arbitrary shapes.

[0102] Exemplarily, the electronic device acquires multiple image data sets, which may include videos and images, etc. The image data sets include multiple types of data sets, including ICDAR2013, ICDAR2015, CTW1500 and TotalText. The text detection boxes of the first two are limited to rectangular boxes, among which the ICDAR2013 data set focuses on horizontal text, and the detection box is parallel to the image border. The ICDAR2015 data set focuses on tilted text, and the detection box allows tilting; the latter two focus on text of arbitrary shapes, and the text detection box can be of arbitrary shapes. Through horizontal and vertical comparisons on different types of data sets, the applicability of the detection method in various situations can be reflected.

[0103] 202. Obtain a first actual image including text label values ​​of pixels where text is located, background label values ​​of pixels where background other than text is located in an image dataset, and a first predicted image including text label values ​​and background label values ​​detected by a text detection model.

[0104] Exemplarily, the electronic device may first process an image data set, mark the corresponding text label values ​​for the pixels where the text is located, for example, the text label value is 1, and mark the corresponding background label values ​​for the pixels where the background other than the text is located in the image data set, for example, the background label value is 0, thereby obtaining a first actual image containing the text label values ​​and the background label values; then, the text label values ​​and the background label values ​​in the image data set are detected by a text detection model, thereby obtaining a first predicted image containing the text label values ​​and the background label values.

[0105] 203. Compare the first actual image with the first predicted image to obtain a first similarity value between the text label value and the background label value in the first actual image and the text label value and the background label value in the first predicted image, and set the first similarity value as the value of the text part loss function.

[0106] Exemplarily, the text part loss function can reflect the information of the location of the text learned by the text detection model. The text part loss function can adopt a dice loss function (Dice Loss). The electronic device can compare the first actual image with the first predicted image for similarity, obtain a first similarity value between the text label value and the background label value in the first actual image and the text label value and the background label value in the first predicted image, and set the first similarity value as the value of the text part loss function. The formula of the dice loss function is as follows:

[0107]

[0108] Among them, P text (i) represents the text label value of the text part at pixel i in the first predicted image, G text (i) represents the text label value of the text part at pixel i in the first actual image, Loss text Represents the first similarity value.

[0109] 204. According to the first similarity value and the text label value of the pixel where the text is located, adjust the text label value of the pixel where the text is located corresponding to the attention mechanism.

[0110] Exemplarily, the electronic device can determine the accuracy of the text detection model predicting the text based on the first similarity value. When the accuracy is high, the text label value of the pixel where the text is located corresponding to the attention mechanism can be adjusted according to the text label value of the pixel where the text is located, so that the attention mechanism focuses on the text label value of the pixel where the text is located.

[0111] 205. Obtain a second actual image including text label values ​​of pixels where the text core is located, background label values ​​of pixels where the background other than the text core is located in the image dataset, and a second predicted image including text label values ​​and background label values ​​detected by a text detection model.

[0112] Exemplarily, the electronic device may first process the image data set, predetermine the text core, mark the corresponding text label value for the pixel where the text core is located, for example, the text label value is 2, and mark the corresponding background label value for the background pixels other than the text core in the image data set, for example, the background label value is 3, thereby obtaining a second actual image containing the text label value and the background label value; then, the text label value and the background label value in the image data set are detected by a text detection model, thereby obtaining a second predicted image containing the text label value and the background label value.

[0113] 206. Compare the second actual image with the second predicted image for similarity, obtain a second similarity value between the text label value and the background label value in the second actual image and the text label value and the background label value in the second predicted image, and set the second similarity value as the value of the loss function of the core part of the text.

[0114] Exemplarily, the text core part loss function can reflect the information of the text detection model learning the location of the text core. The text core part loss function can adopt a dice loss function. The electronic device can compare the second actual image with the second predicted image for similarity, obtain a second similarity value between the text label value and the background label value in the second actual image and the text label value and the background label value in the second predicted image, and set the second similarity value as the value of the text core part loss function. The formula of the dice loss function is as follows:

[0115]

[0116] Among them, P kernel (i) represents the text label value of the text core part at pixel i in the second predicted image, G kernel (i) represents the text label value of the core part of the text at pixel i in the second actual image, Loss kernel Represents the second similarity value.

[0117] 207. According to the second similarity value and the text label value of the pixel where the text core is located, adjust the text label value of the pixel where the text core is located corresponding to the attention mechanism.

[0118] Exemplarily, the electronic device can determine the accuracy of the text detection model in predicting the text core based on the second similarity value. When the accuracy is high, the text label value of the pixel where the text core is located corresponding to the attention mechanism can be adjusted according to the text label value of the pixel where the text core is located, so that the attention mechanism focuses on the text label value of the pixel where the text core is located.

[0119] 208. Determine a third similarity value according to the number of texts in the image data set, the number of pixels in the pixel block where each text is located, and the average value of the feature vector of the pixels of each text, and set the third similarity value as the value of the pixel vector partial loss function.

[0120] For example, the pixel vector partial loss function can reflect the information that the text detection model has learned to distinguish pixels belonging to different texts. The formula of the pixel vector partial loss function is as follows:

[0121] Loss embedding =L agg +L dis

[0122]

[0123]

[0124] Among them, Lagg makes the features of pixels belonging to the same text closer, and L dis Make the features of pixels belonging to different texts farther apart, N represents the number of texts contained in the image dataset, T i represents the number of pixels in the i-th text, μ i represents the average value of the feature vectors of all pixels in the i-th text, η, γ both represent distance thresholds, which are hyperparameters set in advance, and W scale(i) and W dist(i,j) is the correlation coefficient of shape perception, and the specific calculation formula is as follows:

[0125]

[0126]

[0127] Among them, h and w represent the height and width of the image in the image dataset, diag(T i ) represents the text T i The diagonal length of centredist(T i , T j ) represents the text T i and T j The distance from the center point.

[0128] 209. According to the third similarity value, adjust the area of ​​each text corresponding to the attention mechanism.

[0129] Exemplarily, the electronic device may determine the accuracy of the text detection model predicting the text based on the third similarity value, and when the accuracy is high, the attention mechanism may be adjusted to distinguish different pixels so that the attention mechanism focuses on distinguishing texts belonging to different pixels.

[0130] 210. Based on shape-aware loss function, the neural network model is trained using image datasets.

[0131] Exemplarily, the electronic device can train the neural network model based on the shape-aware loss function using four types of data sets in the image data set, and use the accuracy, recall rate, F1 score, and detection speed as indicators for statistical analysis, and compare the optimal model effects with and without each key module (the feature deepening module and the feature fusion module are replaced by multi-layer CNN+pooling layers, and the shape-aware loss function replaces the shape coefficient by setting it to 1). During training, the neural network feature extraction backbone network and related configurations can be set, and the neural network feature extraction backbone network can be set to the residual network ResNet-18. The related configurations are the hyperparameters of each module of the system and the data loading interface settings. The attention mechanism is used to guide the training of the feature deepening module and the feature fusion module, strengthen the learning of image text features, and based on the shape-aware loss function, the image data set is used to train the neural network model, wherein the feature deepening module and the feature fusion module are CNN+attention pooling structures.

[0132] like Figure 3 As shown, Figure 3 A structural diagram of an attention module provided in this application includes: low-level features and deepening features. A series of vector merging, pooling layers and multi-layer CNN processing steps are performed on the low-level features to obtain deepening features; Figure 4 As shown, Figure 4 The structural diagram of a feature deepening module and a feature fusion module provided in this application includes: original features, feature deepening module, feature fusion module and features obtained after fusion. The hierarchical feature graph is obtained through backbone network extraction. In order to improve its quality, the attention mechanism, feature deepening module and feature fusion module are introduced. The core module (i.e., attention module) of the attention mechanism is used according to Figure 2 In the design, high-level features and low-level features are merged and passed through the pooling layer and multi-layer CNN. After the sigmoid function is activated, the attention coefficient is multiplied by the original feature, which is equivalent to weighted screening of useful information in the original features and combining high-level features as deeper features. Figure 3 The structure will Figure 2 The attention modules and convolutional layers in the network are combined hierarchically to obtain the feature deepening module and feature fusion module applied in the system.

[0133] 211. Determine a first similarity value, a second similarity value, and a third similarity value included in a shape-aware loss function, and determine a weighted sum of the first similarity value, the second similarity value, and the third similarity value based on a hyperparameter corresponding to the second similarity value and a hyperparameter corresponding to the third similarity value.

[0134] Exemplarily, the electronic device may first determine the first similarity value, the second similarity value, and the third similarity value included in the shape perception loss function, and determine the weighted sum of the first similarity value, the second similarity value, and the third similarity value according to the hyperparameter corresponding to the second similarity value and the hyperparameter corresponding to the third similarity value. The formula for calculating the weighted sum is as follows:

[0135] Loss=Loss text +αLoss kernel +βLoss embedding

[0136] Among them, α and β are hyperparameters for balancing various loss functions, α is the hyperparameter corresponding to the second similarity value, and β is the hyperparameter corresponding to the third similarity value.

[0137] 212. When the weighted sum reaches the minimum value, the trained text detection model is obtained.

[0138] Exemplarily, the test results of accuracy, recall rate, F1 score, detection speed, feature deepening module and feature fusion module, and shape perception loss function are as follows:

[0139] Table 1 Test results of ICDAR2013 dataset

[0140]

[0141] Table 2 Test results of ICDAR2015 dataset

[0142]

[0143] Table 3 Test results of CTW1500 dataset

[0144]

[0145] Table 4 TotalText dataset test results

[0146]

[0147] From the test results of the above four data sets, we can see that:

[0148] In terms of speed, the feature deepening module, feature fusion module, and shape-aware loss function have greatly accelerated text detection, shortening the detection time of an image by 4-6ms. The combined speed-up of the two does not increase linearly, but increases slightly to 4-7ms.

[0149] In terms of indicators, the impact of the two modules is not significant, so that the overall method remains within a relatively stable level of fluctuation, and the accuracy can be maintained at a high level of 85-90%, which fully meets the needs of text detection in real life. From the perspective of the more comprehensive F1 score, the combination of the two modules brings about a 1% improvement. Therefore, when the weighted sum reaches the minimum value, the neural network model converges, and then a trained text detection model is obtained.

[0150] 213. Obtain a video to be detected, where the video to be detected includes multiple texts.

[0151] In one example, the video to be detected is a real-time video or an offline video.

[0152] Exemplarily, the real-time video may be a video input by a real-time camera, and the offline video may be a cached video.

[0153] 214. Use the attention mechanism of the preset text detection model to detect the text in the video to be detected.

[0154] Exemplarily, since the target data that the attention mechanism focuses on has been gradually adjusted according to the loss function during the process of training the text detection model, the electronic device can use the attention mechanism of the preset text detection model to detect the text in the video to be detected.

[0155] 215. The area of ​​the pixel block where the text is located is determined using a preset shape-aware loss function.

[0156] For example, since the font of the text may be very large or very small, in order to achieve a better detection effect, the electronic device may use a preset shape-aware loss function to determine the area of ​​the pixel block where the text is located.

[0157] 216. Generate a text detection frame equal to the area of ​​the pixel block where the text is located according to the area of ​​the pixel block; and output a video including the text detection frame according to the text detection frame.

[0158] Step 216 includes two methods:

[0159] The first method of step 216: the video to be detected is a real-time video; as the real-time video plays each frame, the real-time picture containing the text detection frame corresponding to each frame is output to obtain a video containing the text detection frame; wherein the text detection frame is specifically used to mark the position of the text in each frame.

[0160] The second method of step 216: the video to be detected is an offline video; according to the text detection frame contained in each frame of the offline video, a video containing the text detection frame is output; wherein the text detection frame is specifically used to mark the position of the text in each frame.

[0161] Exemplarily, the electronic device may generate a text detection frame having an area equal to the area of ​​the pixel block where the text is located, and then output a video containing the text detection frame according to the text detection frame. There are two ways to output the video containing the text detection frame:

[0162] In the first method, the video to be detected is a real-time video; when the video is input and played in real time according to a camera or other device, as each frame of the real-time video is played, the electronic device outputs a real-time picture containing a text detection box corresponding to each frame to obtain a video containing a text detection box, wherein the text detection box is specifically used to indicate the position of the text in each frame.

[0163] In the second method, the video to be detected is an offline video; when the offline video is played, the electronic device outputs a video containing a text detection frame based on the text detection frame contained in each frame of the offline video, wherein the text detection frame is specifically used to mark the position of the text in each frame.

[0164] In an embodiment of the present application, multiple image data sets are obtained, and the image data sets include multiple texts; a first actual image including text label values ​​of pixels where texts are located, background label values ​​of pixels where backgrounds other than texts are located in the image data set, and a first predicted image including text label values ​​and background label values ​​detected by a text detection model are obtained. The first actual image is compared with the first predicted image for similarity, and a first similarity value between the text label value and background label value in the first actual image and the text label value and background label value in the first predicted image is obtained, and the first similarity value is set as the value of the text part loss function. According to the first similarity value and the text label value of the pixel where the text is located, the text label value of the pixel where the text corresponding to the attention mechanism is located is adjusted. A second actual image including the text label value of the pixel where the text core is located, the background label value of the pixel where the background other than the text core is located in the image data set, and a second predicted image including text label values ​​and background label values ​​detected by a text detection model are obtained. The second actual image is compared with the second predicted image for similarity, and a second similarity value between the text label value and background label value in the second actual image and the text label value and background label value in the second predicted image is obtained, and the second similarity value is set as the value of the text core part loss function. According to the second similarity value and the text label value of the pixel where the text core is located, the text label value of the pixel where the text core is located corresponding to the attention mechanism is adjusted. According to the number of texts in the image data set, the number of pixels in the pixel block where each text is located, and the average value of the feature vector of the pixel of each text, the third similarity value is determined, and the third similarity value is set as the value of the pixel vector partial loss function. According to the third similarity value, the area of ​​each text corresponding to the attention mechanism is adjusted. Based on the shape-aware loss function, the neural network model is trained using the image data set. The first similarity value, the second similarity value and the third similarity value included in the shape-aware loss function are determined, and the weighted sum of the first similarity value, the second similarity value and the third similarity value is determined according to the hyperparameter corresponding to the second similarity value and the hyperparameter corresponding to the third similarity value. When the weighted sum obtains the minimum value, the trained text detection model is obtained. A video to be detected is obtained, and the video to be detected includes multiple texts. The text in the video to be detected is detected using the attention mechanism of the preset text detection model. The area of ​​the pixel block where the text is located is determined using the preset shape-aware loss function. According to the area of ​​the pixel block where the text is located, a text detection frame equal to the area is generated; according to the text detection frame, a video containing the text detection frame is output. Therefore, using the trained and optimized text detection model to detect the video to be detected can solve the problem that accuracy and speed cannot be taken into account at the same time in text detection. While achieving high-accuracy text detection, it greatly improves the speed of text detection, is more suitable for practical applications, and solves the technical problem of low efficiency in text detection.

[0165] Figure 5A schematic diagram of the structure of a video data processing device provided in an embodiment of the present application is shown in FIG. Figure 5 As shown, the device comprises:

[0166] The first acquisition unit 51 is used to acquire a video to be detected, where the video to be detected includes a plurality of texts.

[0167] The detection unit 52 is used to detect text in the video to be detected according to a preset text detection model, wherein the text detection model is obtained by training a neural network model according to an attention mechanism and a preset shape-aware loss function.

[0168] The output unit 53 is used to output a video containing a text detection frame based on the detected text, where the text detection frame is used to mark the position of the text in the video.

[0169] The device of this embodiment can execute the technical solution in the above method. Its specific implementation process and technical principles are the same and will not be repeated here.

[0170] Figure 6 A schematic diagram of the structure of another video data processing device provided in an embodiment of the present application, in Figure 5 Based on the embodiment shown, Figure 6 As shown, the detection unit 52 includes:

[0171] The detection module 521 is used to detect the text in the video to be detected by using the attention mechanism of the preset text detection model.

[0172] The first determination module 522 is used to determine the area of ​​the pixel block where the text is located by using a preset shape-aware loss function.

[0173] In one example, the output unit 53 includes:

[0174] The generating module 531 is used to generate a text detection frame having the same area as the pixel block where the text is located according to the area of ​​the pixel block where the text is located.

[0175] The output module 532 is used to output a video containing a text detection frame according to the text detection frame.

[0176] In one example, the video to be detected is a real-time video or an offline video.

[0177] In one example, the video to be detected is a real-time video; the output unit 53 is specifically used to:

[0178] As each frame of the real-time video is played, a real-time image containing a text detection frame corresponding to each frame is output to obtain a video containing a text detection frame; wherein the text detection frame is specifically used to mark the position of the text in each frame.

[0179] In one example, the video to be detected is an offline video; the output unit 53 is specifically used to:

[0180] According to the text detection frame contained in each frame of the offline video, a video containing the text detection frame is output; wherein the text detection frame is specifically used to mark the position of the text in each frame.

[0181] In one example, the device further includes:

[0182] The second acquisition unit 61 is used to acquire multiple image data sets, where the image data sets include multiple texts.

[0183] The setting unit 62 is used to set the shape-aware loss function; wherein the shape-aware loss function includes a text part loss function, a text core part loss function and a pixel vector part loss function.

[0184] The training unit 63 is used to train the neural network model based on the shape-aware loss function using the image data set until the shape-aware loss function reaches a minimum value, thereby obtaining a trained text detection model.

[0185] In one example, the image dataset includes horizontal text, tilted text, and text of arbitrary shapes.

[0186] In one example, the setting unit 62 includes:

[0187] The first acquisition module 621 is used to obtain a first actual image containing text label values ​​of pixels where text is located, background label values ​​of pixels where background other than text is located in an image data set, and a first predicted image containing text label values ​​and background label values ​​detected by a text detection model.

[0188] The first setting module 622 is used to compare the first actual image with the first predicted image to obtain a first similarity value between the text label value and the background label value in the first actual image and the text label value and the background label value in the first predicted image, and set the first similarity value as the value of the text part loss function.

[0189] The first adjustment module 623 is used to adjust the text label value of the pixel where the text is located corresponding to the attention mechanism according to the first similarity value and the text label value of the pixel where the text is located.

[0190] The second acquisition module 624 is used to obtain a second actual image including text label values ​​of pixels where the text core is located, background label values ​​of pixels where the background other than the text core is located in the image data set, and a second predicted image including text label values ​​and background label values ​​detected by the text detection model.

[0191] The second setting module 625 is used to compare the second actual image with the second predicted image to obtain a second similarity value between the text label value and the background label value in the second actual image and the text label value and the background label value in the second predicted image, and set the second similarity value as the value of the loss function of the text core part.

[0192] The second adjustment module 626 is used to adjust the text label value of the pixel where the text core is located corresponding to the attention mechanism according to the second similarity value and the text label value of the pixel where the text core is located.

[0193] The third setting module 627 is used to determine a third similarity value according to the number of texts in the image data set, the number of pixels in the pixel block where each text is located, and the average number of pixels of each text, and set the third similarity value as the value of the pixel vector partial loss function.

[0194] The third adjustment module 628 is used to adjust the area of ​​each text corresponding to the attention mechanism according to the third similarity value.

[0195] In one example, the training unit 63 includes:

[0196] The training module 631 is used to train the neural network model using the image dataset based on the shape-aware loss function.

[0197] The second determination module 632 is used to determine the first similarity value, the second similarity value and the third similarity value included in the shape-aware loss function, and determine the weighted sum of the first similarity value, the second similarity value and the third similarity value according to the hyperparameters corresponding to the second similarity value and the hyperparameters corresponding to the third similarity value.

[0198] The third determination module 633 is used to obtain the trained text detection model when the weighted sum obtains a minimum value.

[0199] The device of this embodiment can execute the technical solution in the above method. Its specific implementation process and technical principles are the same and will not be repeated here.

[0200] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown in FIG. Figure 7 As shown, the electronic device includes: a memory 71, a processor 72;

[0201] The memory 71 stores a computer program that can be executed on the processor 72 .

[0202] The processor 72 is configured to execute the method provided in the above embodiment.

[0203] The electronic device further includes a receiver 73 and a transmitter 74. The receiver 73 is used to receive instructions and data sent by an external device, and the transmitter 74 is used to send instructions and data to the external device.

[0204] Figure 8 This is a block diagram of an electronic device provided in an embodiment of the present application. The electronic device may be a mobile phone, a computer, a digital broadcast terminal, a message transceiver, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0205] The device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .

[0206] The processing component 802 generally controls the overall operation of the device 800, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above-mentioned method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.

[0207] The memory 804 is configured to store various types of data to support operations on the device 800. Examples of such data include instructions for any application or method operating on the device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0208] The power supply component 806 provides power to the various components of the device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 800.

[0209] The multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.

[0210] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), and when the device 800 is in an operating mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 804 or sent via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.

[0211] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: home button, volume button, start button, and lock button.

[0212] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the device 800. For example, the sensor assembly 814 can detect the open / closed state of the device 800, the relative positioning of components, such as the display and keypad of the device 800, the sensor assembly 814 can also detect the position change of the device 800 or a component of the device 800, the presence or absence of user contact with the device 800, the orientation or acceleration / deceleration of the device 800 and the temperature change of the device 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor or a temperature sensor.

[0213] The communication component 816 is configured to facilitate wired or wireless communication between the device 800 and other devices. The device 800 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0214] In an exemplary embodiment, the apparatus 800 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to perform the above method.

[0215] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the instructions can be executed by the processor 820 of the device 800 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0216] An embodiment of the present application also provides a non-temporary computer-readable storage medium. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the method provided by the above embodiment.

[0217] An embodiment of the present application also provides a computer program product, which includes: a computer program, which is stored in a readable storage medium, and at least one processor of an electronic device can read the computer program from the readable storage medium, and at least one processor executes the computer program so that the electronic device executes the solution provided by any of the above embodiments.

[0218] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0219] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A video data processing method, characterized in that: include: Acquire a video to be detected, wherein the video to be detected includes multiple texts; Detecting text in the video to be detected according to a preset text detection model, wherein the text detection model is obtained by training a neural network model according to an attention mechanism and a preset shape-aware loss function; Outputting a video containing a text detection frame according to the detected text, wherein the text detection frame is used to mark the position of the text in the video; The method further comprises: Acquire multiple image data sets, wherein the image data sets include multiple texts; Setting a shape-aware loss function; wherein the shape-aware loss function includes a text part loss function, a text core part loss function and a pixel vector part loss function; Based on the shape-aware loss function, the neural network model is trained using the image data set until the shape-aware loss function reaches a minimum value, thereby obtaining a trained text detection model; The shape-aware loss function is set, including: Acquire a first actual image including text label values ​​of pixels where text is located, background label values ​​of pixels where background other than the text is located in the image dataset, and a first predicted image including the text label values ​​and the background label values ​​detected by the text detection model; Compare the first actual image with the first predicted image for similarity, obtain a first similarity value between the text label value and the background label value in the first actual image and the text label value and the background label value in the first predicted image, and set the first similarity value as the value of the text part loss function; According to the first similarity value and the text label value of the pixel where the text is located, adjusting the text label value of the pixel where the text is located corresponding to the attention mechanism; Acquire a second actual image including a text label value of a pixel where a text core is located, a background label value of a pixel where a background other than the text core is located in the image data set, and a second predicted image including the text label value and the background label value detected by the text detection model; Compare the second actual image with the second predicted image for similarity, obtain a second similarity value between the text label value and the background label value in the second actual image and the text label value and the background label value in the second predicted image, and set the second similarity value as the value of the text core part loss function; According to the second similarity value and the text label value of the pixel where the text core is located, adjusting the text label value of the pixel where the text core is located corresponding to the attention mechanism; Determine a third similarity value according to the number of the texts in the image data set, the number of pixels in the pixel block where each of the texts is located, and the average value of the feature vectors of the pixels of each of the texts, and set the third similarity value as the value of the pixel vector partial loss function; According to the third similarity value, the area of ​​each of the texts corresponding to the attention mechanism is adjusted.

2. The method according to claim 1, characterized in that Detecting the text in the video to be detected according to a preset text detection model includes: Detecting the text in the video to be detected using the attention mechanism of a preset text detection model; The area of ​​the pixel block where the text is located is determined using a preset shape-aware loss function.

3. The method according to claim 2, characterized in that Based on the detected text, output a video containing the text detection box, including: According to the area of ​​the pixel block where the text is located, generating a text detection frame equal to the area; According to the text detection frame, a video including the text detection frame is output.

4. The method according to claim 1, characterized in that: The video to be detected is a real-time video or an offline video.

5. The method according to claim 4, characterized in that The video to be detected is a real-time video; according to the detected text, a video containing a text detection frame is output, including: As the real-time video plays each frame, a real-time image including a text detection frame corresponding to each frame is output to obtain a video including a text detection frame; wherein the text detection frame is specifically used to mark the position of the text in each frame.

6. The method according to claim 4, characterized in that The video to be detected is an offline video; according to the detected text, a video containing a text detection frame is output, including: According to the text detection frame contained in each frame of the offline video, a video containing the text detection frame is output; wherein the text detection frame is specifically used to mark the position of the text in each frame.

7. The method according to claim 1, characterized in that The image data set includes horizontal text, inclined text and text of arbitrary shape.

8. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program running on the processor, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

Citation Information

Patent Citations

  • Pixel-driven mobile phone operation interface text detection method

    CN110991440A

  • Text box detection method and device, electronic equipment and computer storage medium

    CN112232315A