Model Training Method, Text Detection Method, Dictionary Pen, and Storage Medium
By introducing a guideable module into the text detection model, supervising and learning the text center area and the complete text area is solved, and the problem of inaccurate detection results in mid- and low-end devices is achieved, efficient and accurate text detection is achieved.
Patent Information
- Application Number
- CN202210546240.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-19
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-05-19
AI Technical Summary
The existing fast text detection scheme has low accuracy in detection results in low-end devices and lacks supervision of the complete text area restoration process.
The guideable module is introduced into the text detection model. By predicting the probability map of the text center area and performing binarization processing, combined with the guideable module, the area expansion operation is carried out to realize the supervision and learning of the text center area and the complete text area.
Improve the accuracy and efficiency of text detection, avoid the problem of boundary inaccuracy caused by unsupervised during the restoration of the complete text area, and eliminate the need for complex post-processing processes.
Smart Images

Figure CN114943970B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of computer technologies, and in particular, to a model training method, a text detection method, a dictionary pen, and a storage medium. Background Art
[0002] As a pre-step of computer vision tasks, text detection is used to locate the position of text in a picture, so as to facilitate subsequent text recognition, image search, and so on.
[0003] With the continuous development of text detection technologies, various detection methods applicable to different detection scenarios have evolved to meet different detection requirements. For example: a fast text detection solution suitable for deployment in mid- to low-end devices with limited computing resources (such as dictionary pens, scanning pens, etc.). The characteristics of this solution are: fewer computing parameters involved, a simple post-processing process, and the ability to quickly and efficiently obtain detection results.
[0004] However, in related fast text detection solutions, there is usually a problem of low accuracy of detection results. Summary of the Invention
[0005] In view of this, embodiments of the present application provide a model training method, a text detection method, a dictionary pen, and a storage medium to at least partially solve the above problems.
[0006] According to a first aspect of embodiments of the present application, there is provided a model training method, including:
[0007] Obtaining a sample image containing text and a sample label, where the sample label includes: a complete text region label and a text center region label;
[0008] Inputting the sample image into an initial text detection model, obtaining a sample feature vector through a feature extraction module of the initial text detection model; through a prediction module of the initial text detection model, obtaining a text center region probability map based on the sample feature vector, and performing binarization processing on the text center region probability map to obtain a binarized map; through a derivable module of the initial text detection model, performing a region dilation operation on the binarized map to obtain complete text region prediction information;
[0009] Obtaining a loss value based on the text center region probability map, the text center region label, the complete text region prediction information, and the complete text region label; and training the initial text detection model based on the loss value to obtain a trained text detection model.
[0010] According to a second aspect of embodiments of the present application, there is provided another model training method, including:
[0011] Obtain an initial text detection model, a first set of sample images, and a first set of sample labels. The initial text detection model includes a feature extraction module, a prediction module, and a differentiable module;
[0012] Through the feature extraction module, extract the sample feature vectors of each first sample image in the first set of sample images; through the prediction module, obtain a text center region probability map based on each sample feature vector, and obtain a binary map based on the text center region probability map; through the differentiable module, perform a region dilation operation on each binary map to obtain complete text region prediction information;
[0013] Based on the text center region probability map, the complete text region prediction information, and the first set of sample labels, obtain a loss value; and train the initial text detection model based on the loss value to obtain a transitional text detection model;
[0014] Obtain a second set of sample images and a second set of sample labels, and train the transitional text detection model based on the second set of sample labels to obtain a trained text detection model; the number of complex background images in the first set of sample images is greater than the number of complex background images in the second set of sample images.
[0015] According to the third aspect of the embodiments of the present application, a text detection method is provided, including:
[0016] Obtain a target text image to be detected and a pre-trained text detection model. The text detection model includes: a feature extraction module, a prediction module, and a differentiable module; the differentiable module is a module that can perform gradient transmission during the training process of the text detection model;
[0017] Input the target text image into the text detection model, obtain a target feature vector through the feature extraction module; through the prediction module, obtain a text center region probability map based on the target feature vector, and perform binary processing on the text center region probability map to obtain a binary map; through the differentiable module, perform a region dilation operation on the binary map to obtain a text detection result.
[0018] According to the fourth aspect of the embodiments of the present application, an application to a dictionary pen is provided, including:
[0019] Receive an instruction for indicating text detection, and scan a target region containing text according to the instruction to obtain a target text image;
[0020] Obtain a pre-trained text detection model. The text detection model includes: a feature extraction module, a prediction module, and a differentiable module; the differentiable module is a module that can perform gradient transmission during the training process of the text detection model;
[0021] Input the target text image into the text detection model, and obtain a target feature vector through the feature extraction module; through the prediction module, obtain a text center region probability map based on the target feature vector, and perform binarization processing on the text center region probability map to obtain a binarized map; through the differentiable module, perform a region dilation operation on the binarized map to obtain a text detection result.
[0022] According to the fifth aspect of the embodiments of the present application, a model training device is provided, including:
[0023] A sample acquisition module, configured to acquire a sample image and a sample label containing text, where the sample label includes: a complete text region label and a text center region label;
[0024] A prediction information obtaining module, configured to input the sample image into an initial text detection model, and obtain a sample feature vector through the feature extraction module of the initial text detection model; through the prediction module of the initial text detection model, obtain a text center region probability map based on the sample feature vector, and perform binarization processing on the text center region probability map to obtain a binarized map; through the differentiable module of the initial text detection model, perform a region dilation operation on the binarized map to obtain complete text region prediction information;
[0025] A training module, configured to obtain a loss value based on the text center region probability map, the text center region label, the complete text region prediction information, and the complete text region label; and train the initial text detection model based on the loss value to obtain a trained text detection model.
[0026] According to the sixth aspect of the embodiments of the present application, a model training device is provided, including:
[0027] A first sample acquisition module, configured to acquire an initial text detection model, a first sample image set, and a first sample label set, where the initial text detection model includes a feature extraction module, a prediction module, and a differentiable module;
[0028] A first prediction information obtaining module, configured to extract a sample feature vector of each first sample image in the first sample image set through the feature extraction module; through the prediction module, obtain a text center region probability map based on each sample feature vector, and obtain a binarized map based on the text center region probability map; through the differentiable module, perform a region dilation operation on each binarized map to obtain complete text region prediction information;
[0029] A transition model obtaining module, configured to obtain a loss value based on the text center region probability map, the complete text region prediction information, and the first sample label set; and train the initial text detection model based on the loss value to obtain a transition text detection model;
[0030] A text detection model obtaining module, configured to obtain a second sample image set and a second sample label set, and train the transition text detection model based on the second sample label set to obtain a trained text detection model; the number of complex background images in the first sample image set is greater than the number of complex background images in the second sample image set.
[0031] According to a seventh aspect of the embodiments of the present application, there is provided a text detection device, including:
[0032] A target image obtaining module, configured to obtain a target text image to be detected and a pre-trained text detection model, where the text detection model includes: a feature extraction module, a prediction module, and a differentiable module; the differentiable module is a module capable of gradient transmission during the training process of the text detection model;
[0033] A first result obtaining module, configured to input the target text image into the text detection model, obtain a target feature vector through the feature extraction module; through the prediction module, obtain a text center region probability map based on the target feature vector, and perform binarization processing on the text center region probability map to obtain a binarized map; through the differentiable module, perform a region dilation operation on the binarized map to obtain a text detection result.
[0034] According to an eighth aspect of the embodiments of the present application, there is provided a text detection device applied to a dictionary pen, including:
[0035] An instruction receiving module, configured to receive an instruction for indicating text detection, and scan a target region containing text according to the instruction to obtain a target text image;
[0036] A model obtaining module, configured to obtain a pre-trained text detection model, where the text detection model includes: a feature extraction module, a prediction module, and a differentiable module; the differentiable module is a module capable of gradient transmission during the training process of the text detection model;
[0037] A second result obtaining module, configured to input the target text image into the text detection model, obtain a target feature vector through the feature extraction module; through the prediction module, obtain a text center region probability map based on the target feature vector, and perform binarization processing on the text center region probability map to obtain a binarized map; through the differentiable module, perform a region dilation operation on the binarized map to obtain a text detection result.
[0038] According to a ninth aspect of the embodiments of the present application, a dictionary pen is provided, including: an image collector, a memory, and a processor; the image collector is configured to scan a target area containing text to obtain a target text image; the memory is used to store executable instructions; when the processor executes the executable instructions stored on the memory, the target text image is input into a pre-trained text detection model, and a target feature vector is obtained through a feature extraction module in the text detection model; through a prediction module in the text detection model, a text center region probability map is obtained based on the target feature vector, and the text center region probability map is binarized to obtain a binarized map; through a derivable module in the text detection model, a region dilation operation is performed on the binarized map to obtain a text detection result.
[0039] According to a tenth aspect of the embodiments of the present application, a computer storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, it implements the model training method as described in the first aspect or the second aspect, or the text detection method as described in the third aspect or the fourth aspect.
[0040] According to the model training, text detection method, dictionary pen, and storage medium provided by the embodiments of the present application, a derivable module is added to the text detection model. During model training, on the one hand, a text center region probability map is obtained through a prediction model and then a binarized map (corresponding to a text center region mask) is obtained, and a loss value is calculated based on the center region probability map and the text center region label to perform supervised learning on the prediction process of the text center region; on the other hand, a region dilation is also performed on the binarized map through the derivable module to predict the position information of the complete text region (complete text region prediction information), and a loss value is obtained again based on the predicted position information of the complete text region and the complete text region label to perform supervised learning on the reduction process of restoring the complete text region from the text center region. That is to say, in the embodiments of the present application, during the text detection model training process, not only the prediction process of the text center region in the early stage is supervised, but also the reduction process of the complete text region in the later stage is supervised. Therefore, problems such as inaccurate reduction results and inaccurate complete text region boundaries caused by the unsupervised reduction process of the complete text region can be avoided, and the accuracy of text detection is improved.
[0041] In addition, for text detection based on the text detection model of the embodiments of the present application, after the text image to be detected is input into the text detection model, the text detection result can be obtained through the model without a complex post-processing process. Therefore, the embodiments of the present application improve the efficiency of text detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments described in the embodiments of the present application. For those of ordinary skill in the art, other accompanying drawings can also be obtained based on these drawings.
[0043] Figure 1 It is a flowchart of the steps of a model training method according to Embodiment 1 of the present application;
[0044] Figure 2 It is Figure 1 a schematic diagram of a scenario example in the illustrated embodiment;
[0045] Figure 3 It is a flowchart of the steps of a model training method according to Embodiment 2 of the present application;
[0046] Figure 4 It is Figure 3 a schematic diagram of a scenario example in the illustrated embodiment;
[0047] Figure 5 It is a flowchart of the steps of a model training method according to Embodiment 3 of the present application;
[0048] Figure 6 It is a flowchart of the steps of a text detection method according to Embodiment 4 of the present application;
[0049] Figure 7 It is a flowchart of the steps of a text detection method according to Embodiment 5 of the present application;
[0050] Figure 8 It is a structural block diagram of a model training device according to Embodiment 6 of the present application;
[0051] Figure 9 It is a structural block diagram of a model training device according to Embodiment 7 of the present application;
[0052] Figure 10 It is a structural block diagram of a text detection device according to Embodiment 8 of the present application;
[0053] Figure 11 It is a structural block diagram of a text detection device according to Embodiment 9 of the present application;
[0054] Figure 12 It is a structural schematic diagram of a dictionary pen according to Embodiment 10 of the present application;
[0055] Figure 13 It is a structural schematic diagram of an electronic device according to Embodiment 11 of the present application. Detailed implementation manners
[0056] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art shall fall within the scope of protection of the embodiments of the present application.
[0057] The following further illustrates the specific implementation of the embodiments of the present application in conjunction with the accompanying drawings of the embodiments of the present application.
[0058] Embodiment 1
[0059] Refer to Figure 1 , Figure 1 , which is a step flowchart of a model training method according to Embodiment 1 of the present application. Specifically, the model training method provided in this embodiment includes the following steps:
[0060] Step 102, obtain a sample image containing text and a sample label.
[0061] Among them, the sample label includes: a complete text region label and a text center region label.
[0062] Specifically, the complete text region label can be the true position information representing the region where the complete text is located in the sample image. The text center region label can be the information representing the position of the center region of the text in the sample image. In the embodiments of the present application, the acquisition methods for the complete text region label and the text center region label are not limited. For example: the complete text region label in the sample image can be obtained by manual annotation; for the complete text region label, the reverse operation of the image dilation operation performed by the differentiable model in the embodiments of the present application is then used to obtain the text center region label, and so on.
[0063] Step 104, input the sample image into the initial text detection model, obtain a sample feature vector through the feature extraction module of the initial text detection model; through the prediction module of the initial text detection model, obtain a text center region probability map based on the sample feature vector, and perform binarization processing on the text center region probability map to obtain a binarized map; through the differentiable module of the initial text detection model, perform a region dilation operation on the binarized map to obtain complete text region prediction information.
[0064] Among them, the value of each pixel point in the text center region probability map is used to represent the probability that each pixel point belongs to the text center region.
[0065] In addition, in the embodiments of the present application, when binarizing the probability map of the text center region, it can be performed by using a preset fixed probability threshold, that is: in the probability map of the text center region, the values of the pixel points with probability values greater than the preset fixed probability threshold are converted to 1, and the values of the pixel points with probability values less than or equal to the preset fixed probability threshold are converted to 0, so as to obtain a binarized map; other binarization processing methods can also be used. For example, when obtaining the probability map of the center region, a threshold map of the center region can also be obtained. The value of each pixel point in this map is used to represent the probability threshold when binarizing the corresponding pixel point in the probability map of the text center region. Furthermore, during binarization processing, based on the threshold map of the center region and the probability map of the center region, differentiable binarization processing (Differentiable Binarization algorithm) is performed to obtain a binarized map.
[0066] Among them, by performing differentiable binarization processing with the help of the threshold map of the center region, more accurate text shape information (binarized map) can be obtained from the adaptive threshold of differentiable binarization. Therefore, the efficiency of text detection can be effectively improved.
[0067] The differentiable module in the embodiments of the present application can be any module that can perform image dilation operations and can perform gradient transmission during model training. That is to say, during model training, based on the output information of the differentiable module - the prediction information of the complete text region and the corresponding label - the label of the complete text region, a differentiable function can be formed as part of the loss function for supervised training of the model.
[0068] Step 106, based on the probability map of the text center region, the label of the text center region, the prediction information of the complete text region, and the label of the complete text region, obtain a loss value; and train the initial text detection model based on the loss value to obtain a trained text detection model.
[0069] In the present application, there is no limitation on the specific method of obtaining the loss value based on the probability map of the text center region, the label of the text center region, the prediction information of the complete text region, and the label of the complete text region. It can be calculated according to the set loss function formula according to actual needs.
[0070] Optionally, in some embodiments, the above loss value can be obtained in the following manner, and then a trained text detection model can be obtained:
[0071] Based on the probability map of the text center region and the label of the text center region, obtain a first loss value;
[0072] Based on the prediction information of the complete text region and the label of the complete text region, obtain a second loss value;
[0073] Fuse the first loss value and the second loss value to obtain a fused loss value;
[0074] Train the initial text detection model based on the fused loss value to obtain a trained text detection model.
[0075] See Figure 2 , Figure 2 which is a schematic diagram of the scenario corresponding to the first embodiment of this application. Hereinafter, with reference to the schematic diagram shown in Figure 2 a specific scenario example will be used to illustrate the embodiments of this application:
[0076] After obtaining a sample image, input it into the feature extraction module of the initial text detection model to perform feature extraction through this module, so as to obtain a sample feature vector; then through the prediction module, perform prediction based on the sample feature vector to obtain a text center region probability map; perform a binarization operation on the above text center region probability map to obtain a binarized map (that is, obtain a mask of the text center region); then input the binarized map into the differentiable module, and perform a region dilation operation through the differentiable module to obtain complete text region prediction information (that is, the predicted position information of the complete text region); afterwards, calculate the loss value, and perform training on the parameters of each module in the initial text detection model based on the loss value. Specifically: obtain a first loss value based on the text center region probability map and the text center region label; obtain a second loss value based on the complete text region prediction information and the complete text region label, and then fuse the above first loss value and second loss value (the specific fusion method is not limited) to obtain a fused loss value, so as to train the parameters of each module in the initial text detection model to obtain a trained text detection model.
[0077] According to the model training method provided by the embodiments of the present application, a differentiable module is added to the text detection model. During training, on the one hand, a text center region probability map is obtained through the prediction model, and then a binary map (corresponding to the text center region mask) is obtained. Based on the center region probability map and the text center region label, a loss value is calculated to perform supervised learning on the prediction process of the text center region. On the other hand, the binary map is also dilated in regions through the differentiable module to predict the position information of the complete text region (complete text region prediction information), and a loss value is obtained again based on the predicted complete text region position information and the complete text region label to perform supervised learning on the reduction process of restoring the complete text region from the text center region. That is to say, in the embodiments of the present application, during the text detection model training process, not only the prediction process of the early text center region is supervised, but also the reduction process of the later complete text region is supervised. Therefore, problems such as inaccurate reduction results and inaccurate complete text region boundaries caused by the unsupervised reduction process of the complete text region can be avoided, and the accuracy of text detection is improved.
[0078] In addition, when performing text detection based on the text detection model trained according to the embodiments of the present application, after the text image to be detected is input into the text detection model, the text detection result can be obtained through the model without a complex post-processing process. Therefore, the embodiments of the present application improve the efficiency of text detection.
[0079] The model training method of this embodiment can be executed by any suitable electronic device with data processing capabilities, including but not limited to: servers, PCs, etc.
[0080] Embodiment 2
[0081] Refer to Figure 3 , Figure 3 which is a flowchart of the steps of a model training method according to Embodiment 2 of the present application. Specifically, the model training method provided in this embodiment includes the following steps:
[0082] Step 302, obtain a sample image containing text and a sample label.
[0083] The sample label includes: a complete text region label, a text center region label, and a threshold map label.
[0084] Among them, the complete text region label can be the true position information representing the region where the complete text is located in the sample image. The text center region label can be the information representing the position of the center region of the text in the sample image; after obtaining the complete text region label and the text center region label, based on these two labels, a corresponding threshold map calculation method can be used to obtain the threshold map label. The specific acquisition method can refer to the related technology of the threshold map label and will not be elaborated here.
[0085] Step 304: Input the sample image into the initial text detection model. Obtain the sample feature vector through the feature extraction module of the initial text detection model; through the prediction module of the initial text detection model, based on the sample feature vector, obtain the text center region probability map and the text center region threshold map; based on the text center region threshold map, perform differentiable binarization on the text center region probability map to obtain a binarized map; and perform binarization on the text center region probability map to obtain a binarized map; through the differentiable module of the initial text detection model, perform region dilation operation on the binarized map to obtain the complete text region prediction information.
[0086] Among them, the value of each pixel point in the text center region probability map is used to represent the probability that the pixel point belongs to the text center region; the value of each pixel point in the center region threshold map is used to represent the probability threshold when performing binarization on the corresponding pixel point in the text center region probability map.
[0087] The differentiable module in the embodiments of the present application can be any module that can perform image dilation operation and can perform gradient transmission during model training. That is, during model training, based on the output information of the differentiable module - the complete text region prediction information and the corresponding label - the complete text region label, a differentiable function can be formed as part of the loss function for supervised training of the model.
[0088] Optionally, in some embodiments, the differentiable module can be a neural network module including a max pooling layer. Further, to improve computational efficiency, a single max pooling layer can be used as the differentiable module.
[0089] Step 306: Based on the text center region probability map and the text center region label, obtain the first loss value.
[0090] Step 308: Based on the complete text region prediction information and the complete text region label, obtain the second loss value.
[0091] Step 310: Based on the text center region threshold map and the threshold map label, obtain the third loss value.
[0092] Step 312: Fuse the first loss value, the second loss value, and the third loss value to obtain the fused loss value.
[0093] In the embodiments of the present application, no limitation is imposed on the specific fusion method, and the specific fusion method can be set according to the actual situation. For example: corresponding fusion weights can be set in advance for each loss value, and then weighted fusion can be performed to obtain the fused loss value, and so on.
[0094] Step 314: Train the initial text detection model based on the fusion loss value to obtain a trained text detection model.
[0095] See Figure 4 , Figure 4 which is the schematic diagram corresponding to the second embodiment of this application. Hereinafter, with reference to the Figure 4 schematic diagram shown, a specific scenario example will be used to illustrate the embodiments of this application:
[0096] After obtaining the sample image, input it into the feature extraction module of the initial text detection model to extract features through this module, so as to obtain a sample feature vector; then through the prediction module, perform prediction based on the sample feature vector to obtain a text center region probability map. At the same time, a text center region threshold map is also obtained; based on the obtained text center region threshold map, perform a binarization operation (a differentiable binarization operation) on the above-mentioned text center region probability map to obtain a binarized map (that is, obtain the mask of the text center region); then input the binarized map into the differentiable module, and perform a region dilation operation through the differentiable module to obtain complete text region prediction information (that is, the predicted position information of the complete text region); then, calculate the loss value, and train the parameters of each module in the initial text detection model based on the loss value. Specifically: obtain a first loss value based on the text center region probability map and the text center region label; obtain a second loss value based on the complete text region prediction information and the complete text region label; obtain a third loss value based on the text center region threshold map and the threshold map label; then fuse the above-mentioned first loss value, second loss value, and third loss value (the specific fusion method is not limited) to obtain a fusion loss value, so as to train the parameters of each module in the initial text detection model to obtain a trained text detection model.
[0097] In the embodiments of the present application, a differentiable module is added to the text detection model. During training, on the one hand, a probability map of the text center region is obtained through the prediction model, and then a binary map (corresponding to the mask of the text center region) is obtained. Based on the probability map of the center region and the label of the text center region, a loss value is calculated to perform supervised learning on the prediction process of the text center region. On the other hand, the differentiable module is also used to dilate the binary map to predict the position information of the complete text region (complete text region prediction information), and based on the predicted position information of the complete text region and the label of the complete text region, a loss value is obtained again to perform supervised learning on the restoration process of restoring the text center region to the complete text region. That is to say, in the embodiments of the present application, during the training process of the text detection model, not only the prediction process of the text center region in the early stage is supervised, but also the restoration process of the complete text region in the later stage is supervised. Therefore, it is possible to avoid problems such as inaccurate restoration results and inaccurate boundaries of the complete text region caused by the unsupervised restoration process of the complete text region, and improve the accuracy of text detection.
[0098] Meanwhile, for text detection based on the text detection model trained according to the embodiments of the present application, after the text image to be detected is input into the text detection model, the text detection result can be obtained through the model without a complex post-processing process. Therefore, the embodiments of the present application improve the efficiency of text detection.
[0099] In addition, when the probability map of the text center region is output through the prediction model, a threshold map of the text center region for binarizing the probability map of the text center region is also output, thereby converting the non-differentiable binarization process into a differentiable binarization process, facilitating the transmission of gradients during the training process for network supervised learning, and thus helping to generate more accurate prediction results. Therefore, the embodiments of the present application further improve the accuracy of the text detection result.
[0100] The model training method of this embodiment can be executed by any suitable electronic device with data processing capabilities, including but not limited to: servers, PCs, etc.
[0101] Embodiment III
[0102] Refer to Figure 5 , Figure 5 which is a flowchart of the steps of a model training method according to Embodiment III of the present application. Specifically, the model training method provided in this embodiment includes the following steps:
[0103] Step 502, obtain an initial text detection model, a first sample image set, and a first sample label set.
[0104] Among them, the initial text detection model includes a feature extraction module, a prediction module, and a differentiable module;
[0105] Step 504: Use the feature extraction module to extract the sample feature vectors of each first sample image in the first sample image set; use the prediction module to obtain the text center region probability map based on each sample feature vector, and obtain the binary map based on the text center region probability map; use the differentiable module to perform a region dilation operation on each binary map to obtain the complete text region prediction information.
[0106] Step 506: Based on the text center region probability map, the complete text region prediction information, and the first sample label set, obtain the loss value.
[0107] Step 508: Train the initial text detection model based on the loss value to obtain the transitional text detection model.
[0108] In the embodiments of the present application, when using the first sample images in the first sample image set to train the initial text detection model to obtain the transitional text detection model, each step can refer to the corresponding steps of using the sample images to train the initial text detection model to obtain the trained text detection model in Embodiment 1 or Embodiment 2, which will not be elaborated here.
[0109] Step 510: Obtain the second sample image set and the second sample label set.
[0110] Step 512: Train the transitional text detection model based on the second sample label set to obtain the trained text detection model.
[0111] Among them, the number of complex background images in the first sample image set is greater than the number of complex background images in the second sample image set.
[0112] Specifically, for example, the first sample image set may include: a large number of synthetic images containing complex backgrounds or complex patterns; the images in the second sample image set are real images in the actual text detection scenario collected.
[0113] In this step, when using the second sample images in the second sample image set to train the transitional text detection model to obtain the trained text detection model, each step can also refer to the corresponding steps of using the sample images to train the initial text detection model to obtain the trained text detection model in Embodiment 1 or Embodiment 2, which will not be elaborated here either.
[0114] In the embodiment of the present application, a differentiable module is added to the text detection model. During model training, on the one hand, a text center region probability map is obtained through the prediction model, and then a binary map (corresponding to the text center region mask) is obtained. Based on the center region probability map and the text center region label, a loss value is calculated to perform supervised learning on the prediction process of the text center region. On the other hand, the differentiable module is also used to dilate the binary map to predict the complete text region position information (complete text region prediction information), and based on the predicted complete text region position information and the complete text region label, a loss value is obtained again to perform supervised learning on the reduction process of restoring the complete text region from the text center region. That is to say, in the embodiment of the present application, during the text detection model training process, not only the prediction process of the early text center region is supervised, but also the reduction process of the later complete text region is supervised. Therefore, it is possible to avoid the problems of inaccurate reduction results and inaccurate complete text region boundaries caused by the unsupervised reduction process of the complete text region, and improve the accuracy of text detection.
[0115] In addition, in the embodiment of the present application, the model training process is split into two stages: In the first stage, the first sample image set containing more complex background images is used for model training to obtain a transitional text detection model. Since there are more complex background images in the sample images, the trained transitional text detection model is better at handling text detection tasks for complex background images. That is to say, the detection results of the trained transitional text detection model for complex background images are more accurate. In the second stage, the transitional text detection model is trained again using the real image set collected during the text detection process (the number of complex background images in this image set is small), which can make the finally trained text detection model have higher detection accuracy in the real detection scenario while having the ability to process complex background images.
[0116] Embodiment 4
[0117] Refer to Figure 6 , Figure 6 which is a step flowchart of a text detection method according to Embodiment 4 of the present application. Specifically, the text detection method provided in this embodiment includes the following steps:
[0118] Step 602, obtain the target text image to be detected and the pre-trained text detection model.
[0119] The text detection model includes: a feature extraction module, a prediction module, and a differentiable module; the differentiable module is a module that can perform gradient transmission during the training process of the text detection model.
[0120] In the embodiments of the present application, the differentiable module can be any module that can perform image dilation operations and can perform gradient transfer during model training. That is to say, during model training, based on the output information of the differentiable module - the complete text region prediction information and the corresponding label - the complete text region label, a differentiable function can be formed as part of the loss function to perform supervised training of the model.
[0121] Optionally, in some embodiments, the differentiable module can be a neural network module including a max pooling layer. Further, to improve computational efficiency, a single max pooling layer can be used as the differentiable module.
[0122] Step 604, input the target text image into the text detection model, obtain the target feature vector through the feature extraction module; through the prediction module, obtain the text center region probability map based on the target feature vector, and perform binarization processing on the text center region probability map to obtain a binarized map; through the differentiable module, perform region dilation operation on the binarized map to obtain the text detection result.
[0123] In the embodiments of the present application, the steps for text detection of the target text image can refer to the corresponding steps of training the initial text detection model with the sample image in Embodiment 1 or Embodiment 2 to obtain the trained text detection model, which will not be elaborated here.
[0124] The text detection method provided by the embodiments of the present application adds a differentiable module to the text detection model. During model training, on the one hand, the text center region probability map is obtained through the prediction model and then the binarized map (corresponding to the text center region mask) is obtained, and the loss value is calculated based on the center region probability map and the text center region label to perform supervised learning on the prediction process of the text center region; on the other hand, the differentiable module is also used to perform region dilation on the binarized map to predict the position information of the complete text region (complete text region prediction information), and the loss value is obtained again based on the predicted position information of the complete text region and the complete text region label to perform supervised learning on the reduction process of restoring the complete text region from the text center region. That is to say, in the embodiments of the present application, during the training process of the text detection model, not only the prediction process of the text center region in the early stage is supervised, but also the reduction process of the complete text region in the later stage is supervised. Therefore, it is possible to avoid problems such as inaccurate reduction results and inaccurate boundaries of the complete text region caused by the unsupervised reduction process of the complete text region, and improve the accuracy of text detection.
[0125] In addition, through the text detection method of the embodiments of the present application, after inputting the text image to be detected into the text detection model, the text detection result can be obtained through this model without a complex post-processing process. Therefore, the efficiency of text detection is improved.
[0126] Example 5
[0127] Refer to Figure 7 , Figure 7 FIG. is a flowchart of the steps of a text detection method according to Example 5 of the present application. The application scenario of this embodiment can be: a user scans a target area containing text through an offline scanning device (such as a dictionary pen or a scanning pen, etc.) to perform text detection on the scanned target text image to obtain the position information of the text area.
[0128] Specifically, the text detection method provided in this embodiment includes the following steps:
[0129] Step 702, receive an instruction for indicating text detection, and scan a target area containing text according to the instruction to obtain a target text image.
[0130] Step 704, obtain a pre-trained text detection model.
[0131] The text detection model includes: a feature extraction module, a prediction module, and a differentiable module; the differentiable module is a module that can perform gradient transmission during the training process of the text detection model.
[0132] In the embodiments of the present application, the differentiable module can be any module that can perform image dilation operation and can perform gradient transmission during the model training process. That is to say, during the model training process, based on the output information of the differentiable module - the complete text area prediction information and the corresponding label - the complete text area label, a differentiable function can be formed as part of the loss function for supervised training of the model.
[0133] Optionally, in some embodiments, the differentiable module can be a neural network module including a max pooling layer. Further, to improve the calculation efficiency, a single max pooling layer can be used as the differentiable module.
[0134] Step 706, input the target text image into the text detection model, obtain a target feature vector through the feature extraction module; through the prediction module, obtain a text center region probability map based on the target feature vector, and perform binarization processing on the text center region probability map to obtain a binarized map; through the differentiable module, perform region dilation operation on the binarized map to obtain a text detection result.
[0135] In the embodiments of the present application, the steps for text detection of the target text image can refer to the corresponding steps of training the initial text detection model with a sample image in Example 1 or Example 2 to obtain a trained text detection model, which will not be elaborated here.
[0136] The text detection method provided by the embodiments of the present application adds a differentiable module to the text detection model. During model training, on the one hand, a text center region probability map is obtained through the prediction model and then a binary map (corresponding to the text center region mask) is obtained. Based on the center region probability map and the text center region label, a loss value is calculated to perform supervised learning on the prediction process of the text center region. On the other hand, the differentiable module is also used to perform region dilation on the binary map to predict the position information of the complete text region (complete text region prediction information), and based on the predicted position information of the complete text region and the complete text region label, a loss value is obtained again to perform supervised learning on the reduction process of restoring the complete text region from the text center region. That is to say, in the embodiments of the present application, during the text detection model training process, not only the prediction process of the early text center region is supervised, but also the reduction process of the later complete text region is supervised. Therefore, it is possible to avoid problems such as inaccurate reduction results and inaccurate complete text region boundaries caused by the unsupervised reduction process of the complete text region, and improve the accuracy of text detection.
[0137] In addition, through the text detection method of the embodiments of the present application, after the text image to be detected is input into the text detection model, the text detection result can be obtained through the model without a complex post-processing process. Therefore, the efficiency of text detection is improved.
[0138] Embodiment Six
[0139] Refer to Figure 8 , Figure 8 FIG. is a structural block diagram of a model training device according to Embodiment Six of the present application. The model training device provided by the embodiments of the present application includes:
[0140] A sample acquisition module 802, configured to acquire a sample image containing text and a sample label, where the sample label includes: a complete text region label and a text center region label;
[0141] A prediction information obtaining module 804, configured to input the sample image into an initial text detection model, obtain a sample feature vector through the feature extraction module of the initial text detection model; through the prediction module of the initial text detection model, obtain a text center region probability map based on the sample feature vector, and perform binary processing on the text center region probability map to obtain a binary map; through the differentiable module of the initial text detection model, perform a region dilation operation on the binary map to obtain complete text region prediction information;
[0142] A training module 806, configured to obtain a loss value based on the text center region probability map, the text center region label, the complete text region prediction information, and the complete text region label; and train the initial text detection model based on the loss value to obtain a trained text detection model.
[0143] Optionally, in some embodiments, the training module 806 is specifically configured to:
[0144] Obtain a first loss value based on the text center region probability map and the text center region label;
[0145] Obtain a second loss value based on the complete text region prediction information and the complete text region label;
[0146] Fuse the first loss value and the second loss value to obtain a fused loss value;
[0147] Train the initial text detection model based on the fused loss value to obtain a trained text detection model.
[0148] Optionally, in some embodiments, when the prediction information obtaining module 804 executes the step of obtaining the text center region probability map based on the sample feature vector through the prediction module of the initial text detection model and performing binarization processing on the text center region probability map to obtain a binarized map, it is specifically configured to:
[0149] Through the prediction module of the initial text detection model, obtain a text center region probability map and a text center region threshold map based on the sample feature vector; based on the text center region threshold map, perform differentiable binarization processing on the text center region probability map to obtain a binarized map;
[0150] The sample label further includes: a threshold map label;
[0151] The training module 806 is specifically configured to:
[0152] Obtain a first loss value based on the text center region probability map and the text center region label;
[0153] Obtain a second loss value based on the complete text region prediction information and the complete text region label;
[0154] Obtain a third loss value based on the text center region threshold map and the threshold map label;
[0155] Fuse the first loss value, the second loss value, and the third loss value to obtain a fused loss value;
[0156] Train the initial text detection model based on the fused loss value to obtain a trained text detection model.
[0157] Optionally, in some embodiments, the differentiable module includes a max pooling layer.
[0158] The model training device according to an embodiment of the present application is used to implement the corresponding model training method in the first or second method embodiment described above, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated here. In addition, the function implementation of each module in the model training device according to an embodiment of the present application can refer to the description of the corresponding part in the first or second method embodiment described above, which will not be elaborated here either.
[0159] Embodiment VII
[0160] See Figure 9 , Figure 9 FIG. is a structural block diagram of a model training device according to Embodiment VII of the present application. The model training device provided by the embodiment of the present application includes:
[0161] A first sample acquisition module 902, configured to acquire an initial text detection model, a first sample image set, and a first sample label set, where the initial text detection model includes a feature extraction module, a prediction module, and a differentiable module;
[0162] A first prediction information obtaining module 904, configured to extract a sample feature vector of each first sample image in the first sample image set through the feature extraction module; obtain a text center region probability map based on each sample feature vector through the prediction module, and obtain a binary map based on the text center region probability map; perform a region dilation operation on each binary map through the differentiable module to obtain complete text region prediction information;
[0163] A transition model obtaining module 906, configured to obtain a loss value based on the text center region probability map, the complete text region prediction information, and the first sample label set; and train the initial text detection model based on the loss value to obtain a transition text detection model;
[0164] A text detection model obtaining module 908, configured to obtain a second sample image set and a second sample label set, and train the transition text detection model based on the second sample label set to obtain a trained text detection model; the number of complex background images in the first sample image set is greater than the number of complex background images in the second sample image set.
[0165] The model training device according to an embodiment of the present application is used to implement the corresponding model training method in the third method embodiment described above, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated here. In addition, the function implementation of each module in the model training device according to an embodiment of the present application can refer to the description of the corresponding part in the third method embodiment described above, which will not be elaborated here either.
[0166] Embodiment VIII
[0167] See Figure 10 , Figure 10It is a structural block diagram of a text detection device according to Embodiment 8 of the present application. The text detection device provided by the embodiment of the present application includes:
[0168] A target image acquisition module 1002, configured to acquire a target text image to be detected and a pre-trained text detection model. The text detection model includes: a feature extraction module, a prediction module, and a differentiable module; the differentiable module is a module that can perform gradient transmission during the training process of the text detection model;
[0169] A first result obtaining module 1004, configured to input the target text image into the text detection model, obtain a target feature vector through the feature extraction module; through the prediction module, obtain a text center region probability map based on the target feature vector, and perform binarization processing on the text center region probability map to obtain a binarized map; through the differentiable module, perform a region dilation operation on the binarized map to obtain a text detection result.
[0170] The text detection device of the embodiment of the present application is used to implement the corresponding text detection method in the fourth embodiment of the foregoing method, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated here. In addition, the function implementation of each module in the text detection device of the embodiment of the present application can refer to the description of the corresponding part in the fourth embodiment of the foregoing method, which will not be elaborated here either.
[0171] Embodiment 9
[0172] See Figure 11 , Figure 11 It is a structural block diagram of a text detection device according to Embodiment 9 of the present application. The text detection device provided by the embodiment of the present application includes:
[0173] An instruction receiving module 1102, configured to receive an instruction for indicating text detection, and scan a target area containing text according to the instruction to obtain a target text image;
[0174] A model acquisition module 1104, configured to acquire a pre-trained text detection model. The text detection model includes: a feature extraction module, a prediction module, and a differentiable module; the differentiable module is a module that can perform gradient transmission during the training process of the text detection model;
[0175] A second result obtaining module 1106, configured to input the target text image into the text detection model, obtain a target feature vector through the feature extraction module; through the prediction module, obtain a text center region probability map based on the target feature vector, and perform binarization processing on the text center region probability map to obtain a binarized map; through the differentiable module, perform a region dilation operation on the binarized map to obtain a text detection result.
[0176] The text detection device according to the embodiments of the present application is used to implement the corresponding text detection method in the fifth method embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated here. In addition, the function implementation of each module in the text detection device according to the embodiments of the present application can refer to the description of the corresponding part in the fifth method embodiment, which will not be elaborated here either.
[0177] Embodiment Ten
[0178] Referring to Figure 12 , a schematic structural diagram of a dictionary pen according to Embodiment Ten of the present application is shown. Among them, the dictionary pen includes: an image collector 1202, a memory 1204, and a processor 1206;
[0179] The image collector 1202 is configured to scan a target area containing text to obtain a target text image;
[0180] The memory 1204 is used to store executable instructions;
[0181] When the processor 1206 executes the executable instructions stored on the memory 1204, it inputs the target text image into a pre-trained text detection model, and obtains a target feature vector through a feature extraction module in the text detection model; through a prediction module in the text detection model, based on the target feature vector, it obtains a text center area probability map, and performs binarization processing on the text center area probability map to obtain a binarized map; through a differentiable module in the text detection model, it performs a region dilation operation on the binarized map to obtain a text detection result.
[0182] Embodiment Eleven
[0183] Referring to Figure 13 , a schematic structural diagram of an electronic device according to Embodiment Eleven of the present application is shown. The specific implementation of the electronic device is not limited in the specific embodiments of the present application.
[0184] As Figure 13 shown, the electronic device may include: a processor 1302, a communication interface 1304, a memory 1306, and a communication bus 1308.
[0185] Among them:
[0186] The processor 1302, the communication interface 1304, and the memory 1306 communicate with each other through the communication bus 1308.
[0187] The communication interface 1304 is used to communicate with other electronic devices or servers.
[0188] A processor 1302 for executing a program 1230, which can specifically execute the above model training method, or relevant steps in the text detection method embodiments.
[0189] Specifically, the program 1310 may include program code, which includes computer operation instructions.
[0190] The processor 1302 may be a CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application. One or more processors included in the intelligent device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.
[0191] A memory 1306 for storing the program 1310. The memory 1306 may include high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory.
[0192] The program 1310 can specifically be used to cause the processor 1302 to perform the following operations: obtain a sample image containing text and sample labels, where the sample labels include: a complete text region label and a text center region label; input the sample image into an initial text detection model, and obtain a sample feature vector through the feature extraction module of the initial text detection model; through the prediction module of the initial text detection model, obtain a text center region probability map based on the sample feature vector, and perform binarization processing on the text center region probability map to obtain a binarized map; through the derivable module of the initial text detection model, perform a region dilation operation on the binarized map to obtain complete text region prediction information; based on the text center region probability map, the text center region label, the complete text region prediction information, and the complete text region label, obtain a loss value; and based on the loss value, train the initial text detection model to obtain a trained text detection model.
[0193] Alternatively, program 1310 can specifically be used to cause processor 1302 to perform the following operations: obtain an initial text detection model, a first set of sample images, and a first set of sample labels, where the initial text detection model includes a feature extraction module, a prediction module, and a differentiable module; extract sample feature vectors of each first sample image in the first set of sample images through the feature extraction module; obtain a text center region probability map based on each sample feature vector through the prediction module, and obtain a binary map based on the text center region probability map; perform a region dilation operation on each binary map through the differentiable module to obtain complete text region prediction information; obtain a loss value based on the text center region probability map, the complete text region prediction information, and the first set of sample labels; and train the initial text detection model based on the loss value to obtain a transitional text detection model; obtain a second set of sample images and a second set of sample labels, and train the transitional text detection model based on the second set of sample labels to obtain a trained text detection model; the number of complex background images in the first set of sample images is greater than the number of complex background images in the second set of sample images.
[0194] Alternatively, program 1310 can specifically be used to cause processor 1302 to perform the following operations: receive an instruction for indicating text detection, and scan a target region containing text according to the instruction to obtain a target text image; obtain a pre-trained text detection model, where the text detection model includes: a feature extraction module, a prediction module, and a differentiable module; the differentiable module is a module capable of performing gradient transmission during the training process of the text detection model; input the target text image into the text detection model, and obtain a target feature vector through the feature extraction module; obtain a text center region probability map based on the target feature vector through the prediction module, and perform binary processing on the text center region probability map to obtain a binary map; perform a region dilation operation on the binary map through the differentiable module to obtain a text detection result.
[0195] Alternatively, program 1310 can specifically be used to cause processor 1302 to perform the following operations: a sample acquisition module, configured to acquire a sample image containing text and a sample label, where the sample label includes: a complete text region label and a text center region label; a prediction information obtaining module, configured to input the sample image into an initial text detection model, and obtain a sample feature vector through the feature extraction module of the initial text detection model; obtain a text center region probability map based on the sample feature vector through the prediction module of the initial text detection model, and perform binary processing on the text center region probability map to obtain a binary map; perform a region dilation operation on the binary map through the differentiable module of the initial text detection model to obtain complete text region prediction information; a training module, configured to obtain a loss value based on the text center region probability map, the text center region label, the complete text region prediction information, and the complete text region label; and train the initial text detection model based on the loss value to obtain a trained text detection model.
[0196] For the specific implementation of each step in Program 1310, reference can be made to the corresponding descriptions in the above-mentioned embodiments of the model training method or the corresponding steps and units in the embodiments of the text detection method, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described devices and modules can refer to the corresponding process descriptions in the foregoing method embodiments and will not be elaborated here.
[0197] Through the electronic device of this embodiment, a differentiable module is added to the text detection model. During model training, on the one hand, a probability map of the text center region is obtained through the prediction model and then a binary map (corresponding to the mask of the text center region) is obtained. The loss value is calculated based on the probability map of the center region and the label of the text center region to perform supervised learning on the prediction process of the text center region. On the other hand, the differentiable module is also used to dilate the binary map to predict the position information of the complete text region (complete text region prediction information), and the loss value is obtained again based on the predicted position information of the complete text region and the label of the complete text region to perform supervised learning on the restoration process of restoring the complete text region from the text center region. That is to say, in the text detection model training process of this application embodiment, not only the prediction process of the text center region in the early stage is supervised, but also the restoration process of the complete text region in the later stage is supervised. Therefore, it is possible to avoid problems such as inaccurate restoration results and inaccurate boundaries of the complete text region caused by the unsupervised restoration process of the complete text region, and improve the accuracy of text detection.
[0198] In addition, for text detection based on the text detection model of this application embodiment, after the text image to be detected is input into the text detection model, the text detection result can be obtained through this model without a complex post-processing process. Therefore, this application embodiment improves the efficiency of text detection.
[0199] This application embodiment also provides a computer program product, including computer instructions, which direct the computing device to perform the operations corresponding to any one of the above-mentioned model training methods or the operations corresponding to the text detection method.
[0200] It should be noted that according to the needs of implementation, each component / step described in this application embodiment can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of this application embodiment.
[0201] The method according to the embodiments of the present application can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code that is originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through a network and will be stored in a local recording medium, so that the method described herein can be stored on such a software process on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as a RAM, a ROM, a flash memory, etc.) that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, the processor, or the hardware, the model training method, or the text detection method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the model training method, or the text detection method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the model training method, or the text detection method shown herein.
[0202] Those of ordinary skill in the art can realize that the units and method steps of each example described in combination with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.
[0203] The above embodiments are only used to illustrate the embodiments of the present application, rather than to limit the embodiments of the present application. Those of ordinary skill in the relevant technical field can also make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present application. The patent protection scope of the embodiments of the present application shall be defined by the claims.
Claims
1. A model training method, comprising: Obtaining a sample image containing text and a sample label, where the sample label includes: a complete text region label and a text center region label; Inputting the sample image into an initial text detection model, obtaining a sample feature vector through the feature extraction module of the initial text detection model; through the prediction module of the initial text detection model, obtaining a text center region probability map based on the sample feature vector, and performing binarization processing on the text center region probability map to obtain a binarized map; through the differentiable module of the initial text detection model, performing a region dilation operation on the binarized map to obtain complete text region prediction information, where the value of a pixel point in the text center region probability map is used to represent the probability that the pixel point belongs to the text center region, and the differentiable module is used to perform an image dilation operation and is used for gradient transmission during the training process of the initial text detection model; Obtaining a loss value based on the text center region probability map, the text center region label, the complete text region prediction information, and the complete text region label; and training the initial text detection model based on the loss value to obtain a trained text detection model.
2. The method according to claim 1, wherein, The obtaining a loss value based on the text center region probability map, the text center region label, the complete text region prediction information, and the complete text region label; and training the initial text detection model based on the loss value to obtain a trained text detection model includes: Obtaining a first loss value based on the text center region probability map and the text center region label; Obtaining a second loss value based on the complete text region prediction information and the complete text region label; Fusing the first loss value and the second loss value to obtain a fused loss value; Training the initial text detection model based on the fused loss value to obtain a trained text detection model.
3. The method according to claim 1, wherein, The obtaining a text center region probability map based on the sample feature vector through the prediction module of the initial text detection model, and performing binarization processing on the text center region probability map to obtain a binarized map includes: Obtaining a text center region probability map and a text center region threshold map based on the sample feature vector through the prediction module of the initial text detection model; based on the text center region threshold map, performing differentiable binarization processing on the text center region probability map to obtain a binarized map; The sample label further includes: a threshold map label; the obtaining a loss value based on the text center region probability map, the text center region label, the complete text region prediction information, and the complete text region label; and training the initial text detection model based on the loss value to obtain a trained text detection model includes: Obtaining a first loss value based on the text center region probability map and the text center region label; Obtaining a second loss value based on the complete text region prediction information and the complete text region label; Obtaining a third loss value based on the text center region threshold map and the threshold map label; Fuse the first loss value, the second loss value, and the third loss value to obtain a fused loss value; Train the initial text detection model based on the fused loss value to obtain a trained text detection model.
4. The method according to claim 1, wherein, The differentiable module includes a max pooling layer.
5. A model training method, comprising: Obtain an initial text detection model, a first sample image set, and a first sample label set. The initial text detection model includes a feature extraction module, a prediction module, and a differentiable module; Through the feature extraction module, extract the sample feature vectors of each first sample image in the first sample image set; through the prediction module, obtain a text center region probability map based on each sample feature vector, and obtain a binary map based on the text center region probability map; through the differentiable module, perform a region dilation operation on each binary map to obtain complete text region prediction information. Wherein, the value of a pixel point in the text center region probability map is used to represent the probability that the pixel point belongs to the text center region, and the differentiable module is used to perform an image dilation operation and to perform gradient transmission during the training process of the initial text detection model; Based on the text center region probability map, the complete text region prediction information, and the first sample label set, obtain a loss value; and train the initial text detection model based on the loss value to obtain a transitional text detection model; Obtain a second sample image set and a second sample label set, and train the transitional text detection model based on the second sample label set to obtain a trained text detection model; the number of complex background images in the first sample image set is greater than the number of complex background images in the second sample image set.
6. A text detection method, comprising: Obtain a target text image to be detected and a pre-trained text detection model. The text detection model includes: a feature extraction module, a prediction module, and a differentiable module; the differentiable module is used to perform an image dilation operation and to perform gradient transmission during the training process of the text detection model; Input the target text image into the text detection model, obtain a target feature vector through the feature extraction module; through the prediction module, obtain a text center region probability map based on the target feature vector, and perform binary processing on the text center region probability map to obtain a binary map; through the differentiable module, perform a region dilation operation on the binary map to obtain a text detection result. Wherein, the value of a pixel point in the text center region probability map is used to represent the probability that the pixel point belongs to the text center region.
7. A text detection method, applied to a dictionary pen, comprising: Receive an instruction for indicating text detection, and scan a target region containing text according to the instruction to obtain a target text image; Obtain a pre-trained text detection model. The text detection model includes: a feature extraction module, a prediction module, and a differentiable module; the differentiable module is used to perform an image dilation operation and to perform gradient transmission during the training process of the text detection model; Input the target text image into the text detection model, and obtain a target feature vector through the feature extraction module; through the prediction module, obtain a text center region probability map based on the target feature vector, and perform binarization processing on the text center region probability map to obtain a binarized map; through the differentiable module, perform a region dilation operation on the binarized map to obtain a text detection result, where the value of a pixel point in the text center region probability map is used to represent the probability that the pixel point belongs to the text center region.
8. A dictionary pen, comprising: An image collector, a memory, and a processor; The image collector is configured to scan a target region containing text to obtain a target text image; The memory is configured to store executable instructions; When the processor executes the executable instructions stored on the memory, it inputs the target text image into a pre-trained text detection model, and obtains a target feature vector through the feature extraction module in the text detection model; through the prediction module in the text detection model, obtains a text center region probability map based on the target feature vector, and performs binarization processing on the text center region probability map to obtain a binarized map; through the differentiable module in the text detection model, performs a region dilation operation on the binarized map to obtain a text detection result, where the value of a pixel point in the text center region probability map is used to represent the probability that the pixel point belongs to the text center region, and the differentiable module is used to perform an image dilation operation and is used for gradient transmission during the training process of the text detection model.
9. A computer storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the model training method described in any one of claims 1-5, or implements the text detection method described in any one of claims 6-7.
10. A computer program product, including computer instructions, and the computer instructions direct a computing device to perform the operations corresponding to the model training method described in any one of claims 1-5, or perform the operations corresponding to the text detection method described in any one of claims 6-7.
Citation Information
Patent Citations
Text detection method and device, electronic equipment and computer storage medium
CN111652217A
Method and apparatus for training a convolutional neural network to detect defects
WO2020048119A1