A method and apparatus for training a model and character detection

Through the training stage of synthetic training samples generation and labeling model, filtering and manually labeling the real training sample labels, combining the center line as the label training model, the problem of inaccurate enclosure boxes and high labeling costs in character detection is solved, and more efficient and accurate character detection is achieved.

CN113205095BActive Publication Date: 2025-07-22BEIJING SANKUAI ONLINE TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110392490.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-13
Publication Date
2025-07-22
Estimated Expiration
2041-04-13

AI Technical Summary

Technical Problem

In the existing character detection technology, the accuracy of the enclosure box obtained by text detection is not ideal, resulting in inaccurate final text recognition results and high cost of training sample annotation.

Method used

Through the training stage of the synthetic training sample generation and labeling model, the labeling sample is used to determine the label without manual annotation, and the labeling model is trained to label the real training sample; in the real sample processing stage, the enclosing box output of the annotation model and the manual annotation center line are selected as the real training sample label; in the training stage of the text detection model, the center line is added as the real training sample label to improve the accuracy of the model; in the text detection stage, the expansion enclosing box is determined as the character detection result based on the overlap of the enclosing box and the center line.

Benefits of technology

It reduces the annotation cost of training samples, improves the accuracy of character detection, reduces the impact of synthetic training samples errors, and outputs more accurate character detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113205095B_ABST
    Figure CN113205095B_ABST
Patent Text Reader

Abstract

This specification discloses a method and apparatus for training a model and character detection. The annotation model is trained based on synthetic training samples, the real training samples are annotated according to the output of the trained annotation model, and the character detection model is trained based on the synthetic training samples. The image to be detected is subjected to feature extraction by the trained character detection model, and the bounding boxes of each character in the image and each center line in the image are determined. And according to the overlapping degree between each center line and each bounding box, and the bounding boxes overlapping with the same center line, a bounding box group is determined, and according to the geometric position features of each bounding box in each bounding box group, each center line is dilated around to obtain each dilated bounding box, which is used as the character detection result of the image. The accurate bounding box and center line can be output by the trained character detection model to determine the accurate dilated bounding box as the character detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular, to a method and device for training a model and character detection. Background Art

[0002] Optical Character Recognition (OCR) technology is a technology that can convert text in an image into a text format. With the development of OCR technology, OCR-based text recognition technology has been widely applied. The text recognition technology performs text detection on an image to determine the bounding box of a character string (e.g., all letters of a word form a character string) from the image, so as to locate the character string in the image. After obtaining the bounding boxes of each character string through text detection, the text recognition technology can recognize the text in the bounding box based on the obtained bounding boxes of each character string to obtain the text in the image.

[0003] Currently, the accuracy of the bounding boxes of each character string obtained through text detection has a great impact on the accuracy of the final text recognition result. However, in the existing text detection technology, the accuracy of the bounding boxes of each character string obtained through text detection is not ideal. Summary of the Invention

[0004] This specification provides a method and device for training a model and character detection to partially solve the above problems existing in the prior art.

[0005] This specification adopts the following technical solutions:

[0006] This specification provides a method for training a character detection model, including:

[0007] Obtain a plurality of images from an image dataset as training samples, and for each training sample, determine the bounding box of each character in the image corresponding to the training sample as the first label of the training sample, and determine the center line of each character string in the image corresponding to the training sample as the second label of the training sample;

[0008] Input the training sample into the feature extraction network of the character detection model to be trained, and determine a plurality of feature maps corresponding to the training sample;

[0009] Use the plurality of feature maps corresponding to the training sample as inputs, input them into the geometric feature detection network of the character detection model to be trained to obtain each predicted bounding box, and input them into the line feature detection network of the character detection model to be trained to obtain each predicted center line;

[0010] Determine a first loss based on the differences between the obtained predicted bounding boxes and the first label of the training sample, and determine a second loss based on the differences between the obtained predicted centerlines and the second label of the training sample;

[0011] Based on the first loss and the second loss, determine the total loss of the character detection model, and with the goal of minimizing the total loss, adjust the parameters of the character detection model to be trained. The character detection model is used to determine the bounding boxes and centerlines of each character in the image to be detected, and expand the centerlines around according to each bounding box to obtain each expanded bounding box as the character detection result of the image to be detected.

[0012] Optionally, the first label of the training sample further includes the types of characters in each bounding box in the image corresponding to the training sample;

[0013] Taking a plurality of feature maps corresponding to the training sample as inputs, input them into the geometric feature detection network of the character detection model to be trained to obtain each predicted bounding box, specifically including:

[0014] Taking a plurality of feature maps corresponding to the training sample as inputs, input them into the geometric feature detection network of the character detection model to be trained to obtain each predicted bounding box and the confidence of the prediction result of the image within each predicted bounding box on each predicted type dimension.

[0015] Optionally, determining the first loss based on the differences between the obtained predicted bounding boxes and the first label of the training sample specifically includes:

[0016] Determine the geometric position features of each obtained predicted bounding box and the confidence of the prediction result of the image within each predicted bounding box on each predicted type dimension, and determine the geometric position features of each bounding box in the first label of the training sample and the feature values of the types to which the characters within each bounding box belong;

[0017] For each predicted bounding box, determine the regression loss of the predicted bounding box according to the difference between the geometric position features of the predicted bounding box and the geometric position features of the bounding box corresponding to the predicted bounding box in the first label of the training sample;

[0018] According to the feature value of the type to which the bounding box corresponding to the predicted bounding box in the first label of the training sample belongs, and the confidence of the prediction result of the image within the predicted bounding box on each predicted type dimension, determine the classification loss of the predicted bounding box;

[0019] Determine the first loss according to the regression loss of each predicted bounding box and the classification loss of each predicted bounding box.

[0020] Optionally, the geometric feature detection network includes a region detection network and a region correction network;

[0021] Taking a plurality of feature maps corresponding to the training sample as inputs, and inputting them into the geometric feature detection network of the character detection model to be trained to obtain each predicted bounding box, specifically including:

[0022] Taking a plurality of feature maps corresponding to the training sample as inputs, and inputting them into the region detection network to determine each initial predicted bounding box corresponding to each character in the output;

[0023] For each character, according to each initial predicted bounding box corresponding to the character, performing feature sampling on the region enclosed by each initial predicted bounding box to determine a plurality of feature matrices corresponding to the character;

[0024] According to the plurality of feature matrices corresponding to the character obtained, through the region correction network, determining the position offset feature of each initial predicted bounding box, and correcting each initial predicted bounding box according to the position offset feature of each initial predicted bounding box to determine the predicted bounding box of the character in the training sample.

[0025] Optionally, determining a first loss according to the difference between each predicted bounding box obtained and the first label of the training sample, specifically including:

[0026] For each initial predicted bounding box output by the region detection network, according to the geometric position feature of the initial predicted bounding box, determining the bounding box corresponding to the initial predicted bounding box in the first label;

[0027] According to the initial predicted bounding box and the bounding box in the first label corresponding to it, determining the first regression loss of the initial predicted bounding box;

[0028] For each predicted bounding box output by the region correction network, according to the geometric position feature of the predicted bounding box, determining the bounding box corresponding to the predicted bounding box in the first label;

[0029] According to the predicted bounding box and the bounding box in the first label corresponding to it, determining the second regression loss of the predicted bounding box;

[0030] Determining the first loss according to each first regression loss and each second regression loss.

[0031] Optionally, the first label of the training sample further includes the type of the character in each bounding box in the image corresponding to the training sample, and the region detection network and the region correction network also respectively output the confidence of the prediction result of the image in the initial predicted bounding box in each predicted type dimension, and the confidence of the prediction result of the image in the predicted bounding box in each predicted type dimension;

[0032] Optionally, the method further includes:

[0033] For each predicted bounding box output by the region rectification network, determine the bounding box in the first label corresponding to the predicted bounding box according to the geometric position characteristics of the predicted bounding box;

[0034] According to the initial predicted bounding box and the corresponding bounding box in the first label, determine the first regression loss of the initial predicted bounding box;

[0035] According to the confidence of the prediction result of the image within the initial predicted bounding box in each prediction type dimension and the eigenvalue of the corresponding type of the bounding box in the first label, determine the first classification loss of the predicted bounding box;

[0036] Determine the initial loss according to each first regression loss and each first classification loss;

[0037] For each predicted bounding box output by the region rectification network, determine the bounding box in the first label corresponding to the predicted bounding box according to the geometric position characteristics of the predicted bounding box;

[0038] According to the predicted bounding box and the corresponding bounding box in the first label, determine the second regression loss of the predicted bounding box;

[0039] For each predicted bounding box output by the region rectification network, determine the second classification loss of the predicted bounding box according to the confidence of the prediction result of the image within the predicted bounding box in each prediction type dimension and the eigenvalue of the corresponding type of the bounding box in the first label;

[0040] Determine the rectification loss according to each second regression loss and each second classification loss;

[0041] Determine the first loss according to the initial loss and the rectification loss.

[0042] Optionally, determine the second loss according to the difference between each obtained predicted center line and the second label of the training sample, specifically including:

[0043] According to each obtained predicted center line, determine the image containing each predicted center line as the center line map of the training sample;

[0044] Determine the type eigenvalue of each pixel point in the center line map of the training sample;

[0045] For each pixel point, determine the loss corresponding to the pixel point according to the type eigenvalue of the pixel point and the type eigenvalue of the corresponding pixel point in the second label of the training sample;

[0046] Determine the second loss of the training sample according to the losses corresponding to each pixel point.

[0047] Optionally, obtain a number of images from the image dataset as training samples, and for each training sample, determine the bounding boxes of each character in the image corresponding to the training sample as the first label of the training sample, specifically including:

[0048] Obtain a number of images from the image dataset as training samples, and for each training sample, input the image corresponding to the training sample into the trained annotation model, and determine the bounding boxes output by the annotation model, the confidence levels of the prediction results of the images within each bounding box in each preset type dimension, and the center lines of each string in the image corresponding to the training sample;

[0049] According to the confidence levels of the prediction results of the images within each bounding box in each preset type dimension, determine the type corresponding to each bounding box, so as to determine each initial annotated bounding box from each bounding box according to the type corresponding to each bounding box;

[0050] According to each initial annotated bounding box and the center lines of each string, determine each annotated bounding box from each initial annotated bounding box;

[0051] Take each annotated bounding box and the type corresponding to each annotated bounding box as the first label of the training sample.

[0052] Optionally, adopt the following method to determine the training samples for training the annotation model:

[0053] Obtain a number of background images and a number of element images from the image material library, and the element images at least include images corresponding to each character type and images corresponding to each string;

[0054] According to the obtained background images and the element images, synthesize a number of synthetic images as synthetic training samples;

[0055] For each synthetic image, according to the sizes and positions of the element images in the synthetic image, determine the bounding boxes of each character in the synthetic image and the type of the character within each bounding box as the first label of the synthetic training sample corresponding to the synthetic image, and determine the center lines of each string in the synthetic image as the second label of the synthetic training sample;

[0056] Train the annotation model to be trained according to the synthetic training samples to obtain the trained annotation model, and the annotation model is used to annotate the training samples determined from the image dataset.

[0057] Optionally, take the trained annotation model as the character detection model to be trained, and the annotation model is trained in the following manner:

[0058] For each synthetic training sample, use the first label and the second label of the synthetic training sample as the label of the synthetic training sample;

[0059] Input the synthetic training sample into the feature extraction network of the annotation model to determine several feature maps corresponding to the synthetic training sample;

[0060] Use the several feature maps corresponding to the synthetic training sample as inputs, input them into the geometric feature detection network of the annotation model to obtain each predicted bounding box, the confidence of the prediction result of the image within each predicted bounding box in each predicted type dimension, and input them into the line feature detection network of the annotation model to obtain each predicted center line;

[0061] Determine the first loss according to each obtained predicted bounding box and the difference between the confidence of the prediction result of the image within each predicted bounding box in each predicted type dimension and the first label of the synthetic training sample, and determine the second loss according to the difference between each obtained predicted center line and the second label of the training sample;

[0062] Determine the total loss of the annotation model according to the first loss and the second loss, and use minimizing the total loss as the training objective to adjust the parameters of the annotation model.

[0063] This specification provides a method for character detection, including:

[0064] Obtain the image to be detected, input the image into the feature extraction network in a pre-trained character detection model to determine several feature maps corresponding to the image;

[0065] Use the several feature maps corresponding to the image as inputs, and input them into the geometric feature detection network and the line feature detection network in the character detection model respectively. Through the geometric feature detection network, determine the bounding boxes of each character in the image, and through the line feature detection network, determine each center line in the image;

[0066] For each bounding box, determine the center line corresponding to the bounding box according to the overlapping degree between each center line and the bounding box;

[0067] Determine that the bounding boxes corresponding to the same center line are a group of bounding boxes;

[0068] For each group of bounding boxes, determine the dilation distance according to the geometric position features of the bounding boxes in the group of bounding boxes. According to the dilation distance, dilate the center line corresponding to the group of bounding boxes around to determine the dilated bounding box of the group of bounding boxes as the character detection result of the image.

[0069] Optionally, the geometric feature detection network includes a region detection network and a region correction network;

[0070] Through the geometric feature detection network, determine the bounding boxes of each character in the image, specifically including:

[0071] Through the region detection network, determine the initial bounding boxes of each character in the image;

[0072] According to each initial bounding box, perform feature sampling on the region enclosed by each initial bounding box to determine a number of feature matrices;

[0073] According to the obtained number of feature matrices, through the region correction network, determine the position offset features of each initial bounding box, and according to the position offset features of each initial bounding box, correct each initial bounding box to determine the bounding boxes of each character in the image.

[0074] Optionally, through the line feature detection network, determine each center line in the image, specifically including:

[0075] Through the line feature detection network, perform upsampling on a number of feature maps corresponding to the image to determine a number of feature maps of a specified scale;

[0076] Fuse a number of feature maps of a specified scale, reduce the number of channels of the fused feature map, and perform upsampling on the fused feature map to obtain a probability map consistent with the original scale of the image;

[0077] Perform binarization processing on the probability map to determine each center line corresponding to the image and the center line map corresponding to the image.

[0078] Optionally, the geometric position features of each bounding box at least include side length features;

[0079] According to the geometric position features of each bounding box in the bounding box group, determine the dilation distance, specifically including:

[0080] According to the side length features of each bounding box in the bounding box group, determine the dilation value;

[0081] According to the dilation value and a preset dilation coefficient, determine the dilation distance.

[0082] This specification provides an apparatus for training a character detection model, including:

[0083] A sample label determination module, configured to obtain a number of images from an image dataset as training samples, and for each training sample, determine the bounding boxes of each character in the image corresponding to the training sample as the first label of the training sample, and determine the center lines of each character string in the image corresponding to the training sample as the second label of the training sample;

[0084] A feature extraction module, configured to input the training sample into a feature extraction network of a character detection model to be trained, and determine a plurality of feature maps corresponding to the training sample;

[0085] A prediction module, configured to use the plurality of feature maps corresponding to the training sample as inputs, input them into a geometric feature detection network of the character detection model to be trained to obtain respective predicted bounding boxes, and input them into a line feature detection network of the character detection model to be trained to obtain respective predicted center lines;

[0086] A loss determination module, configured to determine a first loss according to the difference between each obtained predicted bounding box and a first label of the training sample, and determine a second loss according to the difference between each obtained predicted center line and a second label of the training sample;

[0087] A parameter adjustment module, configured to determine a total loss of the character detection model according to the first loss and the second loss, and use minimizing the total loss as a training objective to adjust parameters of the character detection model to be trained. The character detection model is used to determine bounding boxes and center lines of each character in an image to be detected, and to expand each center line around according to each bounding box to obtain respective expanded bounding boxes as character detection results of the image to be detected.

[0088] This specification provides a character detection device, including:

[0089] A feature extraction module, configured to obtain an image to be detected, input the image into a feature extraction network in a pre-trained character detection model, and determine a plurality of feature maps corresponding to the image;

[0090] A feature output module, configured to use the plurality of feature maps corresponding to the image as inputs, and input them into a geometric feature detection network and a line feature detection network in the character detection model respectively. Through the geometric feature detection network, determine bounding boxes of each character in the image, and through the line feature detection network, determine center lines in the image;

[0091] A correspondence determination module, configured to, for each bounding box, determine a center line corresponding to the bounding box according to the overlapping degree between each center line and the bounding box;

[0092] A bounding box group determination module, configured to determine each bounding box corresponding to the same center line as a bounding box group;

[0093] A detection result determination module, configured to, for each bounding box group, determine an expansion distance according to geometric position features of each bounding box in the bounding box group, and expand the center line corresponding to the bounding box group around according to the expansion distance to determine an expanded bounding box of the bounding box group as a character detection result of the image.

[0094] This specification provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the above-mentioned method for training a model and character detection.

[0095] This specification provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the above-mentioned method for training a model and character detection.

[0096] At least one of the above technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects:

[0097] In the embodiments of this specification, a labeling model is trained based on synthetic training samples, real training samples are labeled according to the output of the trained labeling model, and a character detection model is trained according to the synthetic training samples. The character detection model after training extracts features from the image to be detected, and determines the bounding boxes of each character in the image and each center line in the image. And according to the overlap degree between each center line and each bounding box, and the bounding boxes overlapping with the same center line, a bounding box group is determined, and according to the geometric position features of each bounding box in each bounding box group, each center line is expanded around to obtain each expanded bounding box, which is used as the character detection result of the image. The accurate bounding box and center line can be output by the trained character detection model to determine the accurate expanded bounding box as the character detection result. BRIEF DESCRIPTION OF THE DRAWINGS

[0098] The drawings described herein are used to provide a further understanding of this specification, and constitute a part of this specification. The illustrative embodiments of this specification and their descriptions are used to explain this specification and do not constitute an improper limitation to this specification. In the drawings:

[0099] Figure 1 is a schematic flowchart of the synthetic training sample generation and labeling model training stage provided by the embodiments of this specification;

[0100] Figure 2 is a schematic diagram of a synthetic image provided by this specification;

[0101] Figure 3 is a schematic structural diagram of a labeling model provided by this specification;

[0102] Figure 4 is a schematic flowchart of the real sample processing stage provided by the embodiments of this specification;

[0103] Figure 5 is a schematic diagram of screening bounding boxes provided by this specification;

[0104] Figure 6 It is a schematic flowchart of the training stage of the text detection model provided by the embodiments of this specification;

[0105] Figure 7 It is a schematic structural diagram of an annotation model provided by this specification;

[0106] Figure 8 It is a schematic flowchart of the text detection stage provided by the embodiments of this specification;

[0107] Figure 9 It is a schematic diagram of determining an expanded bounding box provided by this specification;

[0108] Figure 10 It is a schematic diagram of a device for training a character detection model provided by this specification;

[0109] Figure 11 It is a schematic diagram of a character detection device provided by this specification;

[0110] Figure 12 It is a schematic structural diagram of an electronic device provided by this specification. Specific embodiments

[0111] In some current methods for detecting characters in images, when training a model for character detection, the cost of annotating training samples is relatively high, resulting in a high cost for training the model. Training the model through weak supervision learning methods can eliminate the need for a large amount of sample annotation, thereby reducing the annotation cost of training samples. However, the detection results of characters in images by the model trained through weak supervision learning methods are not accurate enough.

[0112] To solve the above problems existing in the current character detection methods, the embodiments of this specification provide a method for training a model and character detection. This method includes four stages: the stage of generating and training an annotation model for synthetic training samples, the stage of processing real samples, the stage of training a text detection model, and the text detection stage.

[0113] In the stage of generating synthetic training samples and training the annotation model, to solve the problem of high cost of annotating training samples when training a model for character detection, first, a number of background images and a number of element images can be obtained to synthesize training samples, and the labels of the synthetic training samples can be determined according to the attributes of the element images added to the background images (such as position, size, orientation, etc.). Therefore, accurate sample labels can be determined without manual annotation, avoiding the problem of high manual annotation cost. Then, the annotation model used to annotate real training samples can be trained according to the synthetic training samples, so that in the subsequent stage, the labels of the real training samples used to train the character detection model can be generated by the trained annotation model. Since there are differences between synthetic images and real images, the accuracy of the character detection results output by the annotation model trained according to the synthetic training samples based on the input real images is difficult to guarantee. Therefore, this annotation model is not the final model for character detection, but a model used to assist in annotating unannotated real images in the subsequent stage to reduce the manual annotation cost.

[0114] In the stage of processing real samples, after obtaining the trained annotation model, the unannotated real images are used as real training samples, and the real training samples are annotated through this annotation model. Specifically, the bounding boxes of each character output by the trained annotation model and the center lines of each image corresponding to the training sample can be obtained. By screening the content output by the annotation model, the accurate bounding boxes that can be used as the labels of real training samples can be determined, and the center lines of each string in the image corresponding to the real training sample are annotated through manual annotation. In this stage, the selected bounding boxes and the manually annotated center lines are used as the labels of the real training samples, which can be used to train a more accurate character detection model in the subsequent stage. Moreover, compared with the existing method of annotating bounding boxes for all training samples, the amount of manual annotation and annotation time of training samples in this stage are reduced, greatly reducing the annotation cost.

[0115] In the training stage of the text detection model, to solve the problem that the detection results determined based on the output of the model are inaccurate in current character detection methods, the text detection model can be trained according to the real training samples and the accurate labels of the real training samples obtained in the previous stage. Since the training samples used to train the annotation model are synthetic non-real images and there are certain errors in the output of the annotation model, and the bounding boxes used as labels in the real training samples are determined after screening the output of the annotation model. Therefore, in order to reduce errors and make the output of the character detection model more accurate, in this stage, in addition to using the bounding boxes of each character as the labels of the real training samples, the center lines of each string are also used as the labels of the real training samples to train the character detection model, which can make the output of the trained character detection model more accurate. By adding the center line as the label of the real training sample, the influence brought by the error of the bounding box can be reduced or even eliminated to a great extent.

[0116] In the text detection stage, taking the image to be detected as the input and inputting it into the trained text detection model, the bounding boxes of each character in the image to be detected and the center line corresponding to the image to be detected can be determined. According to the overlapping degree of the obtained bounding boxes and center lines, the center line corresponding to each bounding box is determined to form a bounding box group, and the center lines corresponding to each bounding box group are dilated around, and the dilated bounding boxes that enclose each string can be accurately determined as the final character detection results.

[0117] To make the purpose, technical solutions and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.

[0118] The following will detail the technical solutions provided by each embodiment of this specification in conjunction with the drawings.

[0119] Figure 1 The figure is a schematic flowchart of the synthetic training sample generation and annotation model training stage provided by the embodiment of this specification, which specifically includes the following steps:

[0120] S100: Obtain a number of background images and a number of element images from the image material library, and the element images at least include images corresponding to each character type and images corresponding to each string.

[0121] In one or more embodiments of this specification, the training model and the method for character detection can be executed by a server, which can specifically be a single device or a system composed of multiple devices (such as a distributed system). This specification does not make any limitations and can be set according to requirements.

[0122] In one or more embodiments of this specification, as mentioned above, to solve the problem of high cost in annotating training samples currently, during the stage of generating synthetic training samples and training the annotation model, synthetic training samples can be determined through synthetic images. First, the server can obtain a number of background images and a number of element images from the image material library to prepare synthetic materials for subsequent synthetic images. In this specification, synthetic images can be divided into a background image part and a foreground element part. The element images obtained from the image material library at least include images corresponding to each character type and images corresponding to each string. Each character type can be set according to requirements. For example, Arabic numerals can be regarded as one type, Chinese characters as one type, foreign languages as one or more types, and punctuation marks as one type. Or, the types of characters can be classified more meticulously. For example, the ten Arabic numerals from 0 to 9 can be divided into ten character types, and English can be divided into 26 character types according to the 26 letters, etc.

[0123] Among them, the images corresponding to each character type can also be other types of images, such as element images existing in life scenes like brand logos, vehicles, or plants. The images corresponding to each string can be images with a single string as a unit. For example, the image corresponding to a single string can be an image corresponding to an English word, an image corresponding to a string of numbers, or an image corresponding to a line of text. The images corresponding to each character type can be images with a single character as a unit. For example, it can be an image corresponding to a Chinese character, an image corresponding to a punctuation mark, or an image corresponding to a brand logo, etc.

[0124] S102: Synthesize a number of synthetic images based on the obtained background images and element images as synthetic training samples.

[0125] In one or more embodiments of this specification, after the server obtains a number of background images and a number of element images from the image material library, a number of synthesized images can be synthesized based on each background image and each element image as synthesized training samples. Specifically, the server can match the obtained number of element images with the background images. For each background image, the server can randomly select a certain number of element images from the obtained number of element images and place them on the background image for synthesis to obtain a synthesized image as a synthesized training sample. The server can place the selected element images at any position on the background image and record the positions of the placed element images and the sizes of each element image (such as the length and width of the element image). The server can use the position of the central pixel point of each element image as the position of the element image. Of course, other pixel point positions can also be used as the position of the element image, which can be specifically set according to needs, and this specification does not limit it here.

[0126] Among them, the number of element images in each synthesized image can be the same or different, and each synthesized image includes at least one image corresponding to a string, which can be specifically set according to needs. The number of element images obtained by the server from the image material library can be transparent background images, such as images in Portable Network Graphics (PNG) format, or transparent background images in other formats, and this specification does not limit it here.

[0127] Since in real life, the strings in real images are not all regular and straight, and due to different shooting angles of the images, there will be some perspective situations, and there are also some strings that are designed with deformations, such as the sign names of some merchants, etc. Therefore, in order to make the synthesized images closer to the images taken in real situations, in one or more embodiments of this specification, before matching the obtained number of element images with the background images, some element images can also be preprocessed. For example, some element images can be bent, stretched, blurred, angle-transformed, etc., to imitate some situations of character deformation in real images in real life due to character typesetting, improper shooting, or other reasons. By synthesizing each synthesized image with these preprocessed element images as training samples, the output results of the annotation model trained based on each synthesized training sample in subsequent steps can be more accurate.

[0128] Among them, for each synthesized image, the synthesized image contains at least one preprocessed elemental image. For each image corresponding to a string in the synthesized image, one or more preprocessing operations can be performed on it, or of course, no preprocessing can be performed. That is, in each finally obtained synthesized image, there may be images corresponding to strings that have not been preprocessed and images corresponding to strings that have been (one or more) preprocessed, or each synthesized image may also be entirely composed of images corresponding to strings that have been (one or more) preprocessed.

[0129] S104: For each synthesized image, according to the sizes and positions of the elemental images in the synthesized image, determine the bounding boxes of the characters in the synthesized image and the types of the characters within each bounding box as the first label of the synthesized training sample corresponding to the synthesized image, and determine the center lines of the strings in the synthesized image as the second label of the synthesized training sample.

[0130] Since when performing character detection, the ultimate goal of character detection is not to detect individual characters, nor to determine the bounding boxes of each character, but to determine the overall bounding box of the string composed of multiple characters from the image to be detected. However, if only the bounding boxes of each character are used as labels, if there is a missed detection (i.e., a character is not detected at a position where there is actually a character and no bounding box annotation is made) or a false detection (e.g., a bounding box is annotated for other elements as if they were characters in the string) for one or several characters in a string, it is easy to cause errors in the bounding boxes of the strings obtained based on each bounding box. For example, for a string composed of five characters, assuming that the first and last characters of the string are missed detected, then the overall bounding box of the string finally obtained based on the bounding boxes of the string can only enclose the middle three characters.

[0131] Therefore, in one or more embodiments of this specification, in addition to determining the bounding boxes of each character as the label of the synthesized training sample, the center lines of the strings in the synthesized training sample can also be annotated, so that in subsequent steps, the annotation model is trained based on the synthesized training sample, making the text detection result output by the trained annotation model more accurate.

[0132] Specifically, after synthesizing the training sample, the server can, for each synthesized image, determine the images corresponding to the strings according to the sizes and positions of the elemental images in the synthesized image, and for each image corresponding to a string, perform segmentation on it according to the number of characters it contains and the sizes of each character to determine the bounding boxes of the characters in the image corresponding to the string. The server can determine the bounding boxes of the characters in the synthesized image and the types of the characters within each bounding box as the first label of the synthesized training sample corresponding to the synthesized image, and determine the center lines of the strings in the synthesized image as the second label of the synthesized training sample corresponding to the synthesized image.

[0133] Among them, the size of the element image can be represented by the length and width of the element image. The center line of the string is the line passing through the centers of the characters in the string along the arrangement direction of the characters in the string.

[0134] In one or more embodiments of this specification, after determining the center lines of the strings in the composite image, an image containing each predicted center line can be determined according to each center line as the center line map corresponding to the composite image, and the center line map can be used as the second label of the composite training sample corresponding to the composite image. The center line map is an image with the same resolution as the composite image, and the positions of the center lines in the center line map are the same as the positions of the lines passing through the centers of the strings. In the center line map, the values of the pixel points on each center line are different from the values of the pixel points outside the center line. For example, the values of the pixel points on the center line can be assigned as 1, and the values of the pixel points outside the center line can be assigned as 0. Of course, other values can also be used, which can be specifically set according to needs, and this specification does not limit it here.

[0135] Figure 2 It is a schematic diagram of generating a composite image provided in this specification. As shown in the figure, four large rectangles represent four different stages of generating the composite image. In the first stage, the characters enclosed by the same dotted box represent the characters of the image corresponding to a string, including: "FLORA", "ANT", and three strings. In the second stage, by preprocessing the images corresponding to some strings, the strings "FLORA" and "ANT" are deformed, and the string "All things recover" is not preprocessed. In the third stage, the images corresponding to the strings are placed in the background image, and the characters in each string are segmented. According to the positions and sizes of the characters, the bounding boxes corresponding to the characters are determined, and the center point positions of the bounding boxes are determined. Among them, the smaller solid boxes enclosing the characters are the bounding boxes of the characters. In the fourth stage, for the horizontal strings "ANT" and "All things recover", the center lines of the strings can be determined according to the center point positions of the first and last characters of each string. For the curved string "FLORA", the center lines of the strings can be determined according to the center point positions of the first and last characters and the inflection point characters of each string. Among them, the center point of the letter O is the inflection point of the string. Of course, for each string, the center line of the string can also be determined according to the center point positions of each character in the string. It can be specifically set according to needs, and this specification does not limit it here.

[0136] In one or more embodiments of this specification, after the server segments each image corresponding to a string based on the number of characters it contains and the sizes of the characters, and determines the bounding boxes of the characters in the image corresponding to the string. It can also only determine the bounding boxes of the characters in the composite image as the first label of the composite training sample corresponding to the composite image. That is, for each composite image, the server can also determine the bounding boxes of the characters in the composite image as the first label of the composite training sample corresponding to the composite image according to the sizes and positions of the element images in the composite image, and determine the centerlines of the strings in the composite image as the second label of the composite training sample.

[0137] S106: Train the annotation model to be trained according to the composite training sample to obtain the trained annotation model, where the annotation model is used to annotate the training samples determined from the image dataset.

[0138] In one or more embodiments of this specification, after the server determines the composite training sample and the label of the composite training sample, it can train the annotation model to be trained according to the composite training sample to obtain the trained annotation model, where the annotation model is used to annotate the training samples (i.e., non-composite real training samples) determined from the image dataset.

[0139] The structure of the annotation model is as Figure 3 shown Figure 3 It is a schematic structural diagram of an annotation model provided in this specification. As shown in the figure, the annotation model includes: a feature extraction network, a geometric feature detection network, and a line feature detection network. The feature extraction network is used to extract features from the image to be detected to obtain a number of feature maps, the geometric feature detection network is used to output each bounding box, and the line feature detection network is used to output each centerline.

[0140] Specifically, when training the annotation model, the server can, for each composite training sample, use the first label and the second label of the composite training sample as the label of the composite training sample, input the composite training sample into the feature extraction network of the annotation model, and determine the number of feature maps corresponding to the composite training sample. Then, the server can use the number of feature maps corresponding to the composite training sample as the input, input it into the geometric feature detection network of the annotation model to obtain each predicted bounding box, and input it into the line feature detection network of the annotation model to obtain each predicted centerline of the composite image corresponding to the composite training sample.

[0141] In one or more embodiments of this specification, the server may determine a first loss according to the differences between the obtained predicted bounding boxes and the first labels of the synthetic training samples, and determine a second loss according to the differences between the obtained predicted centerlines and the second labels of the training samples. Based on the first loss and the second loss, the total loss of the annotation model is determined, and the parameters of the annotation model are adjusted with the goal of minimizing the total loss.

[0142] Specifically, when determining the first loss, the server may determine the geometric position features of each bounding box in the first labels of the training samples. For each predicted bounding box, according to the differences between the geometric position features of the predicted bounding box and the geometric position features of the corresponding bounding box in the first labels of the training samples, the regression loss of the predicted bounding box is determined. And the first loss is determined according to the regression losses of the predicted bounding boxes. When determining the second loss, for each pixel point on each predicted centerline, according to the type feature value of the pixel point and the type feature value of the corresponding pixel point in the second labels of the training samples, the loss corresponding to the pixel point is determined. And the second loss is determined according to the losses corresponding to the pixel points.

[0143] Wherein, the geometric position features at least include landmark point position features and side length features. The landmark point is the point that marks the position of the corresponding bounding box, the landmark point position features are the x value and y value in the coordinates of the landmark point. The landmark point can be the center point of the corresponding bounding box or the point where a certain corner is located. Of course, it can also be other points, which can be specifically set according to needs. The side length features are the width w and height h of the bounding box. For the geometric position features of a bounding box, it can be represented by t = {x, y, w, h}. The values of the pixel points on each centerline are different from the values of the pixel points outside the centerline. The type feature value is the value representing the type of the pixel point. As described in step S104, the values of the pixel points on each centerline are different from the values of the pixel points outside the centerline, that is, the pixel points in the centerline diagram are divided into two types: on-line points and off-line points, and the two types of points are assigned different values (type feature values).

[0144] In one or more embodiments of this specification, the server may also determine an image containing each predicted centerline based on the obtained predicted centerlines as the centerline diagram of the corresponding synthetic image of the synthetic training sample. When determining the second loss, the server may also determine the type feature values of each pixel point in the centerline diagram, and for each pixel point in the diagram, according to the type feature value of the pixel point and the type feature value of the corresponding pixel point in the second labels of the training samples, determine the loss corresponding to the pixel point, and determine the second loss of the training sample according to the losses corresponding to the pixel points.

[0145] After obtaining the first loss and the second loss, the server may sum the first loss and the second loss to determine the total loss.

[0146] In one or more embodiments of this specification, in addition to outputting each bounding box, the geometric feature detection network may also output the confidence of the prediction results of the images within each bounding box in each prediction type dimension. After determining the several feature maps corresponding to the synthetic training sample, the server may also use the several feature maps corresponding to the synthetic training sample as inputs and input them into the geometric feature detection network of the annotation model to obtain each predicted bounding box and the confidence of the prediction results of the images within each predicted bounding box in each prediction type dimension.

[0147] In one or more embodiments of this specification, among the confidences of the prediction results of the images within each predicted bounding box in each prediction type dimension, the type corresponding to the highest confidence may be used as the type corresponding to each predicted bounding box.

[0148] In one or more embodiments of this specification, when determining the first loss, the server may also determine the geometric position features of each bounding box in the first label of the training sample and the types of the characters within each bounding box. Then, for each predicted bounding box, the server may determine the regression loss of the predicted bounding box according to the difference between the geometric position features of the predicted bounding box and the geometric position features of the bounding box corresponding to the predicted bounding box in the first label of the training sample. According to the type of the character within the bounding box corresponding to the predicted bounding box in the first label of the training sample and the confidence of the prediction results of the image within the predicted bounding box in each prediction type dimension, the server may determine the classification loss of the predicted bounding box. And the first loss may be determined according to the regression losses of each predicted bounding box and the classification losses of each predicted bounding box.

[0149] In one or more embodiments of this specification, the formula for determining the first loss is specifically as follows:

[0150] L1 = L R + L C

[0151] where L1 represents the first loss, L R represents the total regression loss obtained according to the regression losses of each predicted bounding box, and L C represents the total classification loss obtained according to the classification losses of each predicted bounding box.

[0152] The formula for determining L R is specifically as follows:

[0153]

[0154] Among them, N represents the total number of predicted bounding boxes, I represents the set of all predicted bounding boxes, i represents the i-th predicted bounding box, and L Ri represents the regression loss of the i-th predicted bounding box.

[0155] The formula for determining L Ri is specifically as follows:

[0156]

[0157] Among them, t = t1 - t2, t1 represents the normalized geometric position feature of the i-th predicted bounding box, and t2 represents the normalized geometric position feature of the bounding box corresponding to the i-th predicted bounding box in the first label. t1 = {t x1 , t y1 , t w1 , t h1}, and t x1 , t y1 , t w1 , t h1 respectively represent the values after normalizing the x, y, w, and h values in the geometric position feature of the i-th predicted bounding box. t2 = {t x2 , t y2 , t w2 , t h2}, and t x2 , t y2 , t w2 , t h2 respectively represent the values after normalizing the x, y, w, and h values in the geometric position feature of the bounding box corresponding to the i-th predicted bounding box in the first label.

[0158] In one or more embodiments of this specification, the server can determine the type of the character within the bounding box corresponding to each predicted bounding box in the first label as the target type of the predicted bounding box.

[0159] In one or more embodiments of this specification, the formula for determining L C is specifically as follows:

[0160]

[0161] Among them, N represents the total number of predicted bounding boxes, I represents the set of all predicted bounding boxes, i represents the i-th predicted bounding box, and class j represents the confidence of the target type of the i-th predicted bounding box. R represents the set of confidences of the prediction results of the image within the i-th predicted bounding box in each predicted type dimension, and class r represents the confidence of the r-th predicted type. exp(class jclass representing the natural base e j power.

[0162] In one or more embodiments of this specification, the formula for determining the second loss is specifically as follows:

[0163]

[0164] Wherein, L2 represents the second loss, T represents the total number of pixel points in the center line graph of the predicted output, I represents the set of all pixel points, and i represents the i-th pixel point. q 2i represents the type feature value of the i-th pixel point, q 1i represents the type feature value of the pixel point corresponding to the i-th pixel point in the second label.

[0165] In the stage of generating synthetic training samples and training the annotation model in this specification, synthetic training samples are determined through synthetic images, accurate labels that do not require manual annotation are determined, and the annotation model used to annotate real training samples is trained based on the synthetic training samples, so as to obtain an accurate annotation model, which provides an annotation model for determining the labels of real training samples in the subsequent stage. In this stage, there is no need to manually annotate the synthetic training samples, which avoids the problem of high manual annotation cost and can train an accurate annotation model.

[0166] In one or more embodiments of this specification, in order to make the trained annotation model more accurate, before determining L C , the server can also screen the predicted bounding boxes output by the annotation model, and determine L according to the confidence levels of the prediction results of the images within each of the screened predicted bounding boxes in the predicted type dimension C , that is, I in the formula for determining L C represents the set of each screened predicted bounding box. Specifically, when screening, the server can determine the intersection over union (IoU) of each predicted bounding box and the corresponding bounding box in the first label, and for each predicted bounding box, determine whether the corresponding IoU is greater than the preset ratio upper limit. If so, the predicted bounding box is used as the predicted bounding box in set I; if not, L is not determined according to this predicted bounding box C .

[0167] In one or more embodiments of this specification, the server can also, for each predicted bounding box, determine whether the corresponding IoU is less than the preset ratio lower limit, and the bounding box simultaneously meets the condition of not overlapping with any center line. If so, the predicted bounding box is used as the predicted bounding box in set I; if not, L is not determined according to this predicted bounding box CFor a predicted bounding box with an intersection over union less than the preset lower limit of the ratio and having a center line overlapping with it, it is not regarded as a bounding box in set I, that is, L is not determined based on it. C 。

[0168] In one or more embodiments of the present specification, the geometric feature detection network includes a region detection network and a region correction network. When determining the first loss, the server can also, for each initial predicted bounding box output by the region detection network, determine the bounding box corresponding to the initial predicted bounding box in the first label according to the geometric position feature of the initial predicted bounding box, and determine the first regression loss of the initial predicted bounding box according to the initial predicted bounding box and the bounding box in the corresponding first label. And for each predicted bounding box output by the region correction network, determine the bounding box corresponding to the predicted bounding box in the first label according to the geometric position feature of the predicted bounding box, and determine the second regression loss of the predicted bounding box according to the predicted bounding box and the bounding box in the corresponding first label. Then, the first loss is determined according to each first regression loss and each second regression loss.

[0169] Figure 4 It is a schematic flowchart of the real sample processing stage provided by the embodiments of the present specification, specifically including the following steps:

[0170] S200: Obtain a number of images from the image dataset as training samples, and for each training sample, input the image corresponding to the training sample into the trained annotation model to determine each bounding box output by the annotation model, the confidence of the prediction results of the images within each bounding box in each preset type dimension, and the center lines of each string in the image corresponding to the training sample.

[0171] In one or more embodiments of the present specification, in the real sample processing stage, the server can label the training samples determined according to real non-synthetic images through the trained annotation model.

[0172] Specifically, first, the server can obtain a number of images from the image dataset as training samples, and then for each training sample, input the image corresponding to the training sample into the trained annotation model to determine the bounding boxes of each character output by the annotation model, the confidence of the prediction results of the characters within each bounding box in each preset type dimension, and the center lines of each string in the image corresponding to the training sample, so as to determine the label of the training sample from the output of the annotation model in the subsequent steps.

[0173] S202: Determine the type corresponding to each bounding box according to the confidence of the prediction results of the images within each bounding box in each preset type dimension, so as to determine each initial labeled bounding box from each bounding box according to the type corresponding to each bounding box.

[0174] Since the annotation model is trained based on synthetic images, and there are differences between synthetic images and real images, the bounding boxes output by the annotation model are not completely accurate. Moreover, not all bounding boxes can be used as the annotations of real training samples.

[0175] Therefore, in one or more embodiments of this specification, after the server obtains the bounding boxes of each character output by the annotation model, the intersection-over-union (IoU) between each bounding box and the corresponding bounding box in the first label, the confidence of the characters in each bounding box in the prediction results on each preset type dimension, and the center lines of each string in the image corresponding to the training sample, the server can screen each bounding box according to the confidence of the characters in each bounding box in the prediction results on each preset type dimension, and determine each initial labeled bounding box from the bounding boxes of each character.

[0176] Specifically, for each bounding box, the server can screen out the bounding boxes with an IoU lower than a preset value according to the IoU between each bounding box and the corresponding bounding box in the first label. Then, the server can determine the type with the highest confidence as the type to which the bounding box belongs according to the confidence of the bounding box on each preset type dimension. The server can preset a threshold, screen out the bounding boxes of characters with the highest confidence lower than the threshold, and retain the bounding boxes of characters with the highest confidence higher than the threshold as each initial labeled bounding box. In this way, the bounding boxes with inaccurate output positions and low type accuracy can be screened out.

[0177] In one or more embodiments of this specification, after screening each bounding box according to a preset threshold, the server can also continue to screen each bounding box with the highest confidence higher than the threshold according to the type of each character, and determine the bounding boxes whose type is a character as each initial labeled bounding box. For example, assume that each preset type includes: text, others, background. Then the server can determine the bounding boxes whose type is text as the initial labeled bounding boxes. In this way, the bounding boxes that do not enclose text can be screened out, so that each initial labeled bounding box obtained by screening is more accurate and suitable as the bounding box of the label of the real training sample.

[0178] S204: Determine each labeled bounding box from each initial labeled bounding box according to each initial labeled bounding box and the center lines of each string.

[0179] In one or more embodiments of this specification, the server can also continue to screen each initial labeled bounding box according to each initial labeled bounding box and the center lines of each string in the image corresponding to the training sample, and determine each labeled bounding box that finally serves as the label of the real training sample.

[0180] Specifically, the server can screen out the bounding boxes that do not overlap with each string according to the initial labeled bounding boxes and the center lines of each string, and determine the bounding boxes through which the center lines pass, that is, the bounding boxes of the characters belonging to the same string, and use the bounding boxes through which the center lines pass as the labeled bounding boxes of the labels of the true training samples.

[0181] Figure 5 This is a schematic diagram of screening bounding boxes provided in this specification. As shown in the figure, the screening process of the bounding boxes includes four stages: A, B, C, and D. In stage A, each bounding box output by the labeling model, the confidence of the prediction results of the images within each bounding box in each preset type dimension, and the center lines of each string in the image corresponding to the training sample are obtained. It can be seen that the bounding boxes of each character are not completely accurate, and there are bounding boxes that do not enclose text. In stage B, the bounding boxes with a low intersection over union can be screened out. The bounding box corresponding to the letter U is the bounding box to be screened out. In stage C, screening can be performed according to the type of the image within each bounding box, the confidence corresponding to each bounding box, and the center lines of each string. The bounding box corresponding to the bird in the lower left corner of the image is the bounding box to be screened out. The remaining bounding boxes in stage D are the labeled bounding boxes obtained after screening.

[0182] In one or more embodiments of this specification, in step S200, the labeling model can also only output the bounding boxes of each character and the center lines of each string in the image corresponding to the training sample. When screening each bounding box in steps S202 - S204, the server can screen only according to each bounding box and each center line. Determine the bounding boxes that overlap with the center lines of each string as each labeled bounding box.

[0183] S206: Use each labeled bounding box and the type of the character corresponding to each labeled bounding box as the first label of the training sample.

[0184] In one or more embodiments of this specification, the server can use the finally obtained labeled bounding boxes and the types of the characters corresponding to each labeled bounding box as the first label of the training sample determined from several images obtained from the image dataset.

[0185] In one or more embodiments of this specification, for each training sample, the center lines of each string can also be manually labeled according to each string in the training sample, and the server can use the center lines of each string manually labeled as the second label of the training sample.

[0186] In one or more embodiments of this specification, the server can also determine a centerline map consistent with the resolution of the image corresponding to the training sample based on the centerlines of each string manually marked, and use it as the second label of the training sample. Among them, the positions of the pixel points on each centerline in the centerline map are consistent with the positions of the pixel points on each centerline determined manually on the image corresponding to the training sample.

[0187] In the real sample processing stage of this specification, real training samples are determined from the image dataset, and for each real training sample, various bounding boxes, the confidence levels of each bounding box in each preset type dimension, and each centerline in the image corresponding to the real training sample are obtained through the trained annotation model. And based on the output of the annotation model, each annotated bounding box is determined from the various bounding boxes output by the annotation model, and the first label of the real training sample is determined according to each annotated bounding box. And according to each annotated bounding box, each centerline is manually marked as the second label of the real training sample.

[0188] In this stage, in addition to using each annotated bounding box as the label of the real training sample, by adding the centerline of the string as an annotation, it can be ensured that in subsequent steps, an accurate character detection model can be trained, so that the final obtained character detection result is accurate enough, reducing or even eliminating the influence brought by the error of the annotation model trained according to the synthetic training sample, so as to reduce or even eliminate the influence of the error of the label of the real training sample output by the annotation model on the final character inspection result. And although this stage manually annotates the real training sample, the bounding box annotation of the real training sample by the annotation model greatly reduces the annotation time and annotation cost.

[0189] In one or more embodiments of this specification, the server can also use only each annotated bounding box as the first label of the training sample.

[0190] Figure 6 The figure is a schematic flow chart of the text detection model training stage provided by the embodiments of this specification, which specifically includes the following steps:

[0191] S300: Obtain several images from the image dataset as training samples, and for each training sample, determine the bounding boxes of each character in the image corresponding to the training sample as the first label of the training sample, and determine the centerlines of each string in the image corresponding to the training sample as the second label of the training sample.

[0192] In one or more embodiments of the present specification, during the training phase of the text detection model, the server may use the trained annotation model as the character detection model to be trained and train the character detection model to be trained. First, the server may obtain several images from the image dataset as training samples, and for each training sample, determine the bounding box of each character in the image corresponding to the training sample as the first label of the training sample, and determine the center line of each character string in the image corresponding to the training sample, and determine the center line map according to the center lines of the character strings as the second label of the training sample.

[0193] Wherein, the first label and the second label of the training sample are the labels of the training sample determined through the process of steps S200 to S206.

[0194] S302: Input the training sample into the feature extraction network of the character detection model to be trained, and determine several feature maps corresponding to the training sample.

[0195] In one or more embodiments of the present specification, the server may input the training sample into the feature extraction network of the character detection model to be trained, and determine several feature maps corresponding to the training sample. Among them, the several feature maps corresponding to the training sample are feature maps of different scales. By performing feature extraction on the feature maps of different scales, the server can obtain more comprehensive and richer image features of the image corresponding to the training sample.

[0196] S304: Use the several feature maps corresponding to the training sample as inputs, input them into the geometric feature detection network of the character detection model to be trained to obtain each predicted bounding box, and input them into the line feature detection network of the character detection model to be trained to obtain each predicted center line.

[0197] In one or more embodiments of the present specification, the server may use the several feature maps corresponding to the training sample as inputs, input them into the geometric feature detection network of the character detection model to be trained to obtain each predicted bounding box, and input them into the line feature detection network of the character detection model to be trained to obtain each predicted center line.

[0198] Specifically, the server may use several feature maps corresponding to the training sample as inputs, input them into the region detection network in the geometric feature detection network, determine each initial predicted bounding box corresponding to each character in the output, and for each character, perform feature sampling on the regions enclosed by each initial predicted bounding box according to the initial predicted bounding boxes corresponding to the character to determine several feature matrices corresponding to the character. Then, according to the several feature matrices corresponding to the character obtained, through the region correction network in the geometric feature detection network, determine the position offset features of each initial predicted bounding box, and correct each initial predicted bounding box according to the position offset features of each initial predicted bounding box to determine the predicted bounding box of the character in the training sample.

[0199] Among them, for a character in the image corresponding to the training sample, there are multiple initial predicted bounding boxes in the initial predicted bounding boxes output by the region detection network corresponding to it.

[0200] S306: Determine the first loss according to the difference between each obtained predicted bounding box and the first label of the training sample, and determine the second loss according to the difference between each obtained predicted center line and the second label of the training sample.

[0201] In one or more embodiments of this specification, the server may determine the first loss according to the difference between each obtained predicted bounding box and the first label of the training sample.

[0202] In one or more embodiments of this specification, the server may calculate the regression loss according to the output of the region detection network and the output of the region correction network to determine the first loss.

[0203] Specifically, for each initial predicted bounding box output by the region detection network, the server may determine the bounding box corresponding to the initial predicted bounding box in the first label of the training sample according to the geometric position features of the initial predicted bounding box. And determine the first regression loss of the predicted bounding box according to the initial predicted bounding box and the bounding box in the first label of the training sample corresponding to it. And for each predicted bounding box output by the region correction network, determine the bounding box corresponding to the predicted bounding box in the first label of the training sample according to the geometric position features of the predicted bounding box. To determine the second regression loss of the predicted bounding box according to the predicted bounding box and the bounding box in the corresponding first label. Finally, determine the first loss according to each first regression loss and each second regression loss.

[0204] In one or more embodiments of this specification, the server may also calculate the regression loss and the classification loss according to the output of the region detection network and the output of the region correction network to determine the first loss.

[0205] Specifically, for each predicted bounding box output by the region correction network, the server can determine the bounding box corresponding to the predicted bounding box in the first label of the training sample according to the geometric position characteristics of the predicted bounding box. And determine the first regression loss of the predicted bounding box according to the initial predicted bounding box and the bounding box in the first label of the corresponding training sample. And determine the first classification loss of the predicted bounding box according to the confidence of the prediction result of the image in the initial predicted bounding box in each prediction type dimension and the eigenvalue of the corresponding type of the bounding box in the first label.

[0206] After that, determine the initial loss according to each first regression loss and each first classification loss. The server can also, for each predicted bounding box output by the region correction network, determine the bounding box corresponding to the predicted bounding box in the first label of the training sample according to the geometric position characteristics of the predicted bounding box. And determine the second regression loss of the predicted bounding box according to the predicted bounding box and the bounding box in the corresponding first label. For each predicted bounding box output by the region correction network, determine the second classification loss of the predicted bounding box according to the confidence of the prediction result of the image in the predicted bounding box in each prediction type dimension and the type corresponding to the bounding box in the first label of the corresponding training sample.

[0207] Then determine the correction loss according to each second regression loss and each second classification loss.

[0208] Finally, the server can determine the first loss according to the initial loss and the correction loss.

[0209] In one or more embodiments of this specification, the server can determine the second loss according to the difference between each obtained predicted center line and the second label of the training sample.

[0210] Specifically, the server can determine the type eigenvalue of each pixel point in the center line map corresponding to each obtained predicted center line. And for each pixel point, determine the loss corresponding to the pixel point according to the type eigenvalue of the pixel point and the type eigenvalue of the pixel point corresponding to the pixel point in the second label of the training sample. To determine the second loss according to the losses corresponding to each pixel point.

[0211] S308: Determine the total loss of the character detection model according to the first loss and the second loss, and use the minimum total loss as the training target to adjust the parameters of the character detection model to be trained. The character detection model is used to determine the bounding boxes and center lines of each character in the image to be detected, and expand each center line around to obtain each expanded bounding box as the character detection result of the image to be detected.

[0212] In one or more embodiments of this specification, after determining the first loss and the second loss, the server may determine the total loss of the character detection model according to the determined first loss and the second loss. Taking the minimum total loss as the training objective, the parameters of the character detection model to be trained are adjusted. The character detection model is used to determine the bounding boxes of each character and the center lines of each character string in the image to be detected, and expand the center lines of each character string around according to each bounding box to obtain each expanded bounding box as the character detection result of the image to be detected.

[0213] It should be noted that the process of training the character detection model in steps S302 - S308 is the same as the process of training the annotation model in step S106 during the training stage of the synthetic training sample generation and annotation model, and this specification will not elaborate here. The specific formulas for determining the first loss, the second loss, and the total loss in steps S307 and S308 may be the same as the formulas in step S106.

[0214] In the text detection model training stage of this specification, the trained annotation model is used as the character detection model to be trained. According to the true training samples and the labels of the true training samples determined in the previous stage, the character detection model to be trained is trained to further obtain a more accurate character detection model, so that the character detection result obtained from each bounding box and each center line output by the trained character detection model is more accurate.

[0215] In this stage, an accurate character detection model is obtained through training, so that in the subsequent stage, a more accurate character detection result is determined based on the output of the character detection model. Even if there are cases where individual bounding boxes of the character detection model are missed or misdetected for a character string in the image to be detected, the server can still determine the accurate expanded bounding box of the character string according to other accurate bounding boxes and the center line corresponding to the character string. This ensures that the final character detection result is not affected by the missed or misdetected bounding boxes, guaranteeing the accuracy of the character detection result.

[0216] In one or more embodiments of this specification, the specific structure of the annotation model may be as Figure 7 shown.

[0217] Figure 7 FIG. is a schematic structural diagram of an annotation model provided in this specification. As shown in the figure, the annotation model includes: a feature extraction network, a geometric feature detection network, and a line feature detection network. Among them, the feature extraction network includes: a first feature extraction network and a second feature extraction network. The geometric feature detection network includes: a region detection network and a region correction network. The line feature detection network includes: a detection network and a binarization module.

[0218] Figure 8The flowchart of the text detection stage provided by the embodiments of this specification specifically includes the following steps:

[0219] S400: Obtain the image to be detected, input the image into the feature extraction network in the pre-trained character detection model, and determine a plurality of feature maps corresponding to the image.

[0220] In one or more embodiments of this specification, in the text detection stage, the server can obtain the image to be detected, input the image into the trained character detection model, and determine a plurality of feature maps corresponding to the image through the feature extraction network in the trained character detection model.

[0221] Specifically, the server can input the image into the feature extraction network, and determine a plurality of initial feature maps with different scales through the first feature extraction network in the feature extraction network. Then, input the determined plurality of initial feature maps with different scales into different network layers of the second feature extraction network in the feature extraction network respectively, perform feature extraction on the plurality of initial feature maps with different scales, and determine a plurality of feature maps with different scales corresponding to the image.

[0222] Among them, the first feature extraction network can be a Residual Network (ResNet), and the second feature extraction network can be a Feature Pyramid Networks (FPN).

[0223] S402: Take the plurality of feature maps corresponding to the image as inputs, and input them into the geometric feature detection network and the line feature detection network in the character detection model respectively. Through the geometric feature detection network, determine the bounding boxes of each character in the image, and through the line feature detection network, determine the center lines in the image.

[0224] In one or more embodiments of this specification, after the server obtains the plurality of feature maps with different scales corresponding to the image, it can take the plurality of feature maps with different scales corresponding to the image as inputs, and input them into the geometric feature detection network and the line feature detection network in the character detection model respectively. Through the geometric feature detection network, determine the bounding boxes of each character in the image, and through the line feature detection network, determine the center lines in the image.

[0225] Specifically, the server can determine each initial bounding box in the image through the region detection network in the geometric feature detection network, and perform feature sampling on the regions enclosed by each initial bounding box according to each initial bounding box. For example, the regions enclosed by each initial bounding box can be feature sampled through the Region Of Interest (ROI) calibration method (i.e., ROI Align) to determine a number of feature matrices. Then, the server can determine the position offset features of each initial bounding box through the region correction network in the geometric feature detection network according to the obtained number of feature matrices, and correct each initial bounding box according to the position offset features of each initial bounding box to determine the bounding boxes of each character and the type of each character in the image.

[0226] In one or more embodiments of this specification, through the line feature detection network, the server can upsample a number of feature maps corresponding to the image to determine a number of feature maps of a specified scale. When performing upsampling, the nearest neighbor interpolation algorithm can be used, and of course other methods can also be used, which are not limited in this specification. The specified scale can be set as needed. For example, it can be 1 / 4 scale, 1 / 2 scale, etc. of the original image, which are not limited in this specification. After determining a number of feature maps of a specified scale, the server can fuse the number of feature maps of a specified scale, reduce the number of channels of the fused feature map through convolution operations without changing the scale of the feature map, and perform deconvolution on the fused feature map to perform upsampling to obtain a probability map consistent with the original scale of the image.

[0227] In one or more embodiments of this specification, the probability map is a map in which the values of the pixel points in the image are all between 0 and 1. After obtaining the probability map, the server can perform binary processing on the probability map through a binary module according to a preset probability threshold to determine each center line of the image. For example, when performing binary processing, the pixel point values greater than the probability threshold in the probability map can be set to 1, and the pixel point values less than or equal to the probability threshold can be set to 0, which can be set as needed and are not limited in this specification.

[0228] In one or more embodiments of this specification, the server can also determine the center line map corresponding to the image according to each center line in the image.

[0229] Among them, the position offset feature of each initial bounding box is the offset amount between the geometric position feature of each initial bounding box and the geometric position feature of the bounding box that should be actually marked, that is, the difference between each initial bounding box in the position coordinates of the marking point, the width of the bounding box, the height of the bounding box and the bounding box that should be actually marked. The region detection network can be an RPN (Region Proposal Network) network, and the region correction network can be a network composed of a number of fully connected layers.

[0230] S404: For each bounding box, determine the center line corresponding to the bounding box according to the overlapping degree between each center line and the bounding box.

[0231] Since there may be a situation where multiple center lines pass through the same bounding box, in one or more embodiments of this specification, the server may, for each bounding box, determine the center line corresponding to the bounding box according to the overlapping degree between each center line and the bounding box.

[0232] Specifically, the server may, for each bounding box, determine the center line with the largest overlapping area according to the overlapping areas between each center line and the bounding box, and use it as the center line corresponding to the bounding box. If there is a bounding box that does not overlap with any center line, the bounding box may be ignored.

[0233] In one or more embodiments of this specification, before determining the center line corresponding to the bounding box, the server may also screen each bounding box, and screen out the bounding boxes whose highest confidence level among the confidence levels of the prediction results of the images within the bounding boxes in each prediction type dimension is lower than the screening threshold. Since each bounding box obtained after training is already a relatively accurate bounding box, the screening threshold may be assigned a lower value, such as 0.5, 0.4, etc., which can be specifically set according to needs.

[0234] S406: Determine that the bounding boxes corresponding to the same center line are a bounding box group.

[0235] In one or more embodiments of this specification, the server may take the bounding boxes corresponding to the same center line as a bounding box group, and the area surrounded by the same bounding box group is the area where a string in the image is located.

[0236] S408: For each bounding box group, determine the dilation distance according to the geometric position characteristics of the bounding boxes in the bounding box group, and dilate the center line corresponding to the bounding box group around according to the dilation distance to determine the dilated bounding box of the bounding box group as the character detection result of the image.

[0237] In one or more embodiments of this specification, the server may, for each bounding box group, determine the dilation distance according to the geometric position characteristics of the bounding boxes in the bounding box group, and dilate the center line corresponding to the bounding box group around according to the dilation distance to determine the dilated bounding box of the bounding box group as the character detection result of the image.

[0238] For each bounding box group, the server can determine the dilation value according to the side length characteristics of each bounding box in the bounding box group, and determine the dilation distance according to the dilation distance and a preset dilation coefficient. Specifically, for each bounding box in each bounding box group, the server can determine the side length characteristic value of the bounding box according to the side length characteristics of the bounding box, and determine the dilation value according to the side length characteristic values of each bounding box. Then, the dilation distance can be determined according to the dilation value and the preset dilation coefficient.

[0239] In one or more embodiments of this specification, the formula for determining the side length characteristic value is specifically as follows:

[0240]

[0241] where D i represents the side length characteristic value of the i-th bounding box in the bounding box group, h i represents the height of the i-th bounding box, and w i represents the width of the i-th bounding box.

[0242] In one or more embodiments of this specification, the average value of the side length characteristic values of each bounding box in the bounding box group can be taken to determine the dilation value of the bounding box group Then the dilation distance is where γ represents the preset dilation coefficient, which can be set as needed. For example, it can be 0.55, 0.50, and this specification does not limit it here.

[0243] In one or more embodiments of this specification, after the server determines the dilation distance of each bounding box group, for each bounding box group, according to the dilation distance of the bounding box group, the center line corresponding to the bounding box can be dilated around to determine the dilated bounding box of the bounding box group. The server can use the finally obtained bounding box groups as the character detection result of the image.

[0244] Figure 9 This is a schematic diagram for determining the dilated bounding box provided in this specification. As shown in the figure, as can be seen from Figure A, the image includes two strings "BLOSSOM" and "SUMMER". Through the character detection model, the bounding boxes of each character in the image to be detected and the center lines of each string are obtained. After determining the dilation distance of each bounding box according to the bounding boxes of the characters in each string (i.e., the bounding boxes in each bounding box group), the center lines corresponding to each string can be dilated to obtain the boxes surrounding the two strings in Figure B, that is, the obtained dilated bounding boxes. The finally obtained two dilated bounding boxes are the character detection result of the image.

[0245] In one or more embodiments of the present specification, for each bounding box group, after obtaining the dilation distance of the bounding box group, the server can determine the edge of the center line corresponding to the bounding box group, so as to dilate the edge of the center line to obtain a dilated bounding box. When determining the edge of the center line, for each pixel point on the center line, it can be determined whether the pixel values of the pixel points around the pixel point are the same. If they are not the same, it is determined that the pixel point is a point on the edge.

[0246] In the text detection stage of the present specification, the image to be detected is input into the trained character detection model. The trained character detection model extracts features from the image to be detected, and determines the bounding boxes of each character in the image and each center line in the image. And according to the overlapping degree of each center line and each bounding box, and the bounding boxes overlapping with the same center line, bounding box groups are determined, and according to the geometric position features of each bounding box in each bounding box group, each center line is dilated around to obtain each dilated bounding box, as the character detection result of the image.

[0247] In this stage, an accurate bounding box and center line can be output by the trained character detection model to determine an accurate dilated bounding box as the character detection result. Even if there is an error in the bounding box output by the character detection model, the character detection result can still be determined according to other accurate bounding boxes output by the character detection model and the corresponding center lines, which can greatly reduce or even eliminate the influence of the output error of the character detection model and ensure the accuracy of the character detection result.

[0248] In addition, in one or more embodiments of the present specification, in step S106 and step S306 of the present specification, when determining the first loss, the loss of the region detection network and the loss of the region correction network can also be calculated respectively to determine the first loss.

[0249] The formula for determining the first loss is specifically as follows:

[0250] L1 = L S + L Rs + L Cs

[0251] Wherein, L S represents the loss of the region detection network, and L s = L S1 + L S2 L S1 represents the total classification loss of the region correction network obtained according to the classification loss of each initial predicted bounding box output by the region correction network, and L S2 represents the total classification loss of the region correction network obtained according to the regression loss of each initial predicted bounding box output by the region correction network. L RsRepresents the total regression loss of the region correction network obtained based on the regression losses of each predicted bounding box output by the region correction network, L Cs Represents the total classification loss of the region correction network obtained based on the classification losses of each bounding box output by the region correction network. The specific calculation of L Rs The process is the same as the above calculation of L R The process is not elaborated here in this specification as it is the same as above.

[0252] In one or more embodiments of this specification, the formula for determining L S1 is specifically as follows:

[0253]

[0254] where N represents the total number of initial predicted bounding boxes, I represents the set of all initial predicted bounding boxes, i represents the i-th initial predicted bounding box, and L S1i represents the regression loss of the i-th initial predicted bounding box.

[0255] In one or more embodiments of this specification, for each initial predicted bounding box, the server can take the type corresponding to the highest confidence as the type of the initial predicted bounding box. And based on the intersection over union (IoU) between the initial predicted bounding box and the corresponding bounding box in the first label, determine the eigenvalue of the type to which the corresponding bounding box in the first label belongs. Specifically, it can be determined whether the IoU between the initial predicted bounding box and the corresponding bounding box in the first label is greater than a preset ratio. If so, it is determined that the initial predicted bounding box matches the corresponding bounding box in the first label, and the eigenvalue of the type to which the corresponding bounding box in the first label belongs is determined to be 1. If not, it is determined that the predicted bounding box does not match the corresponding bounding box in the first label, and the eigenvalue of the type to which the corresponding bounding box in the first label belongs is determined to be 0. That is, when determining the loss of the region detection network, the types of each character in the first label are only divided into two types, namely character type and non-character type (background type).

[0256] In one or more embodiments of this specification, the formula for determining L S2 is specifically as follows:

[0257]

[0258] where N represents the total number of initial predicted bounding boxes, I represents the set of all initial predicted bounding boxes, i represents the i-th initial predicted bounding box, t i2 represents the confidence of the type corresponding to the i-th initial predicted bounding box, and t i1 represents the eigenvalue of the type of the bounding box corresponding to the i-th initial predicted bounding box in the first label.

[0259] In one or more embodiments provided in this specification, when determining the second loss in step S106 and step S306, in order to reduce the computational complexity and make the calculated second loss more reasonable, the server can also screen the pixel points in the predicted centerline map according to a preset ratio. Specifically, for example, assuming the preset ratio is 1:2, that is, the total number of pixel points on the centerline obtained after screening the pixel points in the predicted centerline map is 1:2 compared to the total number of pixel points not on the centerline, then the server can determine the total number of pixel points belonging to the centerline among the pixel points in the predicted centerline map, and take the total number of pixel points not on the centerline as twice the total number of pixel points on the centerline as the target to screen the pixel points not on the centerline. And perform a similar operation on the pixel points in the second label, so that the pixel points in the finally screened centerline map correspond one by one to the pixel points in the second label.

[0260] After screening the pixel points, the formula for determining the second loss is as follows:

[0261]

[0262] Where L2 represents the second loss, T represents the total number of pixel points obtained after screening in the output predicted centerline map, I represents the set of all pixel points obtained after screening, and i represents the i-th pixel point. q 2i represents the type feature value of the i-th pixel point, q 1i represents the type feature value of the pixel point corresponding to the i-th pixel point among the pixel points obtained after screening the second label.

[0263] Through the training model and character detection method provided in this specification, the annotation model is pre-trained through synthetic training samples, and the real training samples are annotated through the obtained annotation model, so as to train the character detection model according to the real training samples, which can solve the problem of high annotation cost for training samples, and can train an accurate character detection model to obtain an accurate character detection result according to the output of the character detection model. This weakly supervised method provided in this specification can reduce the annotation cost while making the character detection result output by the trained character detection model accurate enough.

[0264] The above is a method for training a model and character detection provided by one or more embodiments of this specification. Based on the same idea, this specification also provides a corresponding device.

[0265] Figure 10 The following is a schematic diagram of a device for training a character detection model provided by this specification. The device includes:

[0266] The sample label determination module 500 is configured to obtain a plurality of images from the image dataset as training samples, and for each training sample, determine the bounding boxes of each character in the image corresponding to the training sample as the first label of the training sample, and determine the centerlines of each string in the image corresponding to the training sample as the second label of the training sample;

[0267] The feature extraction module 501 is configured to input the training sample into the feature extraction network of the character detection model to be trained, and determine a plurality of feature maps corresponding to the training sample;

[0268] The prediction module 502 is configured to use the plurality of feature maps corresponding to the training sample as inputs, input them into the geometric feature detection network of the character detection model to be trained to obtain each predicted bounding box, and input them into the line feature detection network of the character detection model to be trained to obtain each predicted centerline;

[0269] The loss determination module 503 is configured to determine the first loss according to the difference between each obtained predicted bounding box and the first label of the training sample, and determine the second loss according to the difference between each obtained predicted centerline and the second label of the training sample;

[0270] The parameter adjustment module 504 is configured to determine the total loss of the character detection model according to the first loss and the second loss, and take minimizing the total loss as the training objective to adjust the parameters of the character detection model to be trained. The character detection model is used to determine the bounding boxes and centerlines of each character in the image to be detected, and expand each centerline around it according to each bounding box to obtain each expanded bounding box as the character detection result of the image to be detected.

[0271] Optionally, the prediction module 502 is configured to use the plurality of feature maps corresponding to the training sample as inputs, input them into the geometric feature detection network of the character detection model to be trained to obtain each predicted bounding box, and the confidence levels of the prediction results of the images within each predicted bounding box in each predicted type dimension.

[0272] Optionally, the loss determination module 503 is configured to determine the geometric position features of each obtained predicted bounding box and the confidence of the prediction results of the images within each predicted bounding box in each predicted type dimension, and determine the geometric position features of each bounding box in the first label of the training sample and the eigenvalue of the type to which the characters within each bounding box belong. For each predicted bounding box, according to the difference between the geometric position features of the predicted bounding box and the geometric position features of the bounding box corresponding to the predicted bounding box in the first label of the training sample, determine the regression loss of the predicted bounding box. According to the eigenvalue of the type to which the bounding box corresponding to the predicted bounding box in the first label of the training sample belongs, and the confidence of the prediction results of the images within the predicted bounding box in each predicted type dimension, determine the classification loss of the predicted bounding box. Determine the first loss according to the regression loss of each predicted bounding box and the classification loss of each predicted bounding box.

[0273] Optionally, the prediction module 502 is configured to use a plurality of feature maps corresponding to the training sample as inputs, input them into the region detection network, determine each initial predicted bounding box corresponding to each character in the output, and for each character, according to the respective initial predicted bounding boxes corresponding to the character, perform feature sampling on the regions enclosed by the respective initial predicted bounding boxes to determine a plurality of feature matrices corresponding to the character. According to the obtained plurality of feature matrices corresponding to the character, through the region correction network, determine the position offset features of each initial predicted bounding box, and according to the position offset features of each initial predicted bounding box, correct each initial predicted bounding box to determine the predicted bounding box of the character in the training sample.

[0274] Optionally, for each initial predicted bounding box output by the region detection network, the loss determination module 503 is configured to determine the bounding box corresponding to the initial predicted bounding box in the first label according to the geometric position features of the initial predicted bounding box, and determine the first regression loss of the initial predicted bounding box according to the initial predicted bounding box and the corresponding bounding box in the first label. For each predicted bounding box output by the region correction network, according to the geometric position features of the predicted bounding box, determine the bounding box corresponding to the predicted bounding box in the first label, and determine the second regression loss of the predicted bounding box according to the predicted bounding box and the corresponding bounding box in the first label. Determine the first loss according to the respective first regression losses and the respective second regression losses.

[0275] Optionally, the loss determination module 503 is configured to, for each predicted bounding box output by the region correction network, determine the bounding box corresponding to the predicted bounding box in the first label according to the geometric position feature of the predicted bounding box, determine the first regression loss of the initial predicted bounding box according to the initial predicted bounding box and the corresponding bounding box in the first label, determine the first classification loss of the predicted bounding box according to the confidence of the prediction result of the image within the initial predicted bounding box in each predicted type dimension and the eigenvalue of the corresponding type of the bounding box in the first label, determine the initial loss according to each first regression loss and each first classification loss, for each predicted bounding box output by the region correction network, determine the bounding box corresponding to the predicted bounding box in the first label according to the geometric position feature of the predicted bounding box, determine the second regression loss of the predicted bounding box according to the predicted bounding box and the corresponding bounding box in the first label, for each predicted bounding box output by the region correction network, determine the second classification loss of the predicted bounding box according to the confidence of the prediction result of the image within the predicted bounding box in each predicted type dimension and the corresponding type of the bounding box in the first label, determine the correction loss according to each second regression loss and each second classification loss, and determine the first loss according to the initial loss and the correction loss.

[0276] Optionally, the loss determination module 503 is configured to determine the image containing each predicted center line as the center line map of the training sample according to the obtained predicted center lines, determine the type eigenvalue of each pixel point in the center line map of the training sample, and for each pixel point, determine the loss corresponding to the pixel point according to the type eigenvalue of the pixel point and the type eigenvalue of the corresponding pixel point in the second label of the training sample, and determine the second loss of the training sample according to the losses corresponding to each pixel point.

[0277] Optionally, the sample label determination module 500 is configured to obtain a plurality of images from the image dataset as training samples, and for each training sample, input the image corresponding to the training sample into the trained annotation model, determine each bounding box output by the annotation model, the confidence of the prediction result of the image within each bounding box in each preset type dimension, and the center lines of each string in the image corresponding to the training sample, determine the type corresponding to each bounding box according to the confidence of the prediction result of the image within each bounding box in each preset type dimension, so as to determine each initial annotated bounding box from each bounding box according to the type corresponding to each bounding box, and determine each annotated bounding box from each initial annotated bounding box according to each initial annotated bounding box and the center lines of each string, and use each annotated bounding box and the type corresponding to each annotated bounding box as the first label of the training sample.

[0278] Optionally, the sample label determination module 500 is configured to obtain a plurality of background images and a plurality of element images from an image material library. The element images at least include images corresponding to each character type and images corresponding to each string. According to the obtained background images and element images, a plurality of synthesized images are synthesized as synthesized training samples. For each synthesized image, according to the sizes and positions of the element images in the synthesized image, the bounding boxes of the characters in the synthesized image and the types of the characters in each bounding box are determined as the first label of the synthesized training sample corresponding to the synthesized image, and the center lines of the strings in the synthesized image are determined as the second label of the synthesized training sample. The annotation model to be trained is trained according to the synthesized training samples to obtain the trained annotation model, and the annotation model is used to annotate the training samples determined from the image dataset.

[0279] The device further includes:

[0280] A control module 505, for each synthesized training sample, using the first label and the second label of the synthesized training sample as the label of the synthesized training sample, inputting the synthesized training sample into the feature extraction network of the annotation model to determine a plurality of feature maps corresponding to the synthesized training sample, using the plurality of feature maps corresponding to the synthesized training sample as inputs, inputting them into the geometric feature detection network of the annotation model to obtain each predicted bounding box, the confidence of the prediction result of the image in each predicted bounding box in each predicted type dimension, and inputting them into the line feature detection network of the annotation model to obtain each predicted center line. According to the obtained predicted bounding boxes and the difference between the confidence of the prediction result of the image in each predicted bounding box in each predicted type dimension and the first label of the synthesized training sample, a first loss is determined, and according to the difference between the obtained predicted center lines and the second label of the training sample, a second loss is determined. According to the first loss and the second loss, the total loss of the annotation model is determined, and with the goal of minimizing the total loss, the parameters of the annotation model are adjusted.

[0281] Figure 11 The figure is a schematic diagram of a character detection device provided in this specification. The device includes:

[0282] A feature extraction module 600, configured to obtain an image to be detected, input the image into the feature extraction network in a pre-trained character detection model, and determine a plurality of feature maps corresponding to the image;

[0283] A feature output module 601, configured to use the plurality of feature maps corresponding to the image as inputs, and respectively input them into the geometric feature detection network and the line feature detection network in the character detection model. Through the geometric feature detection network, the bounding boxes of the characters in the image are determined, and through the line feature detection network, the center lines in the image are determined;

[0284] The correspondence determination module 602 is configured to determine, for each bounding box, the center line corresponding to the bounding box according to the overlapping degree between each center line and the bounding box.

[0285] The bounding box group determination module 603 is configured to determine that the bounding boxes corresponding to the same center line are a bounding box group.

[0286] The detection result determination module 604 is configured to, for each bounding box group, determine the dilation distance according to the geometric position features of the bounding boxes in the bounding box group, and dilate the center line corresponding to the bounding box group around according to the dilation distance to determine the dilated bounding box of the bounding box group as the character detection result of the image.

[0287] Optionally, the feature output module 601 is configured to determine, through the region detection network, the initial bounding boxes of the characters in the image, sample the features of the regions surrounded by the initial bounding boxes according to the initial bounding boxes to determine a plurality of feature matrices, determine the position offset features of the initial bounding boxes through the region correction network according to the obtained plurality of feature matrices, and correct the initial bounding boxes according to the position offset features of the initial bounding boxes to determine the bounding boxes of the characters in the image.

[0288] Optionally, the feature output module 601 is configured to upsample the plurality of feature maps corresponding to the image through the line feature detection network to determine a plurality of feature maps with specified scales, fuse the plurality of feature maps with specified scales to reduce the number of channels of the fused feature maps, and upsample the fused feature maps to obtain a probability map consistent with the original scale of the image, and perform binarization processing on the probability map to determine the center lines corresponding to the image and the center line map corresponding to the image.

[0289] Optionally, the detection result determination module 604 is configured to determine the dilation value according to the side length features of the bounding boxes in the bounding box group, and determine the dilation distance according to the dilation value and a preset dilation coefficient.

[0290] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the above-provided method for training a model and character detection.

[0291] This specification also provides Figure 12 a schematic structural diagram of the electronic device shown. As Figure 12As shown in the figure, at the hardware level, the electronic device includes a processor, an internal bus, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above methods for training the model and character detection.

[0292] Of course, in addition to the software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution entity of the following processing flow is not limited to each logic unit, and can also be hardware or a logic device.

[0293] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL), and there is not just one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0294] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that, in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.

[0295] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0296] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0297] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0298] The present invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block of the flowchart illustrations and / or block diagrams, and combinations of flows and / or blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing apparatus create means for implementing the functions specified in the flowchart Figure 1 for one or more flows and / or blocks Figure 1 of the flowchart and / or means for implementing the functions specified in one or more blocks of the block diagram.

[0299] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions specified in the flowchart Figure 1 for one or more flows and / or blocks Figure 1 of the flowchart and / or means for implementing the functions specified in one or more blocks of the block diagram.

[0300] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart Figure 1 for one or more flows and / or blocks Figure 1 of the flowchart and / or means for implementing the functions specified in one or more blocks of the block diagram.

[0301] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0302] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0303] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0304] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.

[0305] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, system or computer program product. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0306] This specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0307] Each embodiment in this specification is described in a progressive manner. For the identical or similar parts among the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the description of the method embodiment.

[0308] The above description is only for the embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various modifications and changes can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.

Claims

1. A method for training a character detection model, characterized in that, Including: Obtain a number of images from an image dataset as training samples, and for each training sample, determine the bounding boxes of each character in the image corresponding to the training sample as the first label of the training sample, and determine the centerlines of each string in the image corresponding to the training sample as the second label of the training sample; Wherein, the first label of the training sample further includes the types of characters within each bounding box in the image corresponding to the training sample; input the training sample into the feature extraction network of the character detection model to be trained to determine a number of feature maps corresponding to the training sample; use the number of feature maps corresponding to the training sample as input and input it into the geometric feature detection network of the character detection model to be trained to obtain each predicted bounding box and the confidence of the prediction results of the images within each predicted bounding box in each predicted type dimension; and input it into the line feature detection network of the character detection model to be trained to obtain each predicted centerline; Determine the first loss according to the differences between the obtained predicted bounding boxes and the first label of the training sample, specifically including: determining the geometric position features of the obtained predicted bounding boxes and the confidence of the prediction results of the images within each predicted bounding box in each predicted type dimension, and determining the geometric position features of each bounding box and the feature values of the types of characters within each bounding box in the first label of the training sample; for each predicted bounding box, determine the regression loss of the predicted bounding box according to the difference between the geometric position features of the predicted bounding box and the geometric position features of the bounding box corresponding to the predicted bounding box in the first label of the training sample; determine the classification loss of the predicted bounding box according to the feature value of the type of the bounding box corresponding to the predicted bounding box in the first label of the training sample and the confidence of the prediction results of the images within the predicted bounding box in each predicted type dimension; determine the first loss according to the regression loss of each predicted bounding box and the classification loss of each predicted bounding box; Determine the second loss according to the differences between the obtained predicted centerlines and the second label of the training sample, specifically including: according to the obtained predicted centerlines, determine the image containing each predicted centerline as the centerline map of the training sample; determine the type feature values of each pixel point in the centerline map of the training sample; for each pixel point, determine the loss corresponding to the pixel point according to the type feature value of the pixel point and the type feature value of the pixel point corresponding to the pixel point in the second label of the training sample; determine the second loss of the training sample according to the losses corresponding to each pixel point; determine the total loss of the character detection model according to the first loss and the second loss, and use the minimum total loss as the training objective to adjust the parameters of the character detection model to be trained. The character detection model is used to determine the bounding boxes and centerlines of each character in the image to be detected, and expand each centerline around it according to each bounding box to obtain each expanded bounding box as the character detection result of the image to be detected.

2. The method according to claim 1, wherein The geometric feature detection network includes a region detection network and a region correction network; taking a plurality of feature maps corresponding to the training sample as inputs and inputting them into the geometric feature detection network of the character detection model to be trained to obtain each predicted bounding box, specifically including: taking a plurality of feature maps corresponding to the training sample as inputs and inputting them into the region detection network to determine each initial predicted bounding box corresponding to each character in the output; for each character, according to each initial predicted bounding box corresponding to the character, performing feature sampling on the regions enclosed by each initial predicted bounding box to determine a plurality of feature matrices corresponding to the character; according to the obtained plurality of feature matrices corresponding to the character, passing them through the region correction network to determine the position offset features of each initial predicted bounding box, and correcting each initial predicted bounding box according to the position offset features of each initial predicted bounding box to determine the predicted bounding box of the character in the training sample.

3. The method according to claim 2, wherein Determining a first loss according to the difference between each obtained predicted bounding box and the first label of the training sample, specifically including: for each initial predicted bounding box output by the region detection network, determining the bounding box corresponding to the initial predicted bounding box in the first label according to the geometric position features of the initial predicted bounding box; determining the first regression loss of the initial predicted bounding box according to the initial predicted bounding box and the bounding box in the first label corresponding to it; for each predicted bounding box output by the region correction network, determining the bounding box corresponding to the predicted bounding box in the first label according to the geometric position features of the predicted bounding box; determining the second regression loss of the predicted bounding box according to the predicted bounding box and the bounding box in the first label corresponding to it; determining the first loss according to each first regression loss and each second regression loss.

4. The method according to claim 3, characterized in that, The first label of the training sample further includes the types of characters in each bounding box in the image corresponding to the training sample. The region detection network and the region correction network also respectively output the confidence levels of the prediction results of the image within the initial predicted bounding box in each predicted type dimension, and the confidence levels of the prediction results of the image within the predicted bounding box in each predicted type dimension. The method further includes: for each predicted bounding box output by the region correction network, determining the bounding box in the first label corresponding to the predicted bounding box according to the geometric position features of the predicted bounding box; determining the first regression loss of the initial predicted bounding box according to the initial predicted bounding box and the corresponding bounding box in the first label; determining the first classification loss of the predicted bounding box according to the confidence levels of the prediction results of the image within the initial predicted bounding box in each predicted type dimension and the feature values of the corresponding types of the bounding boxes in the first label; determining the initial loss according to each first regression loss and each first classification loss; for each predicted bounding box output by the region correction network, determining the bounding box in the first label corresponding to the predicted bounding box according to the geometric position features of the predicted bounding box; determining the second regression loss of the predicted bounding box according to the predicted bounding box and the corresponding bounding box in the first label; for each predicted bounding box output by the region correction network, determining the second classification loss of the predicted bounding box according to the confidence levels of the prediction results of the image within the predicted bounding box in each predicted type dimension and the corresponding type of the bounding box in the first label; determining the correction loss according to each second regression loss and each second classification loss; determining the first loss according to the initial loss and the correction loss.

5. The method according to claim 1, wherein Obtain a number of images from the image dataset as training samples, and for each training sample, determine the bounding boxes of each character in the image corresponding to the training sample as the first label of the training sample. Specifically, it includes: obtaining a number of images from the image dataset as training samples, and for each training sample, inputting the image corresponding to the training sample into the trained annotation model, and determining each bounding box output by the annotation model, the confidence levels of the prediction results of the images within each bounding box in each preset type dimension, and the center lines of each string in the image corresponding to the training sample; determining the type corresponding to each bounding box according to the confidence levels of the prediction results of the images within each bounding box in each preset type dimension, so as to determine each initial annotated bounding box from each bounding box according to the type corresponding to each bounding box; determining each annotated bounding box from each initial annotated bounding box according to each initial annotated bounding box and the center lines of each string; taking each annotated bounding box and the type corresponding to each annotated bounding box as the first label of the training sample.

6. The method according to claim 1, characterized in that, The training samples for training the annotation model are determined by the following method: Obtain a number of background images and a number of element images from the image material library, where the element images at least include images corresponding to each character type and images corresponding to each string; According to the obtained background images and element images, synthesize a number of synthetic images as synthetic training samples; For each synthetic image, according to the sizes and positions of the element images in the synthetic image, determine the bounding boxes of the characters in the synthetic image and the types of the characters within each bounding box as the first label of the synthetic training sample corresponding to the synthetic image, and determine the center lines of the strings in the synthetic image as the second label of the synthetic training sample; Train the annotation model to be trained according to the synthetic training samples to obtain the trained annotation model, and the annotation model is used to annotate the training samples determined from the image dataset.

7. The method according to claim 6, wherein Use the trained annotation model as the character detection model to be trained, and the annotation model is trained in the following manner: For each synthetic training sample, use the first label and the second label of the synthetic training sample as the label of the synthetic training sample; Input the synthetic training sample into the feature extraction network of the annotation model to determine a number of feature maps corresponding to the synthetic training sample; Use the number of feature maps corresponding to the synthetic training sample as the input and input it into the geometric feature detection network of the annotation model to obtain each predicted bounding box, the confidence of the prediction results of the images within each predicted bounding box in each predicted type dimension, and input it into the line feature detection network of the annotation model to obtain each predicted center line; Determine the first loss according to the obtained predicted bounding boxes and the difference between the confidence of the prediction results of the images within each predicted bounding box in each predicted type dimension and the first label of the synthetic training sample, and determine the second loss according to the difference between the obtained predicted center lines and the second label of the training sample; Determine the total loss of the annotation model according to the first loss and the second loss, and use the minimum total loss as the training objective to adjust the parameters of the annotation model.

8. A method for character detection, characterized in that, The character detection model is trained according to the method for training a character detection model described in any one of claims 1 to 7, and includes: obtaining an image to be detected, inputting the image into a feature extraction network in a pre-trained character detection model to determine a plurality of feature maps corresponding to the image; using the plurality of feature maps corresponding to the image as inputs, and respectively inputting them into a geometric feature detection network and a line feature detection network in the character detection model. Through the geometric feature detection network, determine the bounding boxes of each character in the image, and through the line feature detection network, determine the center lines in the image; for each bounding box, determine the center line corresponding to the bounding box according to the overlapping degree between each center line and the bounding box; determine that the bounding boxes corresponding to the same center line are a bounding box group; for each bounding box group, determine an expansion distance according to the geometric position features of the bounding boxes in the bounding box group, and expand the center line corresponding to the bounding box group around according to the expansion distance to determine the expanded bounding box of the bounding box group, which is used as the character detection result of the image.

9. The method according to claim 8, wherein The geometric feature detection network includes a region detection network and a region correction network; determining the bounding boxes of each character in the image through the geometric feature detection network specifically includes: determining each initial bounding box of each character in the image through the region detection network; performing feature sampling on the regions enclosed by each initial bounding box according to each initial bounding box to determine a plurality of feature matrices; according to the obtained plurality of feature matrices, determining the position offset features of each initial bounding box through the region correction network, and correcting each initial bounding box according to the position offset features of each initial bounding box to determine the bounding boxes of each character in the image.

10. The method according to claim 9, characterized in that, Determining the center lines in the image through the line feature detection network specifically includes: performing upsampling on the plurality of feature maps corresponding to the image through the line feature detection network to determine a plurality of feature maps with specified scales; fusing the plurality of feature maps with specified scales, reducing the number of channels of the fused feature map, and performing upsampling on the fused feature map to obtain a probability map consistent with the original scale of the image; performing binarization processing on the probability map to determine the center lines corresponding to the image and the center line map corresponding to the image.

11. The method according to claim 10, characterized in that, The geometric position features of each bounding box at least include side length features; determining the expansion distance according to the geometric position features of the bounding boxes in the bounding box group specifically includes: determining an expansion value according to the side length features of the bounding boxes in the bounding box group; determining the expansion distance according to the expansion value and a preset expansion coefficient.

12. An apparatus for training a character detection model, characterized in that, Including: A sample label determination module, configured to obtain a plurality of images from an image dataset as training samples, and for each training sample, determine the bounding boxes of each character in the image corresponding to the training sample as the first label of the training sample, and determine the centerlines of each string in the image corresponding to the training sample as the second label of the training sample; wherein, the first label of the training sample further includes the types of the characters within each bounding box in the image corresponding to the training sample; A feature extraction module, configured to input the training sample into the feature extraction network of the character detection model to be trained, and determine a plurality of feature maps corresponding to the training sample; A prediction module, configured to input the plurality of feature maps corresponding to the training sample into the geometric feature detection network of the character detection model to be trained, to obtain each predicted bounding box and the confidence of the prediction results of the images within each predicted bounding box in each predicted type dimension; and input the plurality of feature maps corresponding to the training sample into the line feature detection network of the character detection model to be trained, to obtain each predicted centerline; A loss determination module, configured to determine a first loss according to the differences between the obtained predicted bounding boxes and the first label of the training sample, specifically including: determining the geometric position features of the obtained predicted bounding boxes and the confidence of the prediction results of the images within each predicted bounding box in each predicted type dimension, and determining the geometric position features of each bounding box in the first label of the training sample and the feature values of the types of the characters within each bounding box; for each predicted bounding box, determining the regression loss of the predicted bounding box according to the difference between the geometric position features of the predicted bounding box and the geometric position features of the bounding box corresponding to the predicted bounding box in the first label of the training sample; determining the classification loss of the predicted bounding box according to the feature value of the type of the bounding box corresponding to the predicted bounding box in the first label of the training sample and the confidence of the prediction results of the images within the predicted bounding box in each predicted type dimension; determining the first loss according to the regression loss of each predicted bounding box and the classification loss of each predicted bounding box; Determining a second loss according to the differences between the obtained predicted centerlines and the second label of the training sample, specifically including: according to the obtained predicted centerlines, determining the image containing each predicted centerline as the centerline map of the training sample; determining the type feature values of each pixel point in the centerline map of the training sample; for each pixel point, determining the loss corresponding to the pixel point according to the type feature value of the pixel point and the type feature value of the pixel point corresponding to the pixel point in the second label of the training sample; determining the second loss of the training sample according to the losses corresponding to each pixel point; A parameter adjustment module, configured to determine the total loss of the character detection model according to the first loss and the second loss, and taking the minimum total loss as the training objective, adjust the parameters of the character detection model to be trained, and the character detection model is used to determine the bounding boxes and centerlines of each character in the image to be detected, and to expand each centerline around according to each bounding box to obtain each expanded bounding box as the character detection result of the image to be detected.

13. A character detection device, characterized in that, The character detection model is trained based on the device for training the character detection model according to claim 12, and includes: a feature extraction module, configured to obtain an image to be detected, input the image into a feature extraction network in a pre-trained character detection model, and determine a plurality of feature maps corresponding to the image; a feature output module, configured to use the plurality of feature maps corresponding to the image as inputs, and respectively input them into a geometric feature detection network and a line feature detection network in the character detection model, determine a bounding box of each character in the image through the geometric feature detection network, and determine each center line in the image through the line feature detection network; a correspondence determination module, configured to, for each bounding box, determine a center line corresponding to the bounding box according to the degree of overlap between each center line and the bounding box; a bounding box group determination module, configured to determine each bounding box corresponding to the same center line as a bounding box group; a detection result determination module, configured to, for each bounding box group, determine an expansion distance according to the geometric position features of each bounding box in the bounding box group, expand the center line corresponding to the bounding box group around according to the expansion distance, and determine an expanded bounding box of the bounding box group as the character detection result of the image.

14. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 11 above is implemented.

15. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the program, the method according to any one of claims 1 to 11 above is implemented.

Citation Information

Patent Citations

  • Method and device for generating training samples

    CN110991520A

  • Text detection method and device and recognition system

    CN111027563A