A method and apparatus for generating training samples
By using a preset recognition model to identify the annotation image, obtaining supplementary annotation information and generating training samples, the problem of time-consuming and labor-consuming manual annotation in the prior art is solved, and the efficiency of training samples is improved.
Patent Information
- Application Number
- CN202110230635.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-02
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-03-02
AI Technical Summary
In the prior art, the annotation process required to generate training samples consumes a lot of labor costs and time, resulting in inefficiency.
The preset recognition model recognizes the annotated image, obtains the recognition results of text lines and individual text, determines the supplementary annotation information, and generates training samples.
It effectively improves the efficiency of training sample generation, reduces the cost and time of manual annotation, and improves the training speed of recognition models.
Smart Images

Figure CN113011424B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of machine learning, and particularly to a method and apparatus for generating training samples. Background Art
[0002] In practical applications, a service platform needs to identify text in images through a machine learning model, and a large number of training samples are required to train the machine learning model.
[0003] For example, the service platform needs to identify an image of a merchant signboard through a pre-trained recognition model to obtain the image area where the text is located in the signboard image and the content of the text. During the process of training the recognition model, a large number of training samples are required to train the recognition model.
[0004] In the prior art, the training samples required for training the recognition model can be obtained by manually annotating images. The service platform can train the recognition model with the training samples obtained after annotating the images. However, this annotation method consumes a large amount of labor costs and time, greatly reducing the efficiency of generating training samples.
[0005] Therefore, how to improve the efficiency of generating training samples and reduce labor costs is an urgent problem to be solved. Summary of the Invention
[0006] This specification provides a method and apparatus for generating training samples to partially solve the above problems existing in the prior art.
[0007] This specification adopts the following technical solutions:
[0008] This specification provides a method for generating training samples, including:
[0009] Obtaining an image to be annotated and the corresponding text annotation information of the image to be annotated;
[0010] Inputting the image to be annotated into a preset recognition model to obtain an overall recognition result for the text line included in the image to be annotated as the first recognition result, and a single-character recognition result for at least some single characters included in the image to be annotated as the second recognition result;
[0011] Determining, according to the first recognition result and the second recognition result, other annotation information for the image to be annotated except the text annotation information as supplementary annotation information;
[0012] Supplement the text annotation information according to the supplementary annotation information to obtain the supplemented annotation information, and generate a training sample corresponding to the image to be annotated through the supplemented annotation information, so as to train the recognition model with the training sample.
[0013] Optionally, the first recognition result includes: the text content of the text line recognized from the image to be annotated and the image region where the text line is located in the image to be annotated, and the second recognition result includes: at least some single characters recognized from the image to be annotated and the image region where each character in the at least some single characters is located in the image to be annotated.
[0014] Optionally, determine other annotation information for the image to be annotated except the text annotation information according to the first recognition result and the second recognition result as supplementary annotation information, specifically including:
[0015] If it is determined that there is a single character in the at least some single characters that matches the text line, the determined single character that matches the text line is used as the target character and added to the character set;
[0016] If it is determined that at least one of the character set and the text line meets a preset condition, determine the supplementary annotation information according to the first recognition result and / or the single-character recognition result corresponding to each character included in the character set in the second recognition result.
[0017] Optionally, determining the single character that matches the text line as the target character specifically includes:
[0018] For each character in the at least some single characters, determine the degree of overlap between the image region where the character is located in the image to be annotated and the image region where the text line is located in the image to be annotated;
[0019] If it is determined that the degree of overlap is not less than the set degree of overlap, determine that the character matches the text line and use the character as the target character.
[0020] Optionally, determining that at least one of the character set and the text line meets a preset condition specifically includes:
[0021] Sort each character included in the character set according to the order of the characters in the text line to obtain a reference text line;
[0022] Determine the matching degree between the text line and the text annotation information as the first matching degree, and determine the matching degree between the reference text line and the text annotation information as the second matching degree;
[0023] If it is determined that at least one of the first matching degree and the second matching degree is not less than the first set matching degree, it is determined that at least one of the text set and the text line meets the preset condition.
[0024] Optionally, according to the first recognition result and / or the single-character recognition results corresponding to each character included in the text set in the second recognition result, the supplementary annotation information is determined, specifically including:
[0025] If it is determined that the number of characters included in the reference text line is less than the number of characters included in the text annotation information, it is determined whether the character at the set position in the reference text line is the same as the character at the set position in the text annotation information;
[0026] If so, according to the first recognition result, the supplementary annotation information is determined;
[0027] If not, other characters except the target character are recognized from the image to be annotated as supplementary characters, and the supplementary characters are added to the text set, so as to determine the supplementary annotation information according to the single-character recognition results corresponding to each character except the supplementary characters in the supplementary text set and the single-character recognition result corresponding to the supplementary characters.
[0028] Optionally, other characters except the at least partial single characters are recognized from the image to be annotated as supplementary characters, specifically including:
[0029] Determine the overall text slope of the text line in the image to be annotated;
[0030] According to the overall text slope, other characters except the target character are recognized from the image to be annotated as supplementary characters.
[0031] Optionally, according to the first recognition result and / or the single-character recognition results corresponding to each character included in the text set in the second recognition result, the supplementary annotation information is determined, specifically including:
[0032] If it is determined that the number of characters included in the reference text line is equal to the number of characters included in the text annotation information, the supplementary annotation information is determined according to the single-character recognition results corresponding to each character included in the text set in the second recognition result.
[0033] Optionally, the method further includes:
[0034] If it is determined that the number of characters included in the reference text line is greater than the number of characters included in the text annotation information, the text annotation information corresponding to the text line is not supplemented.
[0035] Optionally, if it is determined that neither the set of characters nor the text line satisfies the preset condition, the method further includes:
[0036] Not supplementing the text annotation information corresponding to the text line.
[0037] Optionally, according to the first recognition result and the second recognition result, other annotation information for the to-be-annotated image except the text annotation information is determined as supplementary annotation information, specifically including:
[0038] If it is determined that none of the at least part of the single characters contains a single character that matches the text line, the matching degree between the first recognition result and the text annotation information corresponding to the text line is determined as the third matching degree;
[0039] If it is determined that the third matching degree is not less than the second set matching degree, the supplementary annotation information is determined according to the first recognition result.
[0040] Optionally, the method further includes:
[0041] If it is determined that the third matching degree is less than the second set matching degree, the text annotation information corresponding to the text line is not supplemented.
[0042] This specification provides a training sample generation device, including:
[0043] An acquisition module, configured to acquire a to-be-annotated image and the text annotation information corresponding to the to-be-annotated image;
[0044] A recognition module, configured to input the to-be-annotated image into a preset recognition model to obtain an overall recognition result of the text line included in the to-be-annotated image as the first recognition result, and a single-character recognition result of at least part of the single characters included in the to-be-annotated image as the second recognition result;
[0045] A determination module, configured to determine other annotation information for the to-be-annotated image except the text annotation information as supplementary annotation information according to the first recognition result and the second recognition result;
[0046] A training module, configured to supplement the text annotation information according to the supplementary annotation information to obtain supplemented annotation information, and generate a training sample corresponding to the to-be-annotated image through the supplemented annotation information, so as to train the recognition model through the training sample.
[0047] This specification provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the above-mentioned method for generating training samples.
[0048] This specification provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the above-mentioned method for generating training samples.
[0049] The above-mentioned at least one technical solution adopted in this specification can achieve the following beneficial effects:
[0050] In the method and device for generating training samples provided in this specification, a service platform can obtain an image to be labeled and the corresponding text annotation information of the image to be labeled, and input the image to be labeled into a preset recognition model to obtain an overall recognition result for the text lines included in the image to be labeled as the first recognition result, and a single-character recognition result for at least some of the individual characters included in the image to be labeled as the second recognition result. Then, according to the first recognition result and the second recognition result, other annotation information for the image to be labeled except the text annotation information is determined as supplementary annotation information. According to the supplementary annotation information, the text annotation information is supplemented to obtain the supplemented annotation information, and a training sample corresponding to the image to be labeled is generated through the supplemented annotation information to train the recognition model with the training sample.
[0051] It can be seen from the above method that the image to be labeled corresponds to some annotation information. The service platform needs to determine other annotation information except these annotation information, and label the image to be labeled to obtain a training sample. Therefore, the service platform can use a pre-trained recognition model to determine the relevant information of the text lines and individual characters in the image to be labeled, and directly determine other annotation information except the text annotation information according to the relevant information of the text lines and individual characters in the image to be labeled, and label the image to be labeled. Therefore, compared with the prior art, this method can efficiently generate training samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The drawings described herein are used to provide a further understanding of this specification and form a part of this specification. The illustrative embodiments and descriptions thereof of this specification are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings:
[0053] Figure 1 is a flowchart showing a method for generating a training sample in this specification;
[0054] Figure 2 is a schematic diagram of a recognition model provided in this specification;
[0055] Figure 3 A schematic diagram for matching the first recognition result, the second recognition result and the text annotation information provided in this specification;
[0056] Figure 4 A schematic diagram for determining supplementary annotation information of an image to be annotated provided in this specification;
[0057] Figure 5 A schematic diagram of a generating device for training samples in this specification;
[0058] Figure 6 Corresponding to that provided in this specification Figure 1 Schematic diagram of an electronic device. Specific embodiments
[0059] To make the objectives, technical solutions and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of this specification.
[0060] The following will, with reference to the drawings, elaborate on the technical solutions provided in each embodiment of this specification.
[0061] Figure 1 A flowchart of a method for generating training samples in this specification, specifically including the following steps:
[0062] S101: Obtain an image to be annotated and the text annotation information corresponding to the image to be annotated.
[0063] S102: Input the image to be annotated into a preset recognition model to obtain an overall recognition result for the text lines included in the image to be annotated as the first recognition result, and a single-character recognition result for at least some of the single characters included in the image to be annotated as the second recognition result.
[0064] Currently, machine learning models can be well applied in the text recognition scenario. The business platform needs to determine the positions and contents of all single characters and text lines in the image through the recognition model. Therefore, the business platform needs to obtain a large number of training samples to train the recognition model. In actual applications, the business platform can obtain a large number of images to be annotated that have been marked with certain annotation information. For each image to be annotated, the business platform needs to determine some annotation information other than the already marked annotation information to obtain complete training samples.
[0065] Based on this, the business platform can hold the image to be annotated and the corresponding text annotation information of the image to be annotated, and input the image to be annotated into a preset recognition model to obtain the overall recognition result for the text lines contained in the image to be annotated as the first recognition result, and the single-character recognition result for at least some of the single characters contained in the image to be annotated as the second recognition result. The text annotation information mentioned here can refer to the text content of each text line in the image to be annotated that is marked, or it can refer to the image area where each text line marked is located in the image to be annotated. If what is marked is the text content of the text line, then the text area where the text line is located needs to be determined. If what is marked is the text area where the text line is located, then the text content of the text line needs to be determined.
[0066] Among them, the first recognition result can include the text content of the text lines recognized from the image to be annotated and the image area where the text lines are located in the image to be annotated. The second recognition result can include at least some of the single characters recognized from the image to be annotated and the image area where each of the at least some single characters is located in the image to be annotated. In the subsequent process, the business platform can determine some information that can annotate the image to be annotated in addition to the above text annotation information based on the first recognition result and the second recognition result, that is, the supplementary annotation information that will be mentioned later.
[0067] It should be noted that the above recognition model needs to be trained in advance with a number of images marked with complete annotation information. After the recognition model has a certain ability to recognize the text in the image, the business platform can determine the training samples through this recognition model. The recognition model includes a character recognition sub-model for recognizing single characters and a text line recognition sub-model for recognizing text lines. The character recognition sub-model for recognizing single characters needs to recognize each character in the image and the image area where each character is located, and the text line recognition sub-model needs to recognize each text line in the image and the image area where each text line is located, as Figure 2 shown.
[0068] Figure 2 This is a schematic diagram of a recognition model provided in this specification.
[0069] In Figure 2It can be seen that the text line recognition sub-model includes a first region recognition sub-model and a text recognition sub-model. The character recognition sub-model includes a second region recognition sub-model and a character recognition sub-model. The recognition model also includes a feature extraction model. After the feature extraction model extracts features from the image to be labeled, the first region recognition sub-model determines the image region where the text line is located in the image to be recognized. The text recognition sub-model determines the text content of the text line. The second region recognition sub-model determines the image region where each character is located in the image to be labeled, and the character recognition sub-model determines each character in the image to be recognized.
[0070] Among them, the recognition model can be built through conventional image recognition algorithms, feature extraction algorithms, etc. For example, the feature extraction model can be ResNet50 and FPN. The first region recognition sub-model is a convolutional neural network and a fully connected layer. The text recognition sub-model is composed of the decoder part of Seq2Seq. The second region recognition sub-model and the character recognition sub-model are convolutional neural networks, fully connected layers, etc.
[0071] S103: According to the first recognition result and the second recognition result, determine other annotation information for the image to be labeled except the text annotation information as supplementary annotation information.
[0072] The service platform can determine other annotation information for the image to be labeled for selection except the above text annotation information as supplementary annotation information according to the first recognition result and the second recognition result. In order to fully obtain supplementary annotation information that can supplement the text annotation information from the first recognition result and the second recognition result, the service platform can match the first recognition result, the second recognition result, and the text annotation information in various ways, such as Figure 3 shown.
[0073] Figure 3 This is a schematic diagram provided in this specification for matching the first recognition result, the second recognition result, and the text annotation information.
[0074] From Figure 3 it can be seen that the service platform can determine that at least some of the single characters contain single characters that match the text line, and use the determined single characters that match the text line as target characters and add them to the character set. If the service platform determines that at least one of the character set and the text line meets the preset conditions, it can determine the supplementary annotation information according to the first recognition result and / or the single-character recognition results corresponding to each character included in the character set in the second recognition result.
[0075] Among them, there are various ways for the service platform to determine a single character that matches the text line. For example, the service platform can determine, for each character in at least some of the single characters, the degree of overlap between the image area where the character is located in the image to be labeled and the image area where the text line is located in the image to be labeled. If it is determined that the degree of overlap is not less than the set degree of overlap, it can be determined that the character matches the text line, and the character is used as the target character. That is to say, this is to determine whether there is a relatively large overlapping part between the images occupied by the character and the text line respectively in the image to be labeled. If there is a relatively large part, it can be determined that the character is in the text line, and the character can be added to the character set.
[0076] Of course, the service platform can also determine whether the character matches the text line through other methods. For example, the service platform can directly determine the characters in the text line from at least some of the recognized single characters, and then add them to the text set. If there are multiple text lines in the image, then for a text line, these two methods can also be combined to determine the characters that are both in the text line and have a relatively high degree of overlap with the image area of the text line, and add them to the character set.
[0077] In order to obtain accurate supplementary annotation information, the service platform also needs to match the text annotation information with the text line. If the text line cannot be matched with the text annotation information, the supplementary annotation information cannot be determined through the recognition result related to the text line. Specifically, if the service platform can determine that at least one of the above-mentioned character set and the text line meets the preset conditions, the supplementary annotation information can be determined according to the first recognition result and / or the single-character recognition result corresponding to each character included in the character set in the second recognition result. If the service platform determines that neither the character set nor the text line meets the preset conditions, there is no need to supplement the text annotation information corresponding to the text line.
[0078] Among them, there can be various preset conditions. For example, the service platform can sort each character included in the character set according to the order of the characters in the text line to obtain a reference text line, and determine the matching degree between the text line and the text annotation information as the first matching degree, and determine the matching degree between the reference text line and the text annotation information as the second matching degree.
[0079] In practical applications, there can be multiple ways to determine the first matching degree and the second matching degree. For example, it can be determined by the text edit distance. The larger the text edit distance, the smaller the matching degree. Or first, the feature vectors of each text (i.e., the recognized text line, the reference text line, and the text annotation information) are determined according to a deep learning model, and then the first matching degree and the second matching degree are determined through the feature vectors of each text. If at least one of the first matching degree and the second matching degree is not less than the first set matching degree, it can be determined that at least one of the text set and the text line meets the preset conditions, where the first set matching degree can be set according to the actual situation.
[0080] For another example, after determining the reference text line, if the business platform determines that at least one of the text line and the reference text line is exactly the same as the text annotation information, it can be determined that at least one of the text set and the text line meets the preset conditions. After the business platform determines that at least one of the text set and the text line meets the preset conditions, the text line is matched with the text annotation information. In this way, the image area where the text line is located in the first recognition result is very likely to be the position of the text content annotated for the text line in the text annotation information in the image to be annotated. And the image areas where each character in the text set is located in the image to be annotated may be able to more precisely represent the position where the text line is located. Therefore, it is also necessary to determine whether to use the first recognition result or the single-character recognition result corresponding to each character in the text set to determine the supplementary annotation information.
[0081] Therefore, after determining that at least one of the text set and the text line meets the preset conditions, if the business platform determines that the number of characters included in the reference text line is equal to the number of characters included in the text annotation information, the supplementary annotation information can be determined according to the single-character recognition result corresponding to each character included in the text set in the second recognition result. That is to say, the business platform can determine the characters in the text set that conform to the characters annotated in the text annotation information through the equality of the number of characters. Therefore, the supplementary annotation information can be determined according to the single-character recognition result of each character in the text set, as Figure 4 shown.
[0082] Figure 4 is a schematic diagram for determining the supplementary annotation information of the image to be annotated provided in this specification.
[0083] From Figure 4It can be seen that a line of text in the image to be recognized is: "Absolutely delicious lobster". For the first recognition result, the image area where the recognized text line is located is the rectangular area in the figure. For the second recognition result, it is the image areas where the four characters "absolutely", "delicious", "lobster", and "shrimp" are located in the image to be recognized respectively. It can be clearly seen that if the image areas where each character is located are spliced together to obtain an overall image area, this overall image area can represent the position of the text line in the image more precisely than the rectangular area. Therefore, the business platform can integrate the image areas of each character in the recognized text set in the image to be recognized to obtain an overall image area, and then determine the supplementary annotation information based on this overall image area.
[0084] If the business platform determines that the number of characters included in the reference text line is less than the number of characters included in the text annotation information, it can determine whether the character at the set position in the reference text line is the same as the character at the set position in the text annotation information. Here, the set position can be set according to actual needs. For example, the business platform can determine whether the first character in the reference text line is the same as the first character in the text annotation information, and whether the last character in the reference text line is the same as the last character in the text annotation information. Of course, the set position can also be set according to actual needs. For example, the set position can be the second position or the middle position, etc.
[0085] If they are the same, it may be that some characters in the middle of this text line are not recognized in the second recognition result, and the position of this text line cannot be completely marked through the image areas where each character in the text set is located in the image to be annotated. Therefore, the business platform can determine the supplementary annotation information based on the first recognition result. Among them, the business platform can mark the position of this text line in the image to be annotated according to the recognized image area where this text line is located in the image to be annotated to obtain the supplementary annotation information. Of course, this supplementary annotation information can also include the second recognition result.
[0086] If they are different, the business platform can recognize the other characters in the image to be annotated except the target characters as supplementary characters, and add the supplementary characters to the character set, so as to determine the supplementary annotation information based on the single-character recognition results corresponding to each character except the supplementary characters in the supplementary character set and the single-character recognition result corresponding to the supplementary characters. Consistent with the above method of not supplementing the character set, the business platform can also integrate the image areas of each character in the supplementary character set to obtain an overall image area, and determine the supplementary annotation information based on this overall image area.
[0087] In this specification, there can be multiple ways to determine supplementary text. For example, the service platform can determine the overall text slope of the text line in the image to be annotated, and based on this overall text slope, identify other text in the image to be annotated except for the target text as supplementary text. Among them, the service platform can re-enter the overall text slope and the recognition model to be recognized into the above recognition model, so that the recognition model identifies other text except for the target text as supplementary text. Of course, the service platform can also determine a slope range based on this overall text slope, and identify other text within this slope range among the other text except for the target text as supplementary text.
[0088] After determining the supplementary text in the above manner, the supplementary text can be added to the text set to obtain a supplementary text set after the supplementary text meets other conditions. For example, if the supplementary text is the text included in the text annotation information corresponding to this text line, the supplementary text can be added to the text set. Another example is that if the order of the supplementary text relative to the text in the text line is determined according to the image area where the supplementary text is located in the image to be recognized and is consistent with the order of this supplementary text in the text line in the text annotation information, the supplementary text can be added to the text set.
[0089] It should be noted that if the number of words included in the reference text line is greater than the number of words included in the text annotation information, there is no need to supplement the text annotation information corresponding to this text line, because it can be determined that both the text line and the text set contain more words than the text annotation information. That is to say, when recognizing this text line in the image to be annotated, some words that do not belong to this text line are recognized, and the text set also contains some words that are more than those in the text annotation information. Therefore, the text annotation information corresponding to this text line is not supplemented through the overall recognition result corresponding to this text line and the single-word recognition results corresponding to each word included in the text set.
[0090] In practical applications, if no single character is determined to match the text line, it can be directly determined whether the text line matches the text annotation information. If there is text annotation information that matches the text line, the text annotation information can still be supplemented with the first recognition result. Specifically, if the service platform determines that at least some of the single characters do not contain a single character that matches the text line, it can determine the matching degree between the first recognition result and the text annotation information corresponding to the text line as the third matching degree. If it is determined that the third matching degree is not less than the second set matching degree, the supplementary annotation information can be determined according to the first recognition result. If the service platform determines that the third matching degree is less than the second set matching degree, there is no need to supplement the text annotation information corresponding to the text line. Among them, the second set matching degree can be set according to actual needs.
[0091] It should be noted that the image to be recognized may also contain multiple text lines. Then, the recognition model may also recognize the first recognition results corresponding to multiple text lines, as well as the second recognition results corresponding to each character in multiple text lines. The service platform needs to determine the characters belonging to each text line and the text annotation information corresponding to the text line for each text line, so as to obtain the supplementary annotation information corresponding to the text line. Of course, it may not be possible to annotate the text annotation information corresponding to the text line, and it is also possible that the text line has no corresponding text annotation information, and it is also impossible to annotate the text annotation information corresponding to the text line. However, for an image to be annotated, if most of the text annotation information in the image to be annotated is supplemented, the recognition model can be trained with the image to be annotated.
[0092] It should also be noted that even if the text annotation information is supplemented by the image area where the text line in the first recognition result is located in the image to be annotated, the single-character recognition result corresponding to each character in the second recognition result can also supplement the text annotation information. The second recognition result is used to train the character recognition sub-model that recognizes each character in the recognition model.
[0093] It can be seen from the above method that this method can have partial annotation information corresponding to the image to be annotated. The service platform needs to determine the other annotation information except for the partial annotation information, and annotate the image to be annotated to obtain a training sample. Therefore, the service platform can use the pre-trained recognition model to determine the relevant information of the text line and single characters in the image to be annotated, and directly determine the other annotation information except for the partial annotation information according to the relevant information of the text line and single characters in the image to be annotated, and annotate the image to be annotated. Therefore, compared with the prior art, training samples can be efficiently generated by this method.
[0094] The above is the method for generating training samples provided by one or more embodiments of this specification. Based on the same idea, this specification also provides a corresponding device for generating training samples, as shown in Figure 5 as follows.
[0095] Figure 5 The figure is a schematic diagram of a device for generating training samples provided by this specification, specifically including:
[0096] An acquisition module 501, configured to acquire an image to be annotated and the text annotation information corresponding to the image to be annotated;
[0097] An identification module 502, configured to input the image to be annotated into a preset identification model to obtain an overall identification result for the text line included in the image to be annotated as the first identification result, and a single-character identification result for at least some of the single characters included in the image to be annotated as the second identification result;
[0098] A determination module 503, configured to determine other annotation information for the image to be annotated except the text annotation information as supplementary annotation information according to the first identification result and the second identification result;
[0099] A training module 504, configured to supplement the text annotation information according to the supplementary annotation information to obtain the supplemented annotation information, and generate a training sample corresponding to the image to be annotated through the supplemented annotation information, so as to train the identification model through the training sample.
[0100] Optionally, the first identification result includes: the text content of the text line recognized from the image to be annotated and the image area where the text line is located in the image to be annotated, and the second identification result includes: at least some of the single characters recognized from the image to be annotated and the image area where each of the at least some single characters is located in the image to be annotated.
[0101] Optionally, the determination module 503 is specifically configured to, if it is determined that at least some of the single characters include single characters that match the text line, use the determined single characters that match the text line as target characters and add them to the character set; if it is determined that at least one of the character set and the text line meets a preset condition, determine the supplementary annotation information according to the first identification result and / or the single-character identification results corresponding to each character included in the character set in the second identification result.
[0102] Optionally, the determining module 503 is specifically configured to, for each of the at least part of the single characters, determine the degree of overlap between the image region where the character is located in the image to be annotated and the image region where the text line is located in the image to be annotated; if it is determined that the degree of overlap is not less than a set degree of overlap, determine that the character matches the text line, and use the character as the target character.
[0103] Optionally, the determining module 503 is specifically configured to sort each character included in the character set in the order of the characters in the text line to obtain a reference text line; determine the matching degree between the text line and the text annotation information as a first matching degree, and determine the matching degree between the reference text line and the text annotation information as a second matching degree; if it is determined that at least one of the first matching degree and the second matching degree is not less than a first set matching degree, determine that at least one of the character set and the text line meets the preset condition.
[0104] Optionally, the determining module 503 is specifically configured to, if it is determined that the number of characters included in the reference text line is less than the number of characters included in the text annotation information, determine whether the character at a set position in the reference text line is the same as the character at the set position in the text annotation information; if so, determine the supplementary annotation information according to the first recognition result; if not, recognize other characters in the image to be annotated except the target character as supplementary characters, and add the supplementary characters to the character set, so as to determine the supplementary annotation information according to the single-character recognition results corresponding to each character except the supplementary characters in the supplementary character set and the single-character recognition result corresponding to the supplementary characters.
[0105] Optionally, the determining module 503 is specifically configured to determine the overall text slope of the text line in the image to be annotated; according to the overall text slope, recognize other characters in the image to be annotated except the target character as supplementary characters.
[0106] Optionally, the determining module 503 is specifically configured to, if it is determined that the number of characters included in the reference text line is equal to the number of characters included in the text annotation information, determine the supplementary annotation information according to the single-character recognition results corresponding to each character included in the character set.
[0107] Optionally, the determining module 503 is further configured to, if it is determined that the number of characters included in the reference text line is greater than the number of characters included in the text annotation information, not supplement the text annotation information corresponding to the text line.
[0108] Optionally, if the determination module 503 determines that neither the set of characters nor the text line satisfies the preset condition, the determination module 503 is further configured not to supplement the text annotation information corresponding to the text line.
[0109] Optionally, the determination module 503 is specifically configured to, if it is determined that none of the at least part of the single characters contains a single character that matches the text line, determine the matching degree between the first recognition result and the text annotation information corresponding to the text line as a third matching degree; if it is determined that the third matching degree is not less than the second set matching degree, determine the supplementary annotation information according to the first recognition result.
[0110] Optionally, the determination module 503 is further configured to, if it is determined that the third matching degree is less than the second set matching degree, not supplement the text annotation information corresponding to the text line.
[0111] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the above Figure 1 shown method for generating training samples.
[0112] This specification also provides Figure 6 a schematic structural diagram of the electronic device shown. As Figure 6 described above, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 shown method for generating training samples. Of course, in addition to the software implementation manner, this specification does not exclude other implementation manners, such as a logic device or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and may also be a hardware or a logic device.
[0113] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is an integrated circuit whose logic function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, today, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compilers used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL), and there is not just one type of HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.
[0114] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an Application Specific Integrated Circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, ASICs, programmable logic controllers, and embedded microcontrollers to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.
[0115] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0116] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0117] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0118] The present invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowchart illustrations and / or block diagrams, and combinations of flows and / or blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing apparatus create means for implementing the functions specified in the flowchart Figure 1 for one or more of the flows and / or blocks Figure 1 and / or means for implementing the functions specified in one or more of the blocks.
[0119] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions specified in the flowchart Figure 1 for one or more of the flows and / or blocks Figure 1 and / or means for implementing the functions specified in one or more of the blocks.
[0120] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart Figure 1 for one or more of the flows and / or blocks Figure 1 and / or means for implementing the functions specified in one or more of the blocks.
[0121] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0122] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory. The memory is an example of computer-readable media.
[0123] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0124] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0125] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0126] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0127] The various embodiments in this specification are described in a progressive manner. For the parts that are the same or similar among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.
[0128] The above description is only for the embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various modifications and changes can be made to this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.
Claims
1. A method for generating training samples, characterized in that, Including: Obtain the image to be annotated and the corresponding text annotation information of the image to be annotated; Input the image to be annotated into a preset recognition model to obtain the overall recognition result of the text lines included in the image to be annotated as the first recognition result, and the single-character recognition result of at least some of the single characters included in the image to be annotated as the second recognition result; determine, according to the first recognition result and the second recognition result, other annotation information for the image to be annotated except the text annotation information as supplementary annotation information, specifically including: If it is determined that at least some of the single characters include single characters that match the text line, determine the determined single characters that match the text line as target characters and add them to the character set; Sort each character included in the character set according to the character order in the text line to obtain a reference text line; Determine the matching degree between the text line and the text annotation information as the first matching degree, and determine the matching degree between the reference text line and the text annotation information as the second matching degree; If it is determined that at least one of the first matching degree and the second matching degree is not less than the first set matching degree, determine that at least one of the character set and the text line meets the preset conditions; if it is determined that the number of characters included in the reference text line is less than the number of characters included in the text annotation information, determine whether the character at the set position in the reference text line is the same as the character at the set position in the text annotation information; If so, determine the supplementary annotation information according to the first recognition result; If not, identify other characters except the target characters from the image to be annotated as supplementary characters, and add the supplementary characters to the character set, so as to determine the supplementary annotation information according to the single-character recognition results corresponding to each character except the supplementary characters in the character set and the single-character recognition result corresponding to the supplementary characters; Supplement the text annotation information according to the supplementary annotation information to obtain the supplemented annotation information, and generate a training sample corresponding to the image to be annotated through the supplemented annotation information, so as to train the recognition model through the training sample; The first recognition result includes: the text content of the text line recognized from the image to be annotated and the image area where the text line is located in the image to be annotated, and the second recognition result includes: at least some of the single characters recognized from the image to be annotated and the image area where each of the at least some single characters is located in the image to be annotated.
2. The method according to claim 1, wherein Determine the single characters that match the text line as target characters, specifically including: For each of the at least some single characters, determine the overlap degree between the image area where the character is located in the image to be annotated and the image area where the text line is located in the image to be annotated; If it is determined that the degree of coincidence is not less than the set degree of coincidence, it is determined that the text matches the text line, and the text is used as the target text.
3. The method according to claim 1, wherein Identify other texts in the to-be-annotated image except the target text as supplementary texts, specifically including: Determine the overall text slope of the text line in the to-be-annotated image; According to the overall text slope, identify other texts in the to-be-annotated image except the target text as supplementary texts.
4. The method according to claim 1, wherein Determine the supplementary annotation information according to the first recognition result and / or the single-character recognition results corresponding to each text included in the text set in the second recognition result, specifically including: If it is determined that the number of texts included in the reference text line is equal to the number of texts included in the text annotation information, determine the supplementary annotation information according to the single-character recognition results corresponding to each text included in the text set in the second recognition result.
5. The method according to claim 1, wherein The method further includes: If it is determined that the number of texts included in the reference text line is greater than the number of texts included in the text annotation information, do not supplement the text annotation information corresponding to the text line.
6. The method according to claim 1, characterized in that, If it is determined that both the text set and the text line do not meet the preset conditions, the method further includes: Do not supplement the text annotation information corresponding to the text line.
7. The method according to claim 1, wherein Determine other annotation information for the to-be-annotated image except the text annotation information as supplementary annotation information according to the first recognition result and the second recognition result, specifically including: If it is determined that at least some of the single characters do not contain a single character that matches the text line, determine the matching degree between the first recognition result and the text annotation information corresponding to the text line as the third matching degree; If it is determined that the third matching degree is not less than the second set matching degree, determine the supplementary annotation information according to the first recognition result.
8. The method according to claim 7, wherein The method further includes: If it is determined that the third matching degree is less than the second set matching degree, do not supplement the text annotation information corresponding to the text line.
9. A generating device for training samples, characterized in that, Including: An acquisition module for acquiring a to-be-annotated image and the text annotation information corresponding to the to-be-annotated image; An identification module for inputting the to-be-annotated image into a preset identification model to obtain an overall recognition result for the text line included in the to-be-annotated image as the first recognition result, and a single-character recognition result for at least some of the single characters included in the to-be-annotated image as the second recognition result; A determination module for determining other annotation information for the to-be-annotated image except the text annotation information as supplementary annotation information according to the first recognition result and the second recognition result, specifically including: If it is determined that at least some of the single characters contain a single character that matches the text line, use the determined single character that matches the text line as the target text and add it to the text set; Sort each text included in the text set according to the text order in the text line to obtain a reference text line; Determine the matching degree between the text line and the text annotation information as the first matching degree, and determine the matching degree between the reference text line and the text annotation information as the second matching degree; If it is determined that at least one of the first matching degree and the second matching degree is not less than the first set matching degree, determine that at least one of the character set and the text line meets the preset condition; if it is determined that the number of characters included in the reference text line is less than the number of characters included in the text annotation information, determine whether the character at the set position in the reference text line is the same as the character at the set position in the text annotation information; If so, determine the supplementary annotation information according to the first recognition result; If not, identify other characters except the target character from the image to be annotated as supplementary characters, and add the supplementary characters to the character set, so as to determine the supplementary annotation information according to the single-character recognition result corresponding to each character except the supplementary character in the character set and the single-character recognition result corresponding to the supplementary character; A training module, configured to supplement the text annotation information according to the supplementary annotation information to obtain the supplemented annotation information, and generate a training sample corresponding to the image to be annotated through the supplemented annotation information, so as to train the recognition model through the training sample; The first recognition result includes: the text content of the text line recognized from the image to be annotated and the image area where the text line is located in the image to be annotated, and the second recognition result includes: at least some single characters recognized from the image to be annotated and the image area where each character in the at least some single characters is located in the image to be annotated.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 8 above is implemented.
11. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the method described in any one of claims 1 to 8 above is implemented.
Citation Information
Patent Citations
Training data generation method and device and model training method and device
CN109978044A
Training data processing method, device and equipment and computer readable storage medium
CN110321788A