Text recognition method, model training method and device
By analyzing the category attributes of the initial text in OCR recognition technology and performing error correction, the problem of text recognition errors has been solved, achieving higher accuracy and reliability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING BAIDU NETCOM SCI & TECH CO LTD
- Filing Date
- 2022-03-01
- Publication Date
- 2026-04-24
AI Technical Summary
Existing OCR recognition technology is prone to text recognition errors during the text recognition process, resulting in low accuracy of the text content.
By analyzing and processing the initial text, determining its category attributes, correcting erroneous text, generating accurate text content, and using a text position discriminator and a masked language recall model for error correction.
It improves the accuracy and reliability of text recognition and avoids the drawbacks of OCR recognition technology, such as text errors and repetitions.
Smart Images

Figure CN114663886B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to image processing, deep learning, and natural language understanding in artificial intelligence technologies, and particularly to a text recognition method, a model training method, and an apparatus. Background Technology
[0002] Optical Character Recognition (OCR) refers to the process by which electronic devices (such as scanners or digital cameras) examine characters printed on paper, determine their shapes by detecting dark and light patterns, and then translate the shapes into computer text using character recognition methods.
[0003] In existing technologies, OCR recognition technology is typically used to obtain the text content in the image to be recognized.
[0004] However, OCR recognition technology may have text recognition errors, resulting in low accuracy of the obtained text content. Summary of the Invention
[0005] This disclosure provides a text recognition method, a model training method, and an apparatus for improving the accuracy of text recognition.
[0006] According to a first aspect of this disclosure, a text recognition method is provided, comprising:
[0007] Optical character recognition is performed on the acquired image to be recognized to obtain the initial text of the image to be recognized;
[0008] The initial text is analyzed and processed to obtain the category attribute of the initial text. If the category attribute of the initial text indicates that the initial text is incorrect, then the incorrect text is corrected to obtain the correct text used to correct the incorrect text.
[0009] Based on the initial text and the correct text, the text content of the image to be recognized is generated.
[0010] According to a second aspect of this disclosure, a method for training a model is provided, comprising:
[0011] Obtain a first sample dataset, wherein the first sample dataset includes initial point-of-interest (POI) name text and variant POI name text obtained by modifying the initial POI name text, wherein the variant POI name text includes at least one incorrect character.
[0012] The initial model parameters are obtained by training based on the first sample dataset, and the text position discriminator is obtained by training based on the initial model parameters. The text position discriminator is used to analyze and process the initial text of the image to be recognized to obtain the category attribute of the initial text.
[0013] According to a third aspect of this disclosure, a text recognition device is provided, comprising:
[0014] The recognition unit is used to perform optical character recognition on the acquired image to be recognized to obtain the initial text of the image to be recognized;
[0015] An analysis unit is used to analyze and process the initial text to obtain the category attribute of the initial text;
[0016] The error correction unit is used to perform error correction processing on the initial text if the category attribute of the initial text indicates that the initial text is an incorrect text, so as to obtain the correct text used to correct the incorrect text.
[0017] The generation unit is used to generate the text content of the image to be recognized based on the initial text and the correct text.
[0018] According to a fourth aspect of this disclosure, a model training apparatus is provided, comprising:
[0019] An acquisition unit is configured to acquire a first sample dataset, wherein the first sample dataset includes an initial point of interest (POI) name text and a variant POI name text obtained by modifying the initial POI name text, wherein the variant POI name text includes at least one incorrect character.
[0020] The first training unit is used to train the initial model parameters based on the first sample dataset.
[0021] The second training unit is used to train a text position discriminator based on the initial model parameters. The text position discriminator is used to analyze and process the initial text of the image to be recognized to obtain the category attribute of the initial text.
[0022] According to a fifth aspect of this disclosure, an electronic device is provided, comprising:
[0023] At least one processor; and
[0024] A memory communicatively connected to the at least one processor; wherein,
[0025] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect or the second aspect.
[0026] According to a sixth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the method according to the first or second aspect.
[0027] According to a seventh aspect of this disclosure, a computer program product is provided, the computer program product comprising: a computer program stored in a readable storage medium, at least one processor of an electronic device being able to read the computer program from the readable storage medium, the at least one processor executing the computer program causing the electronic device to perform the method described in the first aspect or the second aspect.
[0028] This embodiment provides a text recognition method, a model training method, and an apparatus. By determining the initial text category attribute, when the initial text category attribute indicates that the initial text is an erroneous text, the correct text used to correct the error is determined. This allows for the determination of the text content of the image to be recognized by combining the correct text with the technical features of the correct text. This avoids the drawbacks of text errors caused by OCR recognition technology and improves the accuracy and reliability of text recognition.
[0029] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0030] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0031] Figure 1 This is a schematic diagram based on the first embodiment of the present disclosure;
[0032] Figure 2 This is a schematic diagram according to the second embodiment of the present disclosure;
[0033] Figure 3 This is a schematic diagram illustrating the principle of a text recognition method according to an embodiment of the present disclosure;
[0034] Figure 4 This is a schematic diagram according to the third embodiment of the present disclosure;
[0035] Figure 5 This is a schematic diagram according to the fourth embodiment of the present disclosure;
[0036] Figure 6 This is a schematic diagram according to the fifth embodiment of the present disclosure;
[0037] Figure 7 This is a schematic diagram according to the sixth embodiment of the present disclosure;
[0038] Figure 8 This is a schematic diagram according to the seventh embodiment of the present disclosure;
[0039] Figure 9 This is a schematic diagram according to the eighth embodiment of the present disclosure;
[0040] Figure 10 This is a schematic diagram according to the ninth embodiment of the present disclosure;
[0041] Figure 11 This is a block diagram of an electronic device used to implement the text recognition method and model training method of the embodiments of this disclosure. Detailed Implementation
[0042] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0043] With the development of artificial intelligence (AI) technology, the method of manually recognizing text content in images has been replaced by artificial intelligence technology. For example, OCR recognition technology can be used to recognize images and obtain the text content in the images.
[0044] It is understandable that images can be classified based on different dimensions. For example, images can be classified based on how they are formed, such as dividing them into pictures and photographs. Images can also be classified based on their content, such as dividing them into document images (e.g., images of checks, tax stamps, etc.) and sign images (e.g., images of restaurant signs, warning signs, etc.).
[0045] Because OCR recognition technology is limited by the deep learning model itself and by the quality of the image, when the text content in the image is obtained based on OCR recognition technology, there may be missing or extra characters in the text content, or there may be typos in the text content.
[0046] To avoid at least one of the aforementioned technical problems, the inventors of this disclosure, through creative labor, arrived at the inventive concept of this disclosure: determining the initial text in an image based on OCR recognition technology, analyzing and processing the initial text, correcting the erroneous text when it exists, thereby obtaining the correct text, and determining the text content of the image based on the initial text and the correct text.
[0047] Based on the above inventive concept, this disclosure provides a text recognition method, a model training method, and an apparatus, which are applied to image processing, deep learning, and natural language understanding in artificial intelligence technology to improve the accuracy of text recognition.
[0048] Figure 1 This is a schematic diagram based on the first embodiment of the present disclosure, as shown below. Figure 1 As shown, the text recognition method of this disclosure includes:
[0049] S101: Perform optical character recognition on the acquired image to be recognized to obtain the initial text of the image to be recognized.
[0050] For example, the execution subject of this embodiment can be a text recognition device, which can be a server (such as a local server, a cloud server, or a server cluster, etc.), a terminal device, a processor, or a chip, etc. This embodiment does not limit the scope of the application.
[0051] In this context, "to be identified" in the image to be identified is used to distinguish it from other images, such as the sample images mentioned later, and should not be interpreted as a limitation on the image to be identified. Furthermore, the image to be identified can be understood as the image that needs to be identified.
[0052] Based on the above analysis, it can be seen that there are various types of images, such as ticket images and sign images. Accordingly, in this embodiment, the image to be identified can be either a ticket image or a sign image.
[0053] Similarly, the image can be a picture or a photograph. Accordingly, in this embodiment, the image to be identified can be a picture or a photograph.
[0054] For example, the image to be identified can be a picture of a ticket, a photograph of a ticket, a photograph of a signboard, or a picture of a signboard, etc., and so on.
[0055] This step can be understood as follows: After obtaining the image to be recognized, OCR recognition can be performed on the image to obtain the text in the image. In order to distinguish the text obtained by OCR recognition from the text in the following text (such as the correct text), the text obtained based on OCR recognition is called the initial text.
[0056] It is worth noting that this embodiment does not limit the method of acquiring the image to be recognized. For example, the image to be recognized can be acquired using the following example:
[0057] In one example, the text recognition device can be connected to an image acquisition device and receive the image to be recognized sent by the image acquisition device.
[0058] Among them, the image acquisition device refers to a device that can acquire an image to be recognized, such as a camera.
[0059] In another example, the text recognition device can provide a tool for loading images, which the user can use to transfer the image to be recognized to the text recognition device.
[0060] The tool for loading images can be an interface for connecting to external devices, such as an interface for connecting to other storage devices, through which the image to be recognized transmitted by the external device can be obtained; the tool for loading images can also be a display device, such as a text recognition device that can input an interface for loading images on the display device, through which the user can import the image to be recognized into the text recognition device, and the text recognition device can obtain the imported image to be recognized.
[0061] S102: Analyze and process the initial text to obtain the category attribute of the initial text. If the category attribute of the initial text indicates that the initial text is incorrect, then perform error correction processing on the incorrect text to obtain the correct text used to correct the error.
[0062] For example, the category attribute of the initial text can be used to characterize whether the initial text is erroneous, such as the category attribute of the initial text indicating that the initial text is erroneous, or the category attribute of the initial text indicating that the initial text is not erroneous.
[0063] This embodiment does not limit the analysis and processing method. For example, a network model can be used to analyze and process the initial text to obtain the category attribute of the initial text, and the network model can be a classification network model. That is, the analysis and processing can be classification processing. For example, the initial text can be classified based on a classification network model to determine whether the initial text is incorrect (i.e., to obtain the category attribute of the initial text).
[0064] For example, the initial text can be analyzed in conjunction with its context to obtain its category attribute. This analysis could be semantic analysis, where the initial text is analyzed based on the semantics of its context to determine if it is incorrect (i.e., to obtain its category attribute).
[0065] Similarly, this embodiment does not limit the method of error correction.
[0066] S103: Generate the text content of the image to be recognized based on the initial text and the correct text.
[0067] Based on the above analysis, the initial text contains erroneous characters, and the correct text is the text that corrects the erroneous characters. Therefore, when generating the text content of the image to be recognized based on the initial text and the correct text, the generated text content of the image to be recognized can have high accuracy and reliability.
[0068] Based on the above analysis, this disclosure provides a text recognition method, including: performing optical character recognition on an acquired image to be recognized to obtain initial text in the image to be recognized; analyzing and processing the initial text to obtain the category attribute of the initial text; if the category attribute of the initial text indicates that the initial text is incorrect, then performing error correction processing on the incorrect text to obtain correct text for correcting the error; and generating the text content of the image to be recognized based on the initial text and the correct text. In this embodiment, by determining the category attribute of the initial text, when the category attribute of the initial text indicates that the initial text is incorrect, the correct text for correcting the error is determined, so as to combine the correct text to determine the technical feature of the text content of the image to be recognized, thereby avoiding the drawbacks of text errors caused by OCR recognition technology and improving the technical effect of text recognition accuracy and reliability.
[0069] Figure 2 This is a schematic diagram based on the second embodiment of the present disclosure, as shown below. Figure 2 As shown, the text recognition method of this disclosure includes:
[0070] S201: Perform optical character recognition on the acquired image to be recognized to obtain the initial text of the image to be recognized.
[0071] The initial number of characters is multiple.
[0072] It should be understood that, in order to avoid redundant descriptions, the technical features that are the same as those in the above embodiments will not be repeated in this embodiment.
[0073] S202: Analyze and process each initial character one by one to obtain the category attribute corresponding to each initial character.
[0074] Exemplarily, as Figure 3 shown, if the image to be recognized is a signboard image, after performing OCR recognition on the signboard image, the initial characters of the signboard image, namely "Hailu Happy Hot Pot", are obtained. That is, the number of initial characters obtained is multiple, specifically seven.
[0075] Correspondingly, analyze and process the seven initial characters one by one. For example, analyze and process the character "Hai" to obtain the category attribute of the character "Hai"; then analyze and process the character "Lu" to obtain the category attribute of the character "Lu"; and so on, until analyzing and processing the character "Pot" to obtain the category attribute of the character "Pot".
[0076] Among them, the category attribute of the character "Hai" can represent whether "Hai" is a correct character; the category attribute of the character "Lu" can represent whether "Lu" is a correct character; and so on, the category attribute of the character "Pot" can represent whether "Pot" is a correct character.
[0077] In this embodiment, when the number of initial characters is multiple, analyze and process each initial character one by one to determine the category attribute corresponding to each initial character, that is, determine whether each initial character is a correct character one by one, so as to analyze and process all the initial characters, achieving the technical effects of comprehensiveness and integrity in the analysis and processing.
[0078] In some embodiments, each initial character has a position attribute. Correspondingly, S202 may include: determining the category attribute corresponding to each initial character in sequence according to the position attribute corresponding to each initial character.
[0079] Among them, the position attribute may be coordinate information, such as pixel coordinates. That is, each initial character has pixel coordinates. According to the pixel coordinates corresponding to each initial character, the sequential relationship of each initial character based on the image coordinate system on the signboard image can be determined, and based on this sequential relationship, each initial character is analyzed and processed in sequence, so as to obtain the category attribute corresponding to each initial character.
[0080] Exemplarily, combining the above analysis and Figure 3It can be seen that the character "Hai" has pixel coordinates, and the character "Lu" also has pixel coordinates. According to the pixel coordinates of the character "Hai" and the pixel coordinates of the character "Lu", the order relationship of the characters "Hai" and "Lu" in the sign image can be determined. That is, if the character "Hai" is before the character "Lu", then the character "Hai" is analyzed and processed first to obtain the category attribute of the character "Hai", and then the character "Lu" is analyzed and processed to obtain the category attribute of the character "Lu", and so on. Here, they are not listed one by one.
[0081] It is worth noting that by combining the position attribute of the initial text to determine the category attribute of the initial text, since the position attribute is unique, the category attribute of the initial text determined based on the position attribute has the technical effects of high accuracy and reliability.
[0082] For example, when "Hailu Happy Hot Pot" includes two "Le" characters, if the two "Le" characters have the same position attribute, such as the same pixel coordinates, it means that there is a drawback of repeated recognition in the OCR recognition result. Then, it can be determined that the category attribute of the "Le" character indicates that the "Le" character is a redundant character. In order to improve the accuracy of text recognition, one of the "Le" characters can be removed.
[0083] In some embodiments, the type attribute of the initial text can be determined in combination with a network model. For example, a text position discriminator can be pre-trained to perform discriminant processing on the initial text based on the text position discriminator to obtain the category attribute of the initial text.
[0084] Among them, the text position discriminator is trained based on initialized model parameters, and the initialized model parameters are trained based on a first sample dataset. The first sample dataset includes the initial point of interest name text and the variant point of interest name text obtained by modifying the initial point of interest name text. The variant point of interest name text includes at least one incorrect character.
[0085] Similarly, the "first" in the first sample dataset is used to distinguish the first sample dataset from other sample datasets, such as distinguishing the first sample dataset from the second sample dataset in the following text, and cannot be understood as a limitation on the first sample dataset.
[0086] Exemplarily, an initial point of interest (POI) name text can be obtained, and the initial point of interest name text can be modified. For example, one or more characters in the initial point of interest name text are modified into incorrect characters to obtain a variant point of interest name text, and a first sample dataset is constructed based on the initial point of interest name text and the variant point of interest name text. [[ID=二十]]
[0087] In the first sample dataset, both the initial point of interest (POI) name text and the variant POI name text can be referred to as sample data. That is, an initial POI name text is a sample data, and a variant POI name text is also a sample data.
[0088] The number of sample data in the first sample dataset can be determined based on requirements, historical records, and experiments, and this embodiment does not impose any limitations. Furthermore, this embodiment does not limit the number of variant point-of-interest (POI) texts that can be modified from an initial POI text. Also, this embodiment does not limit the number of characters modified from an initial POI text to obtain variant POI texts.
[0089] In this embodiment, the text corresponding to the name of a point of interest can be understood as the text that is being followed. The text position discriminator is trained using the text corresponding to the point of interest name. Compared to training based on everyday conversational corpora or internet text corpora, this allows for more targeted training, resulting in a text position discriminator with higher accuracy and reliability.
[0090] Especially when the image to be identified is a signboard image, the text in the signboard image that needs to be identified is the text corresponding to the name of the signboard image. The method based on this embodiment can determine the text content of the signboard image with strong targeting and reliability.
[0091] In this embodiment, by using a text position discriminator to determine the category attribute of the initial text, the efficiency of determining the category attribute of the initial text can be improved. Furthermore, since the text position discriminator is trained based on the initial point of interest name text and the variant point of interest name text, the text position discriminator can have high accuracy. That is, when using the text position discriminator to determine the category attribute of the initial text, the technical effect of improving the accuracy of determining the category attribute of the initial text can also be achieved.
[0092] For example, combining the above analysis and Figure 3 When performing OCR recognition on the signboard image to obtain "Hailu Kuailele Hotpot", "Hailu Kuailele Hotpot" can be input into the text position discriminator to output the corresponding category attributes of "Hailu Kuailele Hotpot".
[0093] In some embodiments, different category attributes can be represented by different flag bits: flag bit W represents wrong, used to indicate incorrect text; flag bit D represents duplicate, used to indicate redundant text; and flag bit R represents right, used to indicate correct text.
[0094] like Figure 3As shown, the respective category attributes corresponding to the output "Seaside Happy Hot Pot" are RWRRDRR. That is, the flag bits of the characters "sea", "fast", the first "happy", "fire", and "pot" are all R, indicating that they are all correct characters; the flag bit of the character "road" is W, indicating that the character "road" is an incorrect character; the flag bit of the second "happy" is D, indicating that the second "happy" is an extra character.
[0095] Combined with the above analysis, it can be seen that each initial character has a position attribute. Correspondingly, the flag bit can be the flag bit of the initial character representing the position attribute.
[0096] S203: If the category attribute of the initial character indicates that the initial character is an extra character, then perform a deletion process on the extra character.
[0097] Exemplarily, combined with the above analysis, if the second "happy" is an extra character, then perform a deletion process on the second "happy" being an extra character, that is, obtain "Seaside Happy Hot Pot".
[0098] In this embodiment, when the category attribute of the initial character indicates that the initial character is an extra character, perform a deletion process on the extra character to avoid duplicate characters, thereby improving the technical effects of the accuracy and reliability of text recognition.
[0099] S204: If the category attribute of the initial character indicates that the initial character is an incorrect character, then perform a mask (mask) process on the incorrect character in the initial character.
[0100] Exemplarily, combined with the above analysis, it can be seen that the character "road" is an incorrect character. Correspondingly, perform a mask process on the character "road", as Figure 3 shown.
[0101] S205: Perform a prediction on the initial character after the mask process to obtain a candidate set, and obtain the correct character from the candidate set.
[0102] Among them, the candidate set includes error-correction characters for replacing the incorrect character.
[0103] Exemplarily, combined with the above analysis and Figure 3 , when performing a mask process on the character "road", the initial character after the mask process is "Sea [mask] Happy Hot Pot". Perform a prediction on "[mask]" based on "Sea Happy Hot Pot", that is, perform a prediction on the character "road", and obtain a candidate set including error-correction characters for replacing the character "road".
[0104] That is to say, the candidate set includes one or more error-correction words. If there is one error-correction word, the error-correction word can be determined as the correct word. If there are multiple error-correction words, one error-correction word can be obtained from the multiple error-correction words, and the obtained error-correction word can be determined as the correct word, so as to replace the wrong word "road" based on the correct word.
[0105] In this embodiment, by combining "mask processing + prediction" to obtain a candidate set and obtaining the correct word for correcting the wrong word based on the candidate set, the technical effects of improving the accuracy and reliability of text recognition can be achieved.
[0106] In some embodiments, a candidate set can be obtained by combining a network model. For example, a masked language model (MLM) is pre-trained, and the initial text after mask processing is input into the masked language model to output the candidate set.
[0107] Exemplarily, combining the above embodiments and Figure 3 input the initial text "Sea
mask
[0108] Among them, the masked language model is trained and generated based on a second sample data set, and the second sample data set includes sample point-of-interest name texts.
[0109] Similarly, in this embodiment, by combining a network model to obtain a candidate set, the technical effects of improving the efficiency and accuracy of the candidate set can be achieved.
[0110] The sample data in the second sample data set may be the same as or different from the sample data in the first sample data set, and this embodiment does not make a limitation.
[0111] In some embodiments, training the masked language model includes the following steps:
[0112] The first step: Obtain a second sample data set, where the second sample data set includes sample point-of-interest name texts.
[0113] Similarly, in the embodiment, the number of sample point-of-interest name texts can be determined based on requirements, historical records, and experiments, etc., and this embodiment does not make a limitation.
[0114] The second step: Perform mask processing on any word in the sample point-of-interest name text to obtain the sample point-of-interest name text after mask processing.
[0115] The third step: Based on the preset basic network model, predict the masked text in the sample interest point name text after masking to obtain the predicted text.
[0116] Step 4: Calculate the loss value of the predicted text and the labeled text (i.e., the pre-labeled real text) using MLM. loss The parameters of the basic network model are adjusted based on the loss value to train the training mask language recall model.
[0117] If the text to be masked is y, then the loss value MLM can be calculated using Equation 1. loss Formula 1:
[0118] MLM loss =-log(P mask_yi )
[0119] Where mask_yi is the predicted (softmax) probability of the actual text y.
[0120] It is worth noting that in this embodiment, the sample interest point name text is used to train the masked language recall model, which can avoid the tedious calculation of loss values, thereby improving the technical effect of training efficiency and reliability.
[0121] In some embodiments, the base network model can be transformers, which include an encoder. Since the masked language recall model is trained by sampling interest point name samples, the computation of loss values is relatively reduced. Therefore, the encoder in the transformers can have a six-layer structure to obtain the optimal combination of model parameters and inference performance.
[0122] In some embodiments, the number of corrected texts is multiple, and obtaining the correct texts from the candidate set includes the following steps:
[0123] The first step is to obtain the font structure attributes of the erroneous text and the font structure attributes of each corrected text.
[0124] Among them, the font structure attribute is used to characterize the stroke content and / or stroke order of the text.
[0125] The second step is to determine the correct text from among the corrected texts based on the font structure attributes of the incorrect texts and the font structure attributes of each corrected text.
[0126] Exemplarily, in combination with the above analysis, if the incorrect character is "路" and the corrected characters are "鲜, 洋, 盗, 南", then obtain the font structure attribute of the character "路", and obtain the respective font structure attributes of the characters "鲜, 洋, 盗, 南" to determine the correct character from "鲜, 洋, 盗, 南".
[0127] For example, the stroke content of the character "路" can be obtained, and the stroke content of the character "鲜", the stroke content of the character "洋", the stroke content of the character "盗", and the stroke content of the character "南" can be obtained, so as to determine the correct character from "鲜, 洋, 盗, 南" according to the respective stroke contents of "路, 鲜, 洋, 盗, 南".
[0128] Another example, the stroke order of the character "路" can be obtained, and the stroke order of the character "鲜", the stroke order of the character "洋", the stroke order of the character "盗", and the stroke order of the character "南" can be obtained, so as to determine the correct character from "鲜, 洋, 盗, 南" according to the respective stroke orders of "路, 鲜, 洋, 盗, 南".
[0129] Furthermore, the stroke content and stroke order of the character "路" can be obtained, and the stroke content and stroke order of the character "鲜", the stroke content and stroke order of the character "洋", the stroke content and stroke order of the character "盗", and the stroke content and stroke order of the character "南" can be obtained, so as to determine the correct character from "鲜, 洋, 盗, 南" according to the respective stroke content and stroke order of "路, 鲜, 洋, 盗, 南".
[0130] It should be noted that different characters have different font structure attributes. In this embodiment, by combining the font structure attributes to determine the correct character, the technical effect of improving the accuracy and reliability of the determined correct character can be achieved.
[0131] In some embodiments, the second step may include the following sub-steps:
[0132] The first sub-step: For the font structure attribute of each corrected character, calculate the similarity between the font structure attribute of the corrected character and the font structure attribute of the incorrect character.
[0133] The second sub-step: Determine the correct character from each corrected character according to the similarities.
[0134] Exemplarily, in combination with the above analysis, calculate the similarity between the font structure attributes of the character "路" and the font structure attributes of the character "鲜" (for the convenience of distinction, this similarity is referred to as the first similarity); calculate the similarity between the font structure attributes of the character "路" and the font structure attributes of the character "洋" (similarly, for the convenience of distinction, this similarity is referred to as the second similarity); calculate the similarity between the font structure attributes of the character "路" and the font structure attributes of the character "盗" (similarly, for the convenience of distinction, this similarity is referred to as the third similarity); calculate the similarity between the font structure attributes of the character "路" and the font structure attributes of the character "南" (similarly, for the convenience of distinction, this similarity is referred to as the fourth similarity); and determine the correct character according to the first similarity, the second similarity, the third similarity, and the fourth similarity.
[0135] Among them, the calculation method can adopt the shortest edit distance algorithm.
[0136] In this embodiment, by calculating the similarity based on the font structure attributes and determining the correct character based on the similarity, it is possible to relatively appropriately represent the similarity between two characters through the similarity, so that the correct character determined based on the similarity has the technical effects of high reliability and accuracy.
[0137] In some embodiments, the second sub-step may include the following refinement steps:
[0138] The first refinement step: Determine the maximum similarity from each similarity.
[0139] The second refinement step: Extract the error correction character corresponding to the maximum similarity from the candidate set, and determine the error correction character corresponding to the maximum similarity as the correct character.
[0140] Exemplarily, in combination with the above analysis, after calculating the first similarity, the second similarity, the third similarity, and the fourth similarity, the maximum similarity can be determined from the first similarity, the second similarity, the third similarity, and the fourth similarity, so as to determine the correct character.
[0141] For example, if the third similarity is the maximum among the four similarities, it means that the similarity between the character "盗" and the character "路" is greater, then the character "盗" is determined as the correct character.
[0142] S206: Generate the text content of the to-be-recognized image according to the initial text with redundant text removed and the correct text.
[0143] Exemplarily, in combination with the above analysis and Figure 3The initial text after removing redundant characters is "Sea Route Happy Hot Pot", the incorrect character is "Road", and the correct character is "Thief". Therefore, the text content of the image to be identified is "Pirate Happy Hot Pot".
[0144] In other words, the correct text can replace the incorrect text in the initial text to obtain the text content of the image to be recognized, thus avoiding the drawbacks of text errors or repetitions caused by OCR recognition, thereby improving the accuracy and reliability of text recognition.
[0145] Based on the above analysis, it can be seen that the OCR recognition model, text position discriminator, and masked language recall model can be pre-trained separately and combined with the calculation module for determining similarity to recognize the image to obtain the text content of the image. The modules (i.e., OCR recognition model, text position discriminator, masked language recall model, and calculation module) are decoupled from each other, and recognition failures (such as recognition errors) can be traced, so that the text recognition has a high degree of accuracy and reliability.
[0146] Furthermore, the text position discriminator, masked language recall model, and calculation module can be integrated as a whole, such as an error correction module, to correct the recognition results of the OCR recognition model, thereby obtaining the accurate text content of the image to be recognized.
[0147] Figure 4 This is a schematic diagram based on the third embodiment of the present disclosure, as shown below. Figure 4 As shown, the training method for the model in this embodiment includes:
[0148] S401: Obtain the first sample dataset.
[0149] The first sample dataset includes the initial point of interest (POI) name text and variant POI name text obtained by modifying the initial POI name text. The variant POI name text contains at least one incorrect character.
[0150] For example, the execution entity in this embodiment can be a model training device, which can be a server (such as a local server, a cloud server, or a server cluster), a terminal device, a processor, or a chip, etc. This embodiment does not limit the scope of the device.
[0151] Based on the above analysis, it can be seen that the training device for the model can be the same as the text recognition device or a different device; this embodiment does not impose any limitations.
[0152] If the training device of the model and the text recognition device are different devices, after the training device of the model trains and obtains the character position discriminator, the character position discriminator can be transmitted to the text recognition device, so that the text recognition device can deploy the character position discriminator and recognize the image to be recognized to obtain the text content of the image to be recognized.
[0153] Alternatively, after the training device of the model trains and obtains the character position discriminator, when the text recognition device needs to recognize the image to be recognized, it can call the character position discriminator trained by the training device of the model, so as to obtain the text content of the image to be recognized.
[0154] Similarly, in order to avoid redundant statements, the same technical features of this embodiment and the above embodiments will not be elaborated in this embodiment.
[0155] S402: Train and obtain the initial model parameters according to the first sample data set, and train and obtain the character position discriminator according to the initial model parameters.
[0156] The character position discriminator is used to analyze and process the initial characters of the image to be recognized to obtain the category attributes of the initial characters.
[0157] Figure 5 is a schematic diagram according to the fourth embodiment of the present disclosure. As Figure 5 shown, the training method of the model in the embodiment of the present disclosure includes:
[0158] S501: Obtain the initial point of interest name text, and modify the initial point of interest name text to obtain the variant point of interest name text.
[0159] In one example, the radical of the characters in the initial point of interest name text can be modified to obtain the corresponding variant point of interest name text. It should be understood that the above examples are merely illustrative of how variations of point-of-interest (POI) text may be derived from the initial POI text, and should not be construed as limiting the methods for obtaining variations of POI text.
[0164] S502: Construct a first sample dataset including the initial point of interest name text and variant point of interest name text.
[0165] S503: Input the first sample dataset into the initial language discriminator model to perform classification training on the initial language discriminator model and obtain the initial model parameters.
[0166] In this embodiment, the initial language discriminator model is trained using the first sample dataset to obtain initial model parameters. These initial model parameters are then used to train the OCR recognition model. Through two-stage training, a text position discriminator is obtained, which has a high discrimination ability. This improves the accuracy and reliability of determining the category attributes of the initial text based on the text position discriminator.
[0167] Similarly, in some embodiments, the initial language discriminator model also includes an encoder, which is a six-layer encoder.
[0168] In some embodiments, S503 may include the following steps:
[0169] The first step is to input the first sample dataset into the initial language discriminator model to obtain the predicted category attribute of each character in the first sample dataset.
[0170] The second step is to determine the initial model parameters based on the predicted category attribute and the labeled category attribute of each character.
[0171] The predicted category attribute can be understood as the prediction result obtained by predicting the type attribute of each character in the first sample dataset based on the parameters of the initial language discriminator model. For example, predicting that a character's type attribute is incorrect.
[0172] The category attribute annotation can be understood as the actual type attribute of the pre-annotated text. This annotation can be done manually or in other ways; this embodiment does not impose any limitations. For example, manually annotating a text might incorrectly label it as such.
[0173] The initial language discriminator model is trained by combining the predicted category attributes and the labeled category attributes, such as adjusting the parameters of the initial language discriminator model to make the difference value between the predicted type attributes and the labeled type attributes of the same text less than a preset threshold, so as to obtain an optimized language discriminator model, and determining the parameters of the optimized language discriminator model as the initialization model parameters, so that the text position discriminator trained based on the initialization model parameters has relatively accurate and reliable prediction ability, that is, the technical effect of making the accuracy and reliability of the category attributes of the initial text determined based on the text position discriminator. !
[0174] ! In some embodiments, the first step may include the following sub-steps: !
[0175] ! The first sub-step: For each sample data, where the sample data is the initial point of interest name text or the variant point of interest name text, determine the respective position attributes of each character in the sample data. !
[0176] ! The second sub-step: According to the respective position attributes of each character in the sample data, determine the respective predicted category attributes of each character in the sample data one by one. !
[0177] ! Exemplarily, if the sample data is "Yang AAB Store", where "A" is any character and "B" is also any character, then input "Yang AAB Store" into the initial language discriminator model. The language discriminator model can determine the position attributes corresponding to each character in "Yang AAB Store", and based on the determined position attributes, predict each character in "Yang AAB Store" in turn. !
[0178] ! For example, when predicting the character "Yang", the predicted category attribute W corresponding to the character "Yang" is obtained, that is, "Yang" is an incorrect character; when predicting the first "A" character, the predicted category attribute R corresponding to the first "A" character is obtained, that is, the first "A" character is a correct character; when predicting the second "A" character, the predicted category attribute D corresponding to the second "A" character is obtained, that is, the second "A" character is an extra character; when predicting the "B" character, the predicted category attribute R corresponding to the "B" character is obtained, that is, the "B" character is a correct character; when predicting the character "Store", the predicted category attribute R corresponding to the character "Store" is obtained, that is, the character "Store" is a correct character. !
[0179] ! In this embodiment, by combining the position attributes to determine the predicted category attributes, the ability of the initial language discriminator model to accurately identify extra characters can be trained, that is, the technical effect of improving the effectiveness and reliability of identifying extra characters. !
[0180] ! S504: Determine the initialization model parameters as the model parameters of a preset optical character recognition model. !
[0181] S505: Train the optical character recognition model based on the obtained third sample dataset to obtain the character position discriminator.
[0182] The third sample dataset includes sample image text.
[0183] In other words, in this embodiment, after training the initial language discriminator model to obtain an optimized language discriminator model, the parameters of the optimized language discriminator model are determined as the initialization model parameters, and these initialization model parameters are determined as the model parameters of the OCR recognition model, so as to train the OCR recognition model and obtain the text position discriminator.
[0184] Sample image text can be understood as the text within the collected sample images. Similarly, the number of sample image texts can be determined based on requirements, historical records, and experiments, and this embodiment does not impose any limitations.
[0185] In this embodiment, training the OCR recognition model by combining initial model parameters facilitates the convergence of the OCR recognition model training, thereby improving the efficiency of training the text position discriminator. Furthermore, the process of training the text position discriminator can be understood as two-stage training: the first stage involves classification training of the initial language discriminator model to ensure high classification performance; the second stage involves recognition training of the OCR recognition model to enhance recognition performance, thus achieving high accuracy and reliability.
[0186] In some embodiments, S505 includes: adjusting the model parameters of the OCR recognition model based on a third sample dataset to obtain a text position discriminator.
[0187] Similarly, the OCR recognition model has model parameters. The OCR recognition model with model parameters is used to recognize the text in the sample image to obtain the predicted text content of the sample image text. The text content of the sample image text is also labeled in advance to obtain the labeled text content of the sample image text. The model parameters of the OCR recognition model are adjusted based on the loss value between the predicted text content and the labeled text content to obtain the text position discriminator.
[0188] In some embodiments, the sample image text is the image text that is recognized by the OCR recognition model and results in an incorrect recognition result.
[0189] In this embodiment, by using the image text of the incorrect recognition result as the sample image text, the recognition capability of the OCR recognition model can be improved, thereby enabling the text position discriminator to have higher accuracy and reliability.
[0190] Figure 6 This is a schematic diagram based on the fifth embodiment of the present disclosure, as shown below. Figure 6 As shown, the text recognition device 600 of this embodiment includes:
[0191] The recognition unit 601 is used to perform optical character recognition on the acquired image to be recognized to obtain the initial text of the image to be recognized.
[0192] Analysis unit 602 is used to analyze and process the initial text to obtain the category attributes of the initial text.
[0193] The error correction unit 603 is used to perform error correction processing on the erroneous text if the category attribute of the initial text indicates that the initial text is an erroneous text, so as to obtain the correct text used to correct the erroneous text.
[0194] The generation unit 604 is used to generate the text content of the image to be recognized based on the initial text and the correct text.
[0195] Figure 7 This is a schematic diagram based on the sixth embodiment of the present disclosure, as shown below. Figure 7 As shown, the text recognition device 700 of this disclosure embodiment includes:
[0196] The recognition unit 701 is used to perform optical character recognition on the acquired image to be recognized to obtain the initial text of the image to be recognized.
[0197] Analysis unit 702 is used to analyze and process the initial text to obtain the category attributes of the initial text.
[0198] In some embodiments, there are multiple initial characters; the analysis unit 702 is used to analyze and process each initial character one by one to obtain the category attribute corresponding to each initial character.
[0199] In some embodiments, each initial character has a position attribute; the analysis unit 702 is used to determine the category attribute corresponding to each initial character in sequence according to the position attribute corresponding to each initial character.
[0200] In some embodiments, the analysis unit 702 is used to input the initial text into a pre-trained text position discriminator and output the category attribute of the initial text.
[0201] The text position discriminator is trained based on the initial model parameters, which are trained based on the first sample dataset. The first sample dataset includes the initial point of interest name text and the variant point of interest name text obtained by modifying the initial point of interest name text. The variant point of interest name text contains at least one erroneous character.
[0202] The error correction unit 703 is used to perform error correction processing on the erroneous text if the category attribute of the initial text indicates that the initial text is an erroneous text, so as to obtain the correct text used to correct the erroneous text.
[0203] Combination Figure 7 It is understood that, in some embodiments, the error correction unit 703 includes:
[0204] The masking subunit 7031 is used to mask erroneous characters in the initial text.
[0205] Prediction subunit 7032 is used to predict the initial text after masking to obtain a candidate set.
[0206] In some embodiments, the prediction subunit 7032 is used to input the initial text after masking into a pre-trained masked language recall model and output a candidate set.
[0207] The masked language recall model is generated based on a second sample dataset, which includes the text of sample interest point names.
[0208] The first acquisition subunit 7033 is used to acquire the correct text from the candidate set; wherein, the candidate set includes error correction text for replacing the incorrect text.
[0209] In some embodiments, the number of corrected characters is multiple; the first acquisition subunit 7033 includes:
[0210] The acquisition module is used to acquire the font structure attributes of the erroneous text and the font structure attributes of each corrected text. The font structure attributes are used to characterize the stroke content and / or stroke order of the text.
[0211] The first determination module is used to determine the correct text from the corrected text based on the font structure attributes of the incorrect text and the font structure attributes of each corrected text.
[0212] In some embodiments, the first determining module includes:
[0213] The calculation submodule is used to calculate the similarity between the font structure attributes of the corrected text and the font structure attributes of the incorrect text for each corrected text.
[0214] The determination submodule is used to identify the correct text from each error-correcting text based on its similarity score.
[0215] In some embodiments, the determining submodule is used to determine the maximum similarity from the various similarities, extract the correction text corresponding to the maximum similarity from the candidate set, and determine the correction text corresponding to the maximum similarity as the correct text.
[0216] The removal unit 704 is used to remove redundant text if the category attribute of the initial text indicates that the initial text is redundant, so as to obtain the text content of the image to be recognized.
[0217] The generation unit 705 is used to generate the text content of the image to be recognized based on the initial text and the correct text.
[0218] Combination Figure 7 It is understood that in some embodiments, the generation unit 705 is used to replace the incorrect text in the initial text with the correct text to obtain the text content of the image to be recognized.
[0219] Figure 8 This is a schematic diagram based on the seventh embodiment of the present disclosure, as shown below. Figure 8 As shown, the training device 800 for the model in this embodiment includes:
[0220] The acquisition unit 801 is used to acquire a first sample dataset, wherein the first sample dataset includes an initial point of interest name text and a variant point of interest name text obtained by modifying the initial point of interest name text, and the variant point of interest name text includes at least one erroneous character.
[0221] The first training unit 802 is used to train and obtain the initial model parameters based on the first sample dataset.
[0222] The second training unit 803 is used to train a text position discriminator based on the initial model parameters. The text position discriminator is used to analyze and process the initial text of the image to be recognized to obtain the initial text category attributes.
[0223] Figure 9 This is a schematic diagram based on the eighth embodiment of the present disclosure, as shown below. Figure 9 As shown, the training device 900 for the model in this embodiment includes:
[0224] The acquisition unit 901 is used to acquire a first sample dataset, wherein the first sample dataset includes an initial point of interest name text and a variant point of interest name text obtained by modifying the initial point of interest name text, and the variant point of interest name text includes at least one erroneous character.
[0225] Combination Figure 9 It is understood that, in some embodiments, the acquisition unit 901 includes:
[0226] The second acquisition subunit 9011 is used to acquire the initial point of interest name text.
[0227] Modify subunit 9012 to modify the initial point of interest name text to obtain a variant point of interest name text.
[0228] The first training unit 902 is used to train and obtain the initial model parameters based on the first sample dataset.
[0229] In some embodiments, the first training unit 902 is used to input the first sample dataset into the initial language discriminator model to perform classification training on the initial language discriminator model and obtain the initialized model parameters.
[0230] Combination Figure 9 It is understood that, in some embodiments, the first training unit 902 includes:
[0231] Input subunit 9021 is used to input the first sample dataset into the initial language discriminator model to obtain the predicted category attribute of each character in the first sample dataset.
[0232] In some embodiments, the input subunit 9021 includes:
[0233] The second determination module is used to determine the positional attributes of each character in each sample data, which is either the initial point of interest name text or a variant point of interest name text.
[0234] The third determination module is used to determine the predicted category attribute of each character in the sample data according to the position attribute of each character in the sample data.
[0235] The first determining subunit 9022 is used to determine the initial model parameters based on the predicted category attribute and the labeled category attribute corresponding to each character.
[0236] The second training unit 903 is used to train a text position discriminator based on the initial model parameters. The text position discriminator is used to analyze and process the initial text in the image to be recognized to obtain the initial text category attributes.
[0237] Combination Figure 9 It is understood that, in some embodiments, the second training unit 903 includes:
[0238] The second determining subunit 9031 is used to determine the initialization model parameters as the model parameters of the preset optical character recognition model.
[0239] Training subunit 9032 is used to train the optical character recognition model based on the acquired third sample dataset to obtain a text position discriminator, wherein the third sample dataset includes sample image text.
[0240] In some embodiments, the training subunit 9032 is used to adjust the model parameters of the optical character recognition model based on a third sample dataset to obtain a character position discriminator.
[0241] In some embodiments, the sample image text is the image text that is recognized based on the sample image of the optical character recognition model, resulting in an incorrect recognition result.
[0242] Figure 10 This is a schematic diagram based on the ninth embodiment of the present disclosure, as shown below. Figure 10 As shown, the electronic device 1000 in this disclosure may include a processor 1001 and a memory 1002.
[0243] Memory 1002 is used to store programs. Memory 1002 may include volatile memory, such as random-access memory (RAM), such as static random-access memory (SRAM), double data rate synchronous dynamic random-access memory (DDR SDRAM), etc.; memory may also include non-volatile memory, such as flash memory. Memory 1002 is used to store computer programs (such as application programs, functional modules, etc. that implement the above methods), computer instructions, etc. The computer programs, computer instructions, etc., can be partitioned and stored in one or more memories 1002. Furthermore, the computer programs, computer instructions, data, etc., can be accessed by processor 1001.
[0244] The aforementioned computer programs and instructions can be stored in one or more partitions of memory 1002. Furthermore, the aforementioned computer programs and instructions can be invoked by processor 1001.
[0245] The processor 1001 is configured to execute the computer program stored in the memory 1002 to implement the various steps in the methods described in the above embodiments.
[0246] For details, please refer to the relevant descriptions in the preceding method embodiments.
[0247] The processor 1001 and the memory 1002 can be independent structures or integrated structures. When the processor 1001 and the memory 1002 are independent structures, the memory 1002 and the processor 1001 can be coupled together via the bus 1003.
[0248] The electronic device in this embodiment can execute the technical solution in the above method. Its specific implementation process and technical principle are the same, and will not be repeated here.
[0249] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0250] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0251] According to embodiments of this disclosure, this disclosure also provides a computer program product comprising: a computer program stored in a readable storage medium, at least one processor of an electronic device being able to read the computer program from the readable storage medium, and the at least one processor executing the computer program causing the electronic device to perform the scheme provided in any of the above embodiments.
[0252] Figure 11 A schematic block diagram of an example electronic device 1100 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0253] like Figure 11As shown, device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1102 or a computer program loaded from storage unit 1108 into random access memory (RAM) 1103. The RAM 1103 may also store various programs and data required for the operation of device 1100. The computing unit 1101, ROM 1102, and RAM 1103 are interconnected via bus 1104. Input / output (I / O) interface 1105 is also connected to bus 1104.
[0254] Multiple components in device 1100 are connected to I / O interface 1105, including: input unit 1106, such as keyboard, mouse, etc.; output unit 1107, such as various types of monitors, speakers, etc.; storage unit 1108, such as disk, optical disk, etc.; and communication unit 1109, such as network card, modem, wireless transceiver, etc. Communication unit 1109 allows device 1100 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0255] The computing unit 1101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 performs the various methods and processes described above, such as text recognition methods and model training methods. For example, in some embodiments, the text recognition methods and model training methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1100 via ROM 1102 and / or communication unit 1109. When the computer program is loaded into RAM 1103 and executed by the computing unit 1101, one or more steps of the text recognition methods and model training methods described above can be performed. Alternatively, in other embodiments, the computing unit 1101 may be configured by any other suitable means (e.g., by means of firmware) to perform text recognition methods and model training methods.
[0256] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0257] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0258] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0259] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0260] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0261] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.
[0262] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0263] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A text recognition method, comprising: Optical character recognition is performed on the acquired image to be recognized to obtain the initial text of the image to be recognized; The number of initial characters is multiple, and each initial character has a position attribute; Based on the positional attributes of each initial character, the order of each initial character on the image to be recognized is determined. Based on the order, each initial character is sequentially input into a pre-trained character position discriminator, and the corresponding category attribute of each initial character is output. Different category attributes are identified by different flag bits. The first flag bit is used to represent incorrect characters, the second flag bit is used to represent redundant characters, and the third flag bit is used to represent correct characters. If the category attribute of the initial text indicates that the initial text is redundant, then the redundant text is removed to obtain the text content of the image to be recognized. If the category attribute of the initial text indicates that the initial text is an incorrect text, then a candidate set is obtained based on the incorrect text, and the correct text used to correct the incorrect text is obtained from the candidate set; wherein, the candidate set includes error-correcting text used to replace the incorrect text; Based on the initial text and the correct text, generate the text content of the image to be recognized; The number of correction characters is multiple; the correct characters used to correct the errors are obtained from the candidate set, including: Obtain the font structure attributes of the erroneous text, and obtain the font structure attributes of each corrected text, wherein the font structure attributes are used to characterize the stroke content and stroke order of the text. For each corrected text, calculate the similarity between the font structure attributes of the corrected text and the font structure attributes of the incorrect text. The correct text is determined from each corrected text based on its similarity score.
2. The method according to claim 1, wherein, Based on the erroneous text, a candidate set is obtained, including: The erroneous characters in the initial text are masked. The candidate set is obtained by predicting the initial text after masking.
3. The method according to claim 1, wherein, The correct text is determined from each corrected text based on its similarity score, including: Determine the maximum similarity from among all similarities; Extract the text with the highest similarity from the candidate set and determine the text with the highest similarity as the correct text.
4. The method according to claim 2, wherein, Predicting from the initial text after masking yields a candidate set, including: The initial text after masking is input into a pre-trained masked language recall model, and the candidate set is output. The masked language recall model is generated based on a second sample dataset, which includes sample interest point name text.
5. The method according to any one of claims 1-4, wherein, Based on the initial text and the correct text, the text content of the image to be recognized is generated, including: The correct text replaces the incorrect text in the initial text, thus obtaining the text content of the image to be recognized.
6. The method according to claim 1, wherein the character position discriminator is trained by the following method: Obtain the first sample dataset, where, The first sample dataset includes initial point-of-interest (POI) name text and variant POI name text obtained by modifying the initial POI name text, wherein the variant POI name text includes at least one incorrect character; The initial model parameters are obtained by training based on the first sample dataset, and the character position discriminator is obtained by training based on the initial model parameters.
7. The method according to claim 6, wherein, Obtain the first sample dataset, including: Obtain the initial point of interest (POI) name text and modify the initial POI name text to obtain the variant POI name text.
8. The method according to claim 6 or 7, wherein, The initial model parameters are obtained by training based on the first sample dataset, including: The first sample dataset is input into the initial language discriminator model to perform classification training on the initial language discriminator model, thereby obtaining the initial model parameters.
9. The method according to claim 8, wherein, The first sample dataset is input into the initial language discriminator model to perform classification training on the initial language discriminator model, thereby obtaining the initial model parameters, including: The first sample dataset is input into the initial language discriminator model to obtain the predicted category attribute of each character in the first sample dataset. The initial model parameters are determined based on the predicted category attribute and the labeled category attribute of each character.
10. The method according to claim 9, wherein, The first sample dataset is input into the initial language discriminator model to obtain the predicted category attribute corresponding to each character in the first sample dataset, including: For each sample data, which is either the initial point of interest name text or a variant point of interest name text, determine the positional attribute of each character in the sample data. Based on the positional attributes of each character in the sample data, the predicted category attribute of each character in the sample data is determined one by one.
11. The method according to any one of claims 6-7 and 9-10, wherein, A character position discriminator is trained based on the initialization model parameters, including: The initialization model parameters are determined to be the model parameters of a preset optical character recognition model; The optical character recognition model is trained based on the obtained third sample dataset to obtain the text position discriminator, wherein the third sample dataset includes sample image text.
12. The method according to claim 11, wherein, The optical character recognition model is trained based on the obtained third sample dataset to obtain the character position discriminator, including: The model parameters of the optical character recognition model are adjusted based on the third sample dataset to obtain the character position discriminator.
13. The method according to claim 11, wherein, The sample image text is an image text that has been recognized based on the sample image of the optical character recognition model, resulting in an incorrect recognition result.
14. A text recognition device, comprising: The recognition unit is used to perform optical character recognition on the acquired image to be recognized to obtain the initial text of the image to be recognized; The number of initial characters is multiple, and each initial character has a position attribute; The analysis unit is used to determine the order of each initial character in the image to be recognized based on the position attributes of each initial character. Based on the order, each initial character is input into a pre-trained character position discriminator and the category attribute of each initial character is output. Different category attributes are identified by different flag bits. The first flag bit is used to represent erroneous characters, the second flag bit is used to represent redundant characters, and the third flag bit is used to represent correct characters. The elimination unit is used to eliminate redundant text if the category attribute of the initial text indicates that the initial text is redundant text, so as to obtain the text content of the image to be recognized. An error correction unit is configured to, if the category attribute of the initial text indicates that the initial text is an erroneous text, obtain a candidate set based on the erroneous text, and obtain correct text from the candidate set to correct the erroneous text; wherein, the candidate set includes error correction text for replacing the erroneous text; The generation unit is used to generate the text content of the image to be recognized based on the initial text and the correct text. The number of corrected characters is multiple; the correction unit includes: The acquisition module is used to acquire the font structure attributes of the erroneous text and to acquire the font structure attributes of each corrected text, wherein the font structure attributes are used to characterize the stroke content and stroke order of the text. The first determining module includes a calculation submodule and a determining submodule; The calculation submodule is used to calculate the similarity between the font structure attributes of the corrected text and the font structure attributes of the incorrect text for each corrected text. The determination submodule is used to determine the correct text from each error-correcting text based on each similarity score.
15. The apparatus according to claim 14, wherein, The error correction unit includes: A masking subunit is used to mask out erroneous characters in the initial text. The prediction subunit is used to predict the initial text after masking to obtain the candidate set.
16. The apparatus according to claim 14, wherein, The determining submodule is used to determine the maximum similarity from each similarity, extract the correction text corresponding to the maximum similarity from the candidate set, and determine the correction text corresponding to the maximum similarity as the correct text.
17. The apparatus according to claim 15, wherein, The prediction subunit is used to input the initial text after masking into a pre-trained masked language recall model and output the candidate set. The masked language recall model is generated based on a second sample dataset, which includes sample interest point name text.
18. The apparatus according to any one of claims 14-17, wherein, The generation unit is used to replace the incorrect text in the initial text with the correct text to obtain the text content of the image to be recognized.
19. The apparatus according to claim 14, wherein the character position discriminator is trained by a training device, the training device comprising: An acquisition unit is used to acquire a first sample dataset, wherein the first sample dataset includes an initial point of interest (POI) name text and a variant POI name text obtained by modifying the initial POI name text, wherein the variant POI name text includes at least one erroneous character. The first training unit is used to train the initial model parameters based on the first sample dataset. The second training unit is used to train a text position discriminator based on the initial model parameters. The text position discriminator is used to analyze and process the initial text of the image to be recognized to obtain the category attribute of the initial text.
20. The apparatus according to claim 19, wherein, The acquisition unit includes: The second acquisition subunit is used to acquire the initial point of interest name text; The modification subunit is used to modify the initial point of interest name text to obtain the variant point of interest name text.
21. The apparatus according to claim 19 or 20, wherein, The first training unit is used to input the first sample dataset into the initial language discriminator model to perform classification training on the initial language discriminator model and obtain the initial model parameters.
22. The apparatus according to claim 21, wherein, The first training unit includes: The input subunit is used to input the first sample dataset into the initial language discriminator model to obtain the predicted category attribute of each character in the first sample dataset. The first determining subunit is used to determine the initialization model parameters based on the predicted category attribute and the labeled category attribute corresponding to each character.
23. The apparatus according to claim 22, wherein, The input subunit includes: The second determining module is used to determine the positional attribute of each character in each sample data, wherein the sample data is the initial point of interest name text or the variant point of interest name text. The third determination module is used to determine the predicted category attribute of each character in the sample data according to the position attribute of each character in the sample data.
24. The apparatus according to any one of claims 19-20 and 22-23, wherein, The second training unit includes: The second determining subunit is used to determine the initialization model parameters as the model parameters of a preset optical character recognition model; The training subunit is used to train the optical character recognition model based on the acquired third sample dataset to obtain the text position discriminator, wherein the third sample dataset includes sample image text.
25. The apparatus according to claim 24, wherein, The training subunit is used to adjust the model parameters of the optical character recognition model based on the third sample dataset to obtain the character position discriminator.
26. The apparatus according to claim 25, wherein, The sample image text is an image text that has been recognized based on the sample image of the optical character recognition model, resulting in an incorrect recognition result.
27. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-13.
28. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-13.
29. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1-13.
Citation Information
Patent Citations
Text information processing method and device
CN110765996A
Text error correction method, system and device and readable storage medium
CN112016310A
Text error correction method and device, computer equipment and storage medium
CN112396049A
Text recognition method and device
CN113780229A