Artificial intelligence-based text information extraction method, device, equipment and medium
By using a well-trained character recognition model and language model, and by utilizing the center point and boundary information of the recognized characters, the problem of low accuracy in extracting text information caused by the uncertainty of the spacing between printed characters is solved, and higher accuracy in character recognition and semantic analysis is achieved.
Patent Information
- Application Number
- CN202210958496.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-09
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2042-08-09
AI Technical Summary
In existing technologies, the spacing between printed characters in bounding box-based optical character recognition technology is uncertain, which leads to disordered contextual information during character recognition and low accuracy in extracting text information.
The recognition results of printed text are obtained by using a trained character recognition model. The center point of the recognized character is used to determine the adjacent related characters. The boundary information is determined based on the center point of the related characters. When the preset conditions are met, the boundary feature value is determined. The boundary feature vector is then constructed and input into the language model for text information extraction.
This improves the accuracy of character recognition and semantic analysis, thereby enhancing the accuracy of text information extraction.
Smart Images

Figure CN115294578B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and medium for extracting text information based on artificial intelligence. Background Technology
[0002] Currently, with the development of artificial intelligence technology, the digitization of text information has gradually shifted from manual input into computers to machine recognition and extraction. This is usually achieved using bounding box-based optical character recognition (OCR) technology. OCR is the process of translating the shape of printed characters into computer text. In the OCR process, closely spaced printed characters are grouped into the same bounding box, and the printed characters within the same bounding box are treated as a word or a sentence, which can improve the accuracy of text recognition.
[0003] However, the spacing between printed characters in printed text is usually uncertain, and there may be multiple words or sentences within the bounding box. Since the character recognition process takes into account contextual information, such situations can lead to contextual confusion, introducing irrelevant information during character recognition and resulting in low accuracy of text information extraction. Therefore, improving the accuracy of text information extraction has become an urgent problem to be solved. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a method, apparatus, device and medium for extracting text information based on artificial intelligence, in order to solve the problem of low accuracy in text information extraction.
[0005] In a first aspect, embodiments of the present invention provide a text information extraction method based on artificial intelligence, the text information extraction method comprising:
[0006] The acquired image to be processed is input into the trained character recognition model to obtain the recognition result, which includes at least one recognized character and the center point of the corresponding recognized character.
[0007] For any given character, based on the center point of the character, identify the characters adjacent to it as associated characters;
[0008] Based on the center point of the associated character, determine the boundary information of the identified character;
[0009] When the boundary information is detected to meet the preset conditions, the boundary feature value of the identified character is determined to be the first feature value; otherwise, the boundary feature value of the identified character is determined to be the second feature value, thus obtaining the boundary feature value of each identified character.
[0010] The boundary feature vector composed of the boundary feature values of all recognized characters is used as the embedding vector. The character sequence composed of the embedding vector and the recognized characters is input into the trained language model to obtain the text information extraction result.
[0011] Secondly, embodiments of the present invention provide a text information extraction device based on artificial intelligence, the text information extraction device comprising:
[0012] The character recognition module is used to input the acquired image to be processed into the trained character recognition model to obtain the recognition result, which includes at least one recognized character and the center point of the corresponding recognized character;
[0013] The character association module is used to determine, for any given character, an associated character based on the center point of the character.
[0014] A boundary determination module is used to determine the boundary information of the identified character based on the center point of the associated character;
[0015] The feature value determination module is used to determine the boundary feature value of the identified character as a first feature value when the boundary information is detected to meet a preset condition; otherwise, it determines the boundary feature value of the identified character as a second feature value, thereby obtaining the boundary feature value of each identified character.
[0016] The information extraction module is used to take the boundary feature vector composed of the boundary feature values of all recognized characters as the embedding vector, and input the character sequence composed of the embedding vector and the recognized characters into the trained language model to obtain the text information extraction result.
[0017] Thirdly, embodiments of the present invention provide a computer device, the computer device including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the text information extraction method as described in the first aspect.
[0018] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the text information extraction method as described in the first aspect.
[0019] The beneficial effects of the embodiments of the present invention compared with the prior art are as follows:
[0020] The acquired image to be processed is input into a trained character recognition model to obtain recognition results. The recognition results include at least one recognized character and the center point of the corresponding recognized character. For any recognized character, based on the center point of the recognized character, adjacent recognized characters are determined as associated characters. Based on the center points of the associated characters, the boundary information of the recognized character is determined. When the boundary information meets the preset conditions, the boundary feature value of the recognized character is determined as the first feature value; otherwise, the boundary feature value of the recognized character is determined as the second feature value. The boundary feature value of each recognized character is obtained. The boundary feature vector composed of the boundary feature values of all recognized characters is used as the embedding vector. The character sequence composed of the embedding vector and the recognized characters is input into a trained language model to obtain the text information extraction results. Center point prediction is performed for each recognized character, enabling the character recognition model to segment characters more accurately and improving the accuracy of character recognition. At the same time, the boundary feature vector is constructed based on the boundary information of the recognized characters, providing effective positional information for the language model and improving the accuracy of semantic analysis, thereby improving the accuracy of text information extraction. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a schematic diagram of an application environment for a text information extraction method based on artificial intelligence provided in Embodiment 1 of the present invention;
[0023] Figure 2 This is a flowchart illustrating a text information extraction method based on artificial intelligence provided in Embodiment 1 of the present invention;
[0024] Figure 3 This is a flowchart illustrating a text information extraction method based on artificial intelligence provided in Embodiment 2 of the present invention;
[0025] Figure 4 This is a schematic diagram of the structure of a text information extraction device based on artificial intelligence provided in Embodiment 3 of the present invention;
[0026] Figure 5 This is a schematic diagram of the structure of a computer device provided in Embodiment 4 of the present invention. Detailed Implementation
[0027] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of the invention with unnecessary detail.
[0028] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0029] It should also be understood that the term “and / or” as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0030] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."
[0031] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0032] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0033] The embodiments of this invention can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that utilize digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0034] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0035] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0036] To illustrate the technical solution of the present invention, specific embodiments are described below.
[0037] The first embodiment of this invention provides a text information extraction method based on artificial intelligence, which can be applied to, for example... Figure 1 In this application environment, the client and server communicate with each other. Clients include, but are not limited to, PDAs, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud terminal devices, and personal digital assistants (PDAs). The server can be implemented using a standalone server or a server cluster consisting of multiple servers.
[0038] See Figure 2 This is a flowchart illustrating a text information extraction method based on artificial intelligence provided in Embodiment 1 of the present invention. The above-described text information extraction method can be applied to... Figure 1 The client, a computer device connected to the server, obtains the image to be processed for text information extraction. The client's computer device is equipped with a pre-trained character recognition model and a pre-trained language model. The character recognition model can be used to recognize printed text as computer characters, and the language model can be used to analyze the computer characters to obtain semantic information. For example... Figure 2 As shown, this text information extraction method may include the following steps:
[0039] Step S201: Input the acquired image to be processed into the trained character recognition model to obtain the recognition result.
[0040] The image to be processed can be a printed text image that requires semantic extraction, the character recognition model can be an optical character recognition model, the optical character recognition model can include a text detection model and a text recognition model, and the recognition result includes at least one recognized character and the center point of the corresponding recognized character.
[0041] Specifically, before inputting the image to be processed into the trained character recognition model, the image to be processed is normalized in size, that is, the size of the image to be processed is scaled down to a fixed size to adapt to the input parameters of the trained character recognition model.
[0042] Optionally, the trained character recognition model includes a trained text detection model and a trained text recognition model;
[0043] The acquired image to be processed is input into the trained character recognition model, and the recognition results include:
[0044] Input the image to be processed into the trained text detection model to obtain the bounding box localization points;
[0045] The bounding box is determined based on the bounding box positioning points, and the image of the region to be processed is obtained by cropping from the image to be processed based on the bounding box.
[0046] The image of the region to be processed is input into the trained text recognition model to obtain at least one recognized character and the center point of the corresponding recognized character.
[0047] The text detection model can be used to determine the location of the text to be identified in an image. The bounding box location points can include the upper left corner and the lower right corner of the bounding box. The bounding box can refer to a rectangular box in the image to be processed, and the image to be processed can refer to an image containing only the text to be identified.
[0048] Text recognition models can be used to determine the computer characters corresponding to the text to be recognized. The recognized characters can refer to the computer characters corresponding to the text to be recognized, and the center point of the recognized characters can refer to the center point of the corresponding area of the recognized characters in the image to be processed.
[0049] Specifically, text detection models can employ object detection models, such as Faster R-CNN, SSD (Single Shot MultiBox Detector), and FPN (Feature Pyramid Net). These models can obtain bounding box locations. For example, if the top-left corner is (x1, y1) and the bottom-right corner is (x2, y2), then the bottom-left and top-right corners can be determined using these points, i.e., the bottom-left corner is (x1, y2) and the top-right corner is (x2, y1). Any two corners with the same x and y coordinates are connected. For instance, if the bottom-left and top-left corners have the same x-coordinate, they are connected; if both x and y coordinates are different, they are not connected. The resulting connection is the edge of the bounding box.
[0050] It should be noted that in this embodiment, the bounding boxes are all regular rectangles, that is, the sides of the bounding boxes are all parallel to the X-axis or Y-axis of the image coordinate system.
[0051] In this embodiment, to ensure that the cropped image size is consistent, cropping can be achieved using a masking method. That is, a mask image is generated based on the bounding box. The pixel value of the pixels within the bounding box area in the mask image is 1, and the pixel value of other pixels is 0. The mask image is multiplied point by point with the image to be processed. The pixel values of the pixels within the bounding box area in the image to be processed are retained, and the pixel values of other pixels are set to 0. Thus, the cropping effect is achieved while ensuring the cropped image size.
[0052] This embodiment uses bounding boxes to locate the text, thereby cropping the image to be processed. Character recognition is performed only on the regions containing text information, thus isolating irrelevant information in the image to be processed and improving the accuracy of character recognition.
[0053] Optionally, the sample image of the region to be processed can be used as the training sample for training the text recognition model, the actual characters can be used as the training labels for training the text recognition model, and the cross-entropy loss can be used as the loss function for training the text recognition model.
[0054] The training process for a text recognition model includes:
[0055] The sample image of the region to be processed is divided into M sub-sample images according to the preset step size;
[0056] For any given subsample image, input the subsample image into the text recognition model to obtain the initial sample characters;
[0057] Based on the initial sample characters and the actual characters, the cross-entropy loss is calculated. Using the cross-entropy loss as a basis, the parameters of the text recognition model are updated using the gradient descent method until the cross-entropy loss converges, thus obtaining a preliminarily trained text recognition model.
[0058] Among them, the sample image of the region to be processed can refer to the image obtained by processing and cropping the above character detection model using historical printed text as a sample, the sub-sample image can refer to the image after dividing the sample image of the region to be processed according to a preset step size, and the actual character can refer to the real character corresponding to the text, that is, the character corresponding to the text is known.
[0059] The preset step size can refer to the character segmentation step size, the subsample image can refer to an image containing partial information of a single character, the initial sample character can refer to the character recognition result of the corresponding subsample image, and the gradient descent method can refer to the stochastic gradient descent method.
[0060] Specifically, the size of the sample image of the region to be processed is the same as the size of the image of the region to be processed. The characters in the sample image of the region to be processed may be complete characters or incomplete characters. In this case, since the label is a complete actual character, the sample image of the region to be processed with incomplete characters will also be identified as the closest complete character.
[0061] This embodiment uses sample images of the region to be processed for preliminary training of the text recognition model, which can effectively recognize incomplete characters, thereby improving the accuracy of character recognition.
[0062] Optionally, after obtaining the initially trained text recognition model, the following steps are also included:
[0063] For any subsample image, input the subsample image into the pre-trained text recognition model to obtain the updated sample character and the center point of the corresponding updated sample character;
[0064] Sub-sample images belonging to the same updated sample character are merged to obtain N updated sub-sample images. The center point of each updated sub-sample image is extracted as the center point label of the recognized character.
[0065] Based on the center point of the corresponding updated sample character and the center point label of the recognized character, the cross-entropy loss is calculated. Based on the cross-entropy loss, the parameters of the text recognition model are updated using the gradient descent method until the cross-entropy loss converges, thus obtaining the trained text recognition model.
[0066] Here, the updated sample character can refer to the preliminary recognition result corresponding to the subsample image, and the center point of the updated sample character can refer to the prediction result of the center position of the range of the updated sample character.
[0067] Merging can refer to treating the sub-sample images to be merged as a single image, i.e., updating the sub-sample image. Updating the sub-sample image can refer to an image containing complete character information. The center point of the updated sub-sample image can include the mean of the center points of all the sub-sample images to be merged. In this embodiment, the mean can refer to the mean of the horizontal coordinate.
[0068] Specifically, in this embodiment, a center point prediction branch is added, that is, the text recognition model includes a recognition branch and a center point prediction branch. The recognition branch and the center point prediction branch share a feature extraction encoder to obtain text feature vectors, and the text feature vectors are respectively input into the recognition fully connected layer of the recognition branch and the prediction fully connected layer of the center point prediction branch.
[0069] It should be noted that random position perturbation can be applied to the subsample images. In this embodiment, the horizontal coordinate perturbation range is set to [0,2], with the unit being the number of pixels. That is, the subsample images are shifted from 0 to 2 pixels, thereby achieving the effect of data augmentation and increasing the robustness and accuracy of the trained character recognition model.
[0070] In this embodiment, the center point label of the identified character can be calculated from the center points of the sub-sample images to be merged, thus eliminating the need for manual annotation. This improves the prediction content of the character recognition model while also increasing the training efficiency of the character recognition model.
[0071] The steps described above, which involve inputting the acquired image to be processed into a trained character recognition model to obtain the recognition result, convert the image of the printed text to be processed into computer characters. This facilitates semantic extraction of the printed text in computer character form, effectively reduces the analysis difficulty of the subsequent language model, and improves the accuracy of text information extraction.
[0072] Step S202: For any given character, determine the adjacent characters as associated characters based on the center point of the character.
[0073] Among them, the associated character can refer to the character adjacent to the identification character.
[0074] Specifically, after obtaining the center point of each recognized character, for any recognized character, let its center point be (x... a ,y a Then, based on the x-coordinate of its center point... a The search function finds the associated horizontal coordinate with the smallest absolute difference from the given horizontal coordinate, and the character corresponding to the associated horizontal coordinate is the associated character.
[0075] It should be noted that when searching for the associated horizontal coordinate with the smallest absolute difference from the horizontal coordinate, the first two associated horizontal coordinates with the smallest absolute difference are retained, and it is checked whether both associated horizontal coordinates are greater than the horizontal coordinate. If both associated horizontal coordinates are greater than the horizontal coordinate, the character corresponding to the smaller value of the two associated horizontal coordinates is determined as the associated character. If one of the two associated horizontal coordinates is greater than the horizontal coordinate, the characters corresponding to both associated horizontal coordinates are determined as the associated characters.
[0076] The above steps, which determine the adjacent characters as associated characters based on the center point of any given character, facilitate the provision of positional information for subsequent semantic extraction by the language model, thereby improving the accuracy of text information extraction.
[0077] Step S203: Determine the boundary information of the identified character based on the center point of the associated character.
[0078] The boundary information can refer to the dividing line between the identified character and other semantically unrelated identified characters. In this embodiment, the dividing line can refer to a straight line parallel to the Y-axis.
[0079] Optionally, based on the center point of the associated character, the boundary information of the identified character is determined, including:
[0080] Compare the x-coordinate of the center point of the associated character with the x-coordinate of the center point of the recognized character. If the x-coordinate of the center point of the associated character is less than the x-coordinate of the center point of the recognized character, then the associated character is determined to be a left associated character; otherwise, the associated character is determined to be a right associated character.
[0081] Calculate the first mean of the x-coordinate of the center point of the corresponding left associated character and the x-coordinate of the center point of the corresponding recognized character, and determine the first mean as the x-coordinate of the left boundary.
[0082] Calculate the second mean of the x-coordinate of the center point of the corresponding right associated character and the x-coordinate of the center point of the corresponding recognized character, and determine the second mean as the x-coordinate of the right boundary;
[0083] The left and right boundary coordinates are determined as the boundary information for character recognition.
[0084] Among them, the left associated character can refer to the associated character to the left of the recognized character, and the right associated character can refer to the associated character to the right of the recognized character.
[0085] The left boundary x-coordinate can refer to the x-coordinate used to determine the left boundary line, and the right boundary x-coordinate can refer to the x-coordinate used to determine the right boundary line.
[0086] Specifically, let the x-coordinate of the center point of the recognized character be x. aThe x-coordinate of the center point of the associated character is x. b To facilitate calculation, the image coordinate system is set with the bottom left corner of the image as the origin, the ray pointing from the bottom left corner to the top left corner as the Y-axis, and the ray pointing from the bottom left corner to the bottom right corner as the X-axis. This ensures that the x-coordinate of the center point of the recognized character is always greater than 0. For example, if x... a Less than x b The associated character is the right associated character, if x a Greater than x b The associated character is the left associated character, and the mean of the horizontal axis is... Let x be the x-coordinate of the boundary corresponding to the associated character. Since the boundary lines in this embodiment are all parallel to the Y-axis, the boundary lines are represented as follows:
[0087] It should be noted that since the identified character may only have one associated character, in this case, one of the left boundary x-coordinate and the right boundary x-coordinate is used as the boundary information.
[0088] This embodiment further refines the boundary information by classifying the associated characters into categories and calculating the left and right boundary coordinates, thereby improving the ability of the boundary information to represent the character position information and thus improving the accuracy of text information extraction.
[0089] The above steps, which determine the boundary information of the characters based on the center point of the associated characters, use the center point to determine the boundary information, which can more accurately segment the characters according to their position information. This avoids the situation where semantically unrelated characters are extracted from the text information through the language model at the same time, which would cause interference between the characters and result in a low accuracy of text information extraction.
[0090] Step S204: When the boundary information is detected to meet the preset conditions, the boundary feature value of the recognized character is determined to be the first feature value; otherwise, the boundary feature value of the recognized character is determined to be the second feature value, thus obtaining the boundary feature value of each recognized character.
[0091] Here, the boundary feature value can refer to the encoded value used to represent the boundary information of the identified characters. For example, in this embodiment, the first feature value is set to 1, and the second feature value is set to 0. The preset condition can refer to the condition used to determine that the boundary information is a dividing boundary between character sequences.
[0092] Optionally, the process of detecting whether the boundary information meets the preset conditions includes:
[0093] When a character is detected to have a left associated character, the difference between the left boundary x-coordinate of the character and the right boundary x-coordinate of the associated character is calculated to obtain the first difference.
[0094] When a character with a right-linked character is detected, the difference between the right boundary x-coordinate of the character and the left boundary x-coordinate of the right-linked character is calculated to obtain the second difference.
[0095] Calculate the ratio of the first difference and the second difference, and compare the calculation result with a preset threshold. If the calculation result is greater than the preset threshold, then determine the right boundary horizontal coordinate of the left associated character and the left boundary horizontal coordinate of the right associated character as the segmentation boundary information.
[0096] The first difference can be used to characterize the distance between the identified character and the left associated character in the image to be processed, the second difference can be used to characterize the distance between the identified character and the right associated character in the image to be processed, and the preset threshold can be used to determine whether the identified character and the associated character are the same string. For example, in this embodiment, the preset threshold is set to 0.6.
[0097] When the calculation result is greater than the preset threshold, it indicates that the distance between the two recognized characters is large. According to the printing habits of printed text, the distance between characters belonging to the same string is small. Therefore, it can be determined that the recognized characters can be divided from the middle of the two recognized characters, and the character sequence is divided into two strings. Repeat the division judgment for each pair of adjacent recognized characters, and finally, several segments of strings can be obtained from the character sequence.
[0098] In this embodiment, the boundary information is determined to be a segmentation boundary by threshold comparison. The calculation is fast and simple, and the threshold can be flexibly adjusted according to the actual situation, thereby improving the efficiency and accuracy of text information extraction.
[0099] The above steps, which determine the boundary feature value of the identified character as the first feature value when the boundary information meets the preset conditions, and otherwise determine the boundary feature value of the identified character as the second feature value, and obtain the boundary feature value of each identified character, determine whether the boundary information is the segmentation boundary between characters by using preset conditions, and then segment the character sequence, providing character segmentation information for the subsequent language model, thereby improving the accuracy of text information extraction.
[0100] Step S205: Use the boundary feature vector composed of the boundary feature values of all recognized characters as the embedding vector, and input the character sequence composed of the embedding vector and the recognized characters into the trained language model to obtain the text information extraction result.
[0101] Among them, the boundary feature vector can refer to a 1*K dimensional vector composed of boundary feature values, that is, a vector with one row and K columns, where K can refer to the number of characters to be recognized.
[0102] Specifically, the embedding vector is input into the first encoder for feature extraction to obtain the embedded feature vector. The character sequence composed of the recognized characters is input into the second encoder for feature extraction to obtain the character feature vector. The embedded feature vector and the character feature vector are fused together, and the feature fusion result is input into the recurrent network model for semantic extraction to obtain the semantic representation, which is the result of text information extraction. The above recurrent network model can adopt the Long Short Term Memory (LSTM) model.
[0103] The above steps, which use the boundary feature vector composed of the boundary feature values of all recognized characters as the embedding vector, and input the character sequence composed of the embedded vector and the recognized characters into the trained language model to obtain the text information extraction result, improve the richness of the recognized character features through multi-dimensional feature collaboration in text information extraction, thereby providing more feature information for semantic extraction and effectively improving the accuracy of text information extraction.
[0104] This embodiment predicts the center point for each recognized character, enabling the character recognition model to segment characters more accurately and improving the accuracy of character recognition. At the same time, it constructs boundary feature vectors based on the boundary information of the recognized characters, providing effective positional information for the language model and improving the accuracy of semantic analysis, thereby improving the accuracy of text information extraction.
[0105] See Figure 3 This is a flowchart illustrating a text information extraction method based on artificial intelligence provided in Embodiment 2 of the present invention. In this text information extraction method, when inputting a character sequence composed of an embedding vector and recognized characters into a trained language model, the character sequence can be directly used as input to the trained language model, or features can be constructed from the character sequence and then the constructed features can be used as input to the trained language model.
[0106] The process of directly inputting the character sequence into the trained language model is described in Example 1 and will not be repeated here.
[0107] The process of constructing features from a character sequence and then using those features as input to a trained language model includes the following steps:
[0108] Step S301: By using preset title terms, perform term matching on all recognized characters in the character sequence, and assign corresponding identifiers to the successfully matched recognized characters according to the matched title terms to obtain an identifier vector;
[0109] Step S302: Convert the character sequence into text word vectors through word vector embedding;
[0110] Step S303: The text word vectors, identifier vectors and embedding vectors are fused to form features, and the feature fusion result is input into the trained language model to obtain the text information extraction result.
[0111] The preset title terms can refer to a pre-stored database of common titles. Common titles can include basic titles such as "date," "number," and "name," as well as specific application titles such as "quota" and "time limit." Term matching can refer to regular expression matching, and the identifier can refer to encoded information. In this embodiment, encoded information in numerical form is used, meaning different numbers correspond to different terms.
[0112] Word embedding refers to the process of converting computer characters into word vectors. Word embedding can employ models such as Bidirectional Encoder Representation from Transformers (BERT) and Word2Vec.
[0113] In this embodiment, feature fusion can refer to feature concatenation, which is to concatenate different features according to their dimensions. The feature fusion result can be a feature vector obtained by concatenating text word vectors, identifier vectors, and embedding vectors. The text information extraction result can refer to semantic information.
[0114] In one implementation, feature fusion may refer to feature multiplication or feature addition point by point.
[0115] This embodiment constructs features through character sequences and then uses the constructed features as input to the trained language model, thereby improving the representation ability of the features corresponding to the character sequences and avoiding situations where character sequence features with rich information cannot be mapped to semantic information through the language model, thus improving the accuracy of text information extraction.
[0116] Corresponding to the AI-based text information extraction method in the above embodiments, Figure 4 A structural block diagram of an artificial intelligence-based text information extraction device according to Embodiment 3 of the present invention is shown. This text information extraction device is applied to a client. The computer device corresponding to the client connects to the server to obtain the image to be processed for text information extraction. The computer device corresponding to the client is equipped with a trained character recognition model and a trained language model. The trained character recognition model can be used to recognize printed text as computer characters, and the trained language model can be used to analyze semantic information based on the computer characters. For ease of explanation, only the parts relevant to the embodiments of the present invention are shown.
[0117] See Figure 4 The text information extraction device includes:
[0118] The character recognition module 41 is used to input the acquired image to be processed into the trained character recognition model to obtain the recognition result, which includes at least one recognized character and the center point of the corresponding recognized character.
[0119] The character association module 42 is used to determine the adjacent recognition characters as associated characters based on the center point of any recognition character;
[0120] Boundary determination module 43 is used to determine the boundary information of the recognized characters based on the center point of the associated characters;
[0121] The feature value determination module 44 is used to determine the boundary feature value of the recognized character as the first feature value when the detected boundary information meets the preset conditions; otherwise, it determines the boundary feature value of the recognized character as the second feature value, thereby obtaining the boundary feature value of each recognized character.
[0122] The information extraction module 45 is used as an embedding vector composed of the boundary feature values of all recognized characters. The embedding vector and the character sequence composed of the recognized characters are input into the trained language model to obtain the text information extraction result.
[0123] Optionally, the trained character recognition model includes a trained text detection model and a trained text recognition model;
[0124] The character recognition module 41 mentioned above includes:
[0125] The text localization unit is used to input the image to be processed into the trained text detection model to obtain the bounding box localization points;
[0126] The image cropping unit determines the bounding box based on the bounding box positioning points, and crops the region image to be processed from the image to be processed based on the bounding box.
[0127] The text recognition unit is used to input the image of the region to be processed into the trained text recognition model to obtain at least one recognized character and the center point of the corresponding recognized character.
[0128] Optionally, the sample image of the region to be processed can be used as the training sample for training the text recognition model, the actual characters can be used as the training labels for training the text recognition model, and the cross-entropy loss can be used as the loss function for training the text recognition model.
[0129] The aforementioned text information extraction device also includes:
[0130] The sample segmentation module is used to divide the sample image of the region to be processed into M sub-sample images according to a preset step size.
[0131] The first sample recognition module is used to input the sub-sample image into the text recognition model for any sub-sample image to obtain the initial sample character;
[0132] The first training module is used to calculate the cross-entropy loss based on the initial sample characters and the actual characters. Based on the cross-entropy loss, the parameters of the text recognition model are updated using the gradient descent method until the cross-entropy loss converges, thus obtaining the preliminarily trained text recognition model.
[0133] Optionally, the above-mentioned text information extraction device further includes:
[0134] The second sample recognition module is used to input the sub-sample image into the pre-trained text recognition model for any sub-sample image to obtain the updated sample character and the center point of the corresponding updated sample character.
[0135] The sample merging module is used to merge sub-sample images belonging to the same updated sample character to obtain N updated sub-sample images, and extract the center point of each updated sub-sample image as the center point label of the recognized character.
[0136] The second training module is used to calculate the cross-entropy loss based on the center point of the corresponding updated sample character and the center point label of the recognized character. Based on the cross-entropy loss, the parameters of the text recognition model are updated using the gradient descent method until the cross-entropy loss converges, thus obtaining the trained text recognition model.
[0137] Optionally, the boundary determination module 43 mentioned above includes:
[0138] The association determination unit is used to compare the horizontal coordinate of the center point of the associated character with the horizontal coordinate of the center point of the recognized character. If the horizontal coordinate of the center point of the associated character is less than the horizontal coordinate of the center point of the recognized character, the associated character is determined to be a left associated character; otherwise, the associated character is determined to be a right associated character.
[0139] The first mean calculation unit is used to calculate the first mean of the x-coordinate of the center point of the corresponding left associated character and the x-coordinate of the center point of the corresponding recognized character, and to determine the first mean as the x-coordinate of the left boundary.
[0140] The second mean calculation unit is used to calculate the second mean of the x-coordinate of the center point of the corresponding right associated character and the x-coordinate of the center point of the corresponding recognized character, and to determine the second mean as the x-coordinate of the right boundary.
[0141] The boundary information acquisition unit is used to determine the left and right boundary coordinates as the boundary information of the recognized character.
[0142] Optionally, the aforementioned feature value determination module 44 includes:
[0143] The first difference calculation unit is used to calculate the difference between the left boundary horizontal coordinate of the recognized character and the right boundary horizontal coordinate of the left associated character when a left associated character is detected, and obtain the first difference.
[0144] The second difference calculation unit is used to calculate the difference between the right boundary horizontal coordinate of the recognized character and the left boundary horizontal coordinate of the right associated character when a right associated character is detected, and obtain the second difference.
[0145] The threshold comparison unit is used to calculate the ratio of the first difference and the second difference, and compare the calculation result with the preset threshold. If the calculation result is greater than the preset threshold, the right boundary horizontal coordinate of the left associated character and the left boundary horizontal coordinate of the right associated character are determined as the segmentation boundary information.
[0146] Optionally, the information extraction module 45 mentioned above includes:
[0147] The identifier vector determination unit is used to perform term matching on all the recognized characters in the character sequence through preset title terms, and assign corresponding identifiers to the successfully matched recognized characters according to the matched title terms to obtain the identifier vector;
[0148] The word vector transformation unit is used to convert character sequences into text word vectors through word vector embedding;
[0149] The feature fusion unit is used to fuse text word vectors, identifier vectors, and embedding vectors, and input the feature fusion result into the trained language model to obtain the text information extraction result.
[0150] It should be noted that the information interaction and execution process between the above modules and units are based on the same concept as the method embodiments of the present invention. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.
[0151] Figure 5 This is a schematic diagram of the structure of a computer device provided in Embodiment 4 of the present invention. Figure 5 As shown, the computer device of this embodiment includes: at least one processor ( Figure 5 Only one is shown in the diagram), a memory, and a computer program stored in the memory and capable of running on at least one processor, wherein the processor executes the computer program to implement the steps in any of the above-described text information extraction method embodiments.
[0152] This computer device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that... Figure 5The examples of computer devices are merely examples and do not constitute a limitation on computer devices. Computer devices may include more or fewer components than shown in the illustration, or combinations of certain components, or different components, such as network interfaces, displays, and input devices.
[0153] The processor referred to can be a CPU, but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0154] Memory includes readable storage media, internal memory, etc., wherein internal memory can be the RAM of a computer device, providing an environment for the operation of the operating system and computer-readable instructions stored in the readable storage media. The readable storage media can be the hard drive of a computer device, or in other embodiments, it can be an external storage device of the computer device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, memory can include both internal storage units and external storage devices of the computer device. Memory is used to store the operating system, applications, bootloader, data, and other programs, such as program code for computer programs. Memory can also be used to temporarily store data that has been output or will be output.
[0155] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the functions described above can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this invention. The specific working process of the units and modules in the above device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention can implement all or part of the processes in the methods of the above embodiments by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the above method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium can include at least: any entity or device capable of carrying computer program code, a recording medium, a computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0156] The present invention can implement all or part of the processes in the methods of the above embodiments, or it can be accomplished by a computer program product. When the computer program product is run on a computer device, the computer device executes the steps in the above method embodiments.
[0157] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0158] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0159] In the embodiments provided by this invention, it should be understood that the disclosed apparatus / computer devices and methods can be implemented in other ways. For example, the apparatus / computer device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0160] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0161] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. An artificial intelligence-based text information extraction method, characterized by, The method comprises: inputting the obtained to-be-processed image into a trained character recognition model to obtain a recognition result, the recognition result comprising at least one recognized character and a center point of the corresponding recognized character; for any recognized character, determining, according to the center point of the recognized character, a recognized character adjacent to the recognized character as an associated character; determining, according to the center point of the associated character, boundary information of the recognized character; when it is detected that the boundary information meets a preset condition, determining a boundary feature value of the recognized character as a first feature value, or otherwise, determining the boundary feature value of the recognized character as a second feature value, to obtain a boundary feature value of each recognized character; inputting a boundary feature vector composed of the boundary feature values of all the recognized characters as an embedding vector, and inputting a character sequence composed of the embedding vector and the recognized characters into a trained language model to obtain a text information extraction result; the determining, according to the center point of the associated character, of the boundary information of the recognized character comprises: comparing the center point horizontal coordinate of the associated character with the center point horizontal coordinate of the recognized character, and if the center point horizontal coordinate of the associated character is less than the center point horizontal coordinate of the recognized character, determining that the associated character is a left associated character, or otherwise, determining that the associated character is a right associated character; calculating a first average value of the center point horizontal coordinate of the corresponding left associated character and the center point horizontal coordinate of the corresponding recognized character, and determining the first average value as a left boundary horizontal coordinate; calculating a second average value of the center point horizontal coordinate of the corresponding right associated character and the center point horizontal coordinate of the corresponding recognized character, and determining the second average value as a right boundary horizontal coordinate; determining the left boundary horizontal coordinate and the right boundary horizontal coordinate as the boundary information of the recognized character; the process of detecting whether the boundary information meets the preset condition comprises: when it is detected that the recognized character has the left associated character, calculating a difference value between the left boundary horizontal coordinate of the recognized character and the right boundary horizontal coordinate of the left associated character to obtain a first difference value; when it is detected that the recognized character has the right associated character, calculating a difference value between the right boundary horizontal coordinate of the recognized character and the left boundary horizontal coordinate of the right associated character to obtain a second difference value; calculating a ratio of the first difference value and the second difference value, comparing the calculation result with a preset threshold value, and if the calculation result is greater than the preset threshold value, determining that the right boundary horizontal coordinate of the left associated character and the left boundary horizontal coordinate of the right associated character are split boundary information.
2. The text information extraction method according to claim 1, characterized by, the trained character recognition model comprises a trained text detection model and a trained text recognition model; the inputting the obtained to-be-processed image into the trained character recognition model to obtain a recognition result comprises: inputting the to-be-processed image into the trained text detection model to obtain a bounding box positioning point; determining a bounding box according to the bounding box positioning point, and cutting a to-be-processed region image from the to-be-processed image according to the bounding box; inputting the to-be-processed region image into the trained text recognition model to obtain at least one recognized character and a center point of the corresponding recognized character.
3. The text information extraction method according to claim 2, characterized by, The text recognition model is trained by using the sample image of the to-be-processed region as a training sample, using actual characters as training labels, and using a cross-entropy loss as a loss function; The training process of the text recognition model comprises: dividing the sample image of the to-be-processed region into M sub-sample images according to a preset step size; for any sub-sample image, inputting the sub-sample image into the text recognition model to obtain initial sample characters; calculating the cross-entropy loss according to the initial sample characters and the actual characters, and updating the parameters of the text recognition model by using a gradient descent method according to the cross-entropy loss until the cross-entropy loss converges, to obtain a preliminarily trained text recognition model.
4. The text information extraction method according to claim 3, characterized by, After the preliminarily trained text recognition model is obtained, the method further comprises: for any sub-sample image, inputting the sub-sample image into the preliminarily trained text recognition model to obtain updated sample characters and corresponding center points of the updated sample characters; merging sub-sample images belonging to the same updated sample character to obtain N updated sub-sample images, and extracting the center points of each updated sub-sample image as recognition character center point labels; calculating the cross-entropy loss according to the corresponding center points of the updated sample characters and the recognition character center point labels, and updating the parameters of the text recognition model by using a gradient descent method according to the cross-entropy loss until the cross-entropy loss converges, to obtain a trained text recognition model.
5. The text information extraction method according to claim 1, characterized by, The inputting of the character sequence composed of the embedding vector and the recognition characters into the trained language model to obtain a text information extraction result comprises: performing term matching on all recognition characters in the character sequence by using a preset title term, and assigning corresponding identifiers to the recognition characters for which the matching is successful according to the matched title terms, to obtain an identifier vector; converting the character sequence into a text word vector by using word vector embedding; performing feature fusion on the text word vector, the identifier vector and the embedding vector, and inputting the feature fusion result into the trained language model to obtain a text information extraction result.
6. An artificial intelligence-based text information extraction device, characterized by, The text information extraction device comprises: a character recognition module configured to input a to-be-processed image obtained by the image acquisition module into a trained character recognition model to obtain a recognition result, the recognition result comprising at least one recognition character and a center point of the corresponding recognition character; a character association module configured to, for any recognition character, determine, according to the center point of the recognition character, a recognition character adjacent to the recognition character as an associated character; a boundary determination module configured to determine, according to the center points of the associated characters, boundary information of the recognition character; a feature value determination module configured to, when detecting that the boundary information satisfies a preset condition, determine a boundary feature value of the recognition character as a first feature value, or otherwise, determine the boundary feature value of the recognition character as a second feature value, to obtain a boundary feature value of each recognition character; The information extraction module is configured to input a boundary feature vector composed of boundary feature values of all recognized characters and a character sequence composed of the recognized characters into a trained language model as an embedding vector to obtain a text information extraction result. The boundary determination module comprises: The association determination unit is configured to compare a horizontal coordinate of a center point of the associated character with a horizontal coordinate of a center point of the recognized character, and if the horizontal coordinate of the center point of the associated character is less than the horizontal coordinate of the center point of the recognized character, determine the associated character as a left associated character, otherwise, determine the associated character as a right associated character. The first mean value calculation unit is configured to calculate a first mean value of the horizontal coordinate of the center point of the left associated character and the horizontal coordinate of the center point of the recognized character, and determine the first mean value as a left boundary horizontal coordinate. The second mean value calculation unit is configured to calculate a second mean value of the horizontal coordinate of the center point of the right associated character and the horizontal coordinate of the center point of the recognized character, and determine the second mean value as a right boundary horizontal coordinate. The boundary information acquisition unit is configured to determine the left boundary horizontal coordinate and the right boundary horizontal coordinate as boundary information of the recognized character. The feature value determination module comprises: The first difference calculation unit is configured to calculate a difference between the left boundary horizontal coordinate of the recognized character and a right boundary horizontal coordinate of the left associated character when the recognized character is detected to have the left associated character, to obtain a first difference. The second difference calculation unit is configured to calculate a difference between the right boundary horizontal coordinate of the recognized character and a left boundary horizontal coordinate of the right associated character when the recognized character is detected to have the right associated character, to obtain a second difference. The threshold comparison unit is configured to calculate a ratio of the first difference and the second difference, compare the calculation result with a preset threshold, and if the calculation result is greater than the preset threshold, determine the right boundary horizontal coordinate of the left associated character and the left boundary horizontal coordinate of the right associated character as segmentation boundary information.
7. A computer device, characterized by The computer device comprises a processor, a memory, and a computer program stored in the memory and executable on the processor, and the processor implements the text information extraction method according to any one of claims 1 to 5 when executing the computer program.
8. A computer-readable storage medium storing a computer program, the computer-readable storage medium comprising: The computer program is executable on the processor to implement the text information extraction method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Image character segmentation method, device and equipment and storage medium
CN108446702A
Character recognition method and device, electronic equipment and storage medium
CN111428723A