Text relationship detection, model training method, device, equipment and medium
The method enhances text relationship detection in images by using convolutional neural networks for feature extraction and classification, improving the accuracy of structured information extraction for database construction and user profiling.
Patent Information
- Application Number
- CN202310142310.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-09
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-02-09
AI Technical Summary
The prior art is difficult to effectively identify and extract structural relationships in text images, resulting in inaccurate extraction of structured information.
By extracting and classifying text images features, using detection methods corresponding to the text structure relationship category, the structural relationship between text areas is identified, and deep learning models are used for training and optimization to improve detection accuracy.
It realizes more accurate structured information extraction, improves the accuracy and efficiency of text relationship detection, and is suitable for image processing of various text types.
Smart Images

Figure CN116152819B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence, specifically to the fields of deep learning and image processing, and more particularly to a method, apparatus, device, and medium for text relationship detection and model training. Background Art
[0002] Currently, structured documents are used to represent information, which can clearly display user information.
[0003] Obtaining this structured information helps to understand the basic situation of users, conduct targeted analysis and processing, and at the same time, a complete database and user portraits can be established. Summary of the Invention
[0004] The present disclosure provides a method, apparatus, device, and medium for text relationship detection and model training.
[0005] According to one aspect of the present disclosure, a method for text relationship detection is provided, including:
[0006] Performing feature extraction on a text image to obtain text features;
[0007] Classifying the text image according to the text features to obtain the text structure relationship category of the text image;
[0008] Using a detection method corresponding to the text structure relationship category of the text image to perform text relationship detection on the text features to obtain the structural relationship between multiple text regions in the text image.
[0009] According to another aspect of the present disclosure, a method for training a text relationship detection model is provided, including:
[0010] Obtaining training samples, where the training samples include sample images and the standard structural relationship between the standard text regions included in the sample images;
[0011] Performing feature extraction on the sample image through the feature extraction layer in the initial model to obtain sample features;
[0012] Classifying the sample image according to the sample features through the classification layer in the initial model to obtain the text structure relationship category of the sample image;
[0013] Using the detection layer corresponding to the text structure relationship category of the sample image in the initial model and using a detection method corresponding to the text structure relationship category of the sample image to perform text relationship detection on the sample features to obtain the structural relationship between multiple text regions in the sample image;
[0014] Train the initial model based on the difference between the structural relationships among multiple text regions in the sample image and the standard structural relationships among the standard text regions, to obtain a text relationship detection model.
[0015] According to one aspect of the present disclosure, there is provided a text relationship detection device, including:
[0016] A feature extraction module, configured to extract features from a text image to obtain text features;
[0017] An image classification module, configured to classify the text image according to the text features to obtain the text structure relationship category of the text image;
[0018] A structure relationship detection module, configured to perform text relationship detection on the text features by using a detection method corresponding to the text structure relationship category of the text image, to obtain the structural relationship among multiple text regions in the text image.
[0019] According to one aspect of the present disclosure, there is provided a training device for a text relationship detection module, including:
[0020] A training sample acquisition module, configured to acquire training samples, where the training samples include sample images and the standard structural relationships among the standard text regions included in the sample images;
[0021] The feature extraction layer in the initial model, configured to extract features from the sample image to obtain sample features;
[0022] The classification layer in the initial model, configured to classify the sample image according to the sample features to obtain the text structure relationship category of the sample image;
[0023] The detection layer in the initial model that uses a detection method corresponding to the text structure relationship category of the sample image, configured to perform text relationship detection on the sample features by using a detection method corresponding to the text structure relationship category of the sample image, to obtain the structural relationship among multiple text regions in the sample image;
[0024] A model training module, configured to train the initial model based on the difference between the structural relationships among multiple text regions in the sample image and the standard structural relationships among the standard text regions, to obtain a text relationship detection model.
[0025] According to another aspect of the present disclosure, there is provided an electronic device, including:
[0026] At least one processor; and
[0027] A memory communicatively connected to the at least one processor; wherein,
[0028] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the text relationship detection method or the training method of the text relationship detection model according to any embodiment of the present disclosure.
[0029] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the text relationship detection method or the training method of the text relationship detection model according to any embodiment of the present disclosure.
[0030] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, which implements the text relationship detection method or the training method of the text relationship detection model according to any embodiment of the present disclosure when executed by a processor.
[0031] Embodiments of the present disclosure can improve the accuracy of text relationship detection.
[0032] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0034] Figure 1 is a flowchart of a text relationship detection method disclosed according to an embodiment of the present disclosure;
[0035] Figure 2 is a flowchart of another text relationship detection method disclosed according to an embodiment of the present disclosure;
[0036] Figure 3 is a flowchart of another text relationship detection method disclosed according to an embodiment of the present disclosure;
[0037] Figure 4 is a flowchart of a text relationship detection model training method disclosed according to an embodiment of the present disclosure;
[0038] Figure 5 is a scenario diagram of a text relationship detection method disclosed according to an embodiment of the present disclosure;
[0039] Figure 6 is a schematic structural diagram of a text relationship detection device disclosed according to an embodiment of the present disclosure;
[0040] Figure 7It is a schematic structural diagram of an apparatus for training a text relationship detection model disclosed according to an embodiment of the present disclosure;
[0041] Figure 8 It is a block diagram of an electronic device for implementing the text relationship detection method according to an embodiment of the present disclosure. Detailed implementation manners
[0042] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0043] Figure 1 It is a flowchart of a text relationship detection method disclosed according to an embodiment of the present disclosure. This embodiment is applicable to the situation of identifying the structural relationship between texts in an image. The method of this embodiment can be executed by a text relationship detection device, which can be implemented in software and / or hardware and is specifically configured in an electronic device with certain data operation capabilities. The electronic device can be a client device or a server device. Examples of client devices include mobile phones, tablet computers, in-vehicle terminals, and desktop computers.
[0044] S101. Extract features from the text image to obtain text features.
[0045] There is a structural relationship between the texts included in the text image. The text features are used to identify the structural relationship between text regions. The text features can include the positional relationship and / or semantic relationship between texts in the text image, etc. The text image can be processed by convolution to obtain text features. Optionally, the text features can be a three-dimensional feature map. Exemplarily, the text image can be an image of a financial document, an image of a medical document, or an image of a traffic document, etc. The content and style of the text are not limited. It can be Chinese or English, and it can be Song typeface or handwritten, which can be set according to needs.
[0046] S102. Classify the text image according to the text features to obtain the text structure relationship category of the text image.
[0047] The structural relationships between texts can be divided into different text structure relationship categories. The detection methods corresponding to different text structure relationship categories are different. Exemplarily, the text structure relationship category of the text image can be the relationship of key-value pairs, or it can be a table relationship.
[0048] S103. Use a detection method corresponding to the text structure relationship category of the text image to perform text relationship detection on the text features, and obtain the structural relationship between multiple text regions in the text image.
[0049] A text region is an area including text. There is a certain distance between the texts in different text regions. And the meanings of the texts in different text regions are different. The structural relationship between text regions refers to the relationship of a two-dimensional table structure existing between the texts included in different text regions.
[0050] Exemplarily, as shown in Table 1 below:
[0051] Table 1
[0052] Name Age AA BB
[0053] Among them, in the text shown in Table 1, there is a structural relationship between the name and AA, and there is a structural relationship between the age and BB. There is no structural relationship between any two other texts.
[0054] According to the technical solution of the present disclosure, by classifying the text image and adopting an appropriate detection method according to the classification result to detect the structural relationship between text regions in the text image, more accurate structured information extraction can be achieved.
[0055] Figure 2 It is a flowchart of another text relationship detection method disclosed according to an embodiment of the present disclosure, which is further optimized and extended based on the above technical solution and can be combined with the above various optional embodiments. The structural relationship is specifically: a structural relationship in at least one direction.
[0056] S201. Extract features from the text image to obtain text features.
[0057] S202. Classify the text image according to the text features to obtain the text structure relationship category of the text image.
[0058] S203. Use a detection method corresponding to the text structure relationship category of the text image to perform text relationship detection on the text features, and obtain the structural relationship between multiple text regions in the text image; the structural relationship includes: a structural relationship in at least one direction.
[0059] Actually, in a text image, the positions of each text in the text image determine the structural relationship between the texts. The relative positions between the texts can be represented in different directions. Correspondingly, the structural relationship can be distinguished in different directions.
[0060] Correspondingly, the text structure relationship category can be determined according to the number and fixity of the directions of recognizable structure relationships. Exemplarily, for a table document, there are structure relationships in two directions between texts in different cells. For example, the row direction and the column direction. Another example is that for a general document, the direction of the structure relationship between different text pairs is not fixed, or it can be said that the existing structure relationship is a key-value relationship. At this time, it is determined that the structure relationship between texts is a non-fixed direction structure relationship. Exemplarily, the text structure relationship category can include a fixed direction relationship category or a non-fixed direction relationship category; among them, the fixed direction relationship category can be further divided according to the number of directions. For example, the structure in a table document has two directions, namely the row direction and the column direction. Correspondingly, in a table document, there are structure relationships in two directions between texts in different cells, that is, the number of directions is two.
[0061] Optionally, the text structure relationship category includes a non-fixed direction relationship category; the method of using a detection method corresponding to the text structure relationship category of the text image to perform text relationship detection on the text features to obtain the structure relationship between multiple text regions in the text image includes: using a detection method corresponding to the non-fixed direction relationship category to perform text relationship detection on the text features to obtain the probability of whether there is a structure relationship between each text region in the text image and each of the text regions.
[0062] The text structure relationship category includes a non-fixed direction relationship category, and the directions of the structure relationships between different text pairs with existing structure relationships can be the same or different. Exemplarily, there is a horizontal structure relationship between the first text pair, a vertical structure relationship between the second text pair, and a horizontal structure relationship between the third text pair. The non-fixed direction relationship category actually indicates whether there is a structure relationship between two texts, without limiting the direction of this structure relationship.
[0063] The detection method corresponding to the non-fixed direction relationship category is used to detect whether there is a structure relationship in any direction between text regions in the text image, that is, to detect whether there is a structure relationship between text regions in the text image. The result obtained by the detection method is: the probability of whether there is a structure relationship between each text region and these text regions. This includes the probability of whether there is a structure relationship between text region A and text region B. Exemplarily, the greater the probability, the higher the possibility of the existence of a structure relationship. When the probability is greater than or equal to the threshold, it is determined that there is a structure relationship between text region A and text region B; when the probability is less than the threshold, it is determined that there is no structure relationship between text region A and text region B. Among them, this detection method cannot determine in which direction the structure relationship exists. The non-fixed direction relationship category can be adapted to the detection application scenario of the key-value structure relationship.
[0064] By specifying the text structure relationship category as a non-fixed direction relationship category, the structure relationship detection for the application scenario of the non-fixed direction structure relationship can be improved, and the detection accuracy of the non-fixed direction structure relationship can be enhanced.
[0065] Optionally, the text structure relationship category includes a fixed direction relationship category, and the structure relationship includes a structure relationship in the row direction and a structure relationship in the column direction; the text relationship detection of the text feature is performed by using a detection method corresponding to the text structure relationship category of the text image to obtain the structure relationship between multiple text regions in the text image, including: performing text relationship detection on the text feature by using a detection method corresponding to the fixed direction relationship category to obtain the probability of whether there is a structure relationship in the row direction between each text region in the text image and each of the text regions, and the probability of whether there is a structure relationship in the column direction between each text region in the text image and each of the text regions.
[0066] The text structure relationship category includes a fixed direction relationship category, and the range of the direction of the structure relationship between different text pairs with a structure relationship is the row direction and the column direction. Exemplarily, among the text pairs with a structure relationship, there is a structure relationship in the row direction or the column direction between the first text pair, and there is a structure relationship in the row direction or the column direction between the second text pair.
[0067] The detection method corresponding to the non-fixed direction relationship category is used to detect whether there is a structure relationship in the row direction and whether there is a structure relationship in the column direction between the text regions in the text image. The results obtained by the detection method are: the probability of whether there is a structure relationship in the row direction and the probability of whether there is a structure relationship in the column direction between each text region and these text regions. This includes the probability of whether there is a structure relationship in the row direction and the probability of whether there is a structure relationship in the column direction between text region A and text region B. Exemplarily, the greater the probability, the higher the possibility of having a structure relationship. When the probability in the row direction is greater than or equal to the threshold, it is determined that there is a structure relationship in the row direction between text region A and text region B; when the probability in the row direction is less than the threshold, it is determined that there is no structure relationship in the row direction between text region A and text region B; when the probability in the column direction is greater than or equal to the threshold, it is determined that there is a structure relationship in the column direction between text region A and text region B; when the probability in the column direction is less than the threshold, it is determined that there is no structure relationship in the column direction between text region A and text region B. The fixed direction relationship category, and the structure relationship including the structure relationship in the row direction and the structure relationship in the column direction can adapt to the detection application scenario of the structure relationship of the table document.
[0068] By specifying the text structure relationship category as a fixed direction relationship category, the structure relationship detection for the application scenario of the structure relationship in the fixed direction can be carried out, and the detection accuracy of the structure relationship in the fixed direction can be improved.
[0069] In addition, for the fixed direction relationship category, the structure relationship can include the structure relationship in a direction other than the row direction and the column direction. For example, the structure relationship in the diagonal oblique direction; also, it can include the structure relationship in the depth (or page) direction. It can be set according to needs, and no specific limitation is imposed on this.
[0070] Optionally, the method of using the detection method corresponding to the text structure relationship category of the text image to perform text relationship detection on the text features to obtain the structure relationship between multiple text regions in the text image includes: using the detection method corresponding to the text structure relationship category of the text image to perform text relationship detection on the text features to obtain the structure relationship between multiple alternative regions; performing text detection on the text features to obtain the position and confidence of the alternative regions in the text image; and screening out the structure relationship between multiple text regions from the structure relationship between multiple alternative regions according to the position and confidence of each alternative region in the text image.
[0071] Before identifying the structure relationship of the text regions, it is also necessary to detect the text regions. Text detection is used to detect the location where the text is located. Performing text detection on the text features to identify the location of the alternative regions that may include text in the text image. Confidence is used to determine the probability that the alternative region includes text. The alternative regions identified in the text image may be misidentified, that is, there is no text in the alternative region. The position of the alternative region in the text image can be represented by two-dimensional coordinates.
[0072] Generally, the greater the confidence, the greater the probability that the alternative region includes text. The alternative regions with a confidence greater than or equal to the confidence threshold are determined as text regions. Among the structure relationships between the multiple detected alternative regions, query the two text regions with a structure relationship, and finally obtain the structure relationship between multiple text regions.
[0073] By identifying the position and confidence of the alternative regions in the text image, as well as the structure relationship between the alternative regions, and screening out the structure relationship between the text regions including text, text detection can be achieved simultaneously, and the detection accuracy of the text regions can be improved. On the basis of detecting the text regions, the structure information of the text regions is detected, and the accuracy of text structure information extraction is improved.
[0074] Optionally, the text relationship detection method further includes: obtaining a region image at the position of each text region in the text image according to the position and confidence of each candidate region in the text image; performing text recognition on the region image to obtain the text content corresponding to the region image; generating a structured text corresponding to the text image according to the structural relationship between multiple text regions in the text image and the text content corresponding to the region image.
[0075] According to the confidence of each candidate region, the positions of the text regions are preferentially screened out. Among the positions of the detected candidate regions, the positions of the screened text regions are determined. According to the positions of the text regions, the region images corresponding to the positions are determined in the text image. The region image is image data including text. Text recognition is performed on the region image corresponding to the text region to obtain the text content in the region image. Text recognition is used to detect the content of the text. The text content is combined with the structural relationship of the text to generate a structured text.
[0076] By obtaining the region image at the position of the text region in the text image and performing text recognition, the content of the text included in the text region can be obtained. Combining the structural relationship of the text region, the structured text corresponding to the text region can be generated, realizing the extraction of structured text in the image and improving the accuracy of structured text extraction.
[0077] According to the technical solution of the present disclosure, refining the structural relationship into a structural relationship in at least one direction can increase the types of text images that can be recognized, adapt to more application scenarios of text structural relationships, and improve the accuracy of structural relationship recognition.
[0078] Figure 3 It is a flowchart of another text relationship detection method disclosed in an embodiment of the present disclosure, which is further optimized and extended based on the above technical solution and can be combined with each of the above optional implementation manners. A text relationship detection model is used to implement the text relationship detection method.
[0079] S301: Extract features from the text image through the feature extraction layer in the text relationship detection model to obtain text features.
[0080] S302: Classify the text image according to the text features through the classification layer in the text relationship detection model to obtain the text structure relationship category of the text image.
[0081] S303: Perform text relationship detection on the text features through the detection layer corresponding to the text structure relationship category of the text image in the text relationship detection model to obtain the structural relationship between multiple text regions in the text image.
[0082] Optionally, through the text detection layer in the text relationship detection model, text detection is performed on the text features to obtain the positions and confidence levels of the candidate regions in the text image.
[0083] Through the convolutional layer corresponding to the text structure relationship category of the text image in the text relationship detection model, text relationship detection is performed on the text features to obtain the structural relationships between multiple candidate regions.
[0084] Through the screening layer of the text relationship detection model, according to the positions and confidence levels of the candidate regions in the text image, among the structural relationships between multiple candidate regions, the structural relationships between multiple text regions are screened out. Or, the text relationship detection model directly outputs the structural relationships of multiple candidate regions, which are implemented by the user or an additional module, and according to the positions and confidence levels of the candidate regions in the text image, the structural relationships between multiple text regions are screened out among the structural relationships between multiple candidate regions.
[0085] Among them, the convolutional layer can be divided into a first convolutional layer and a second convolutional layer.
[0086] Optionally, the text structure relationship category includes a non-fixed direction relationship category; the performing text relationship detection on the text features through the detection layer corresponding to the text structure relationship category of the text image in the text relationship detection model to obtain the structural relationships between multiple text regions in the text image includes: performing text relationship detection on the text features through the first convolutional layer corresponding to the text structure relationship category of the text image in the text relationship detection model to obtain the probability of whether there is a structural relationship between each text region in the text image and each of the text regions.
[0087] Optionally, the text structure relationship category includes a fixed direction relationship category, and the structural relationships include structural relationships in the row direction and structural relationships in the column direction; the performing text relationship detection on the text features through the detection layer corresponding to the text structure relationship category of the text image in the text relationship detection model to obtain the structural relationships between multiple text regions in the text image includes: performing text relationship detection on the text features through the second convolutional layer corresponding to the text structure relationship category of the text image in the text relationship detection model to obtain the probability of whether there is a structural relationship in the row direction between each text region in the text image and each of the text regions, and the probability of whether there is a structural relationship in the column direction between each text region in the text image and each of the text regions.
[0088] According to the technical solution of the present disclosure, by using a text relationship detection model to identify the structural relationship of text regions, the structured analysis of image text can be parsed, effectively improving the accuracy and practicality of the structured analysis of image text, and improving the efficiency of image structured extraction.
[0089] Figure 4 FIG. 4 is a flowchart of a method for training a text relationship detection model disclosed according to an embodiment of the present disclosure. This embodiment is applicable to the situation of training a text relationship detection model. The method of this embodiment can be executed by a text relationship detection device, which can be implemented in software and / or hardware and is specifically configured in an electronic device with certain data operation capabilities. The electronic device can be a client device or a server device. The client device can be, for example, a mobile phone, a tablet computer, a vehicle-mounted terminal, and a desktop computer, etc.
[0090] S401. Obtain training samples, where the training samples include a sample image and a standard structural relationship between standard text regions included in the sample image.
[0091] A large number of training samples can be obtained to train the text relationship detection model. The sample image is used to be input into the model to be trained to obtain the result output by the model. The standard structural relationship is the correct structural relationship between texts in the sample image.
[0092] S402. Extract features from the sample image through a feature extraction layer in the initial model to obtain sample features.
[0093] S403. Classify the sample image according to the sample features through a classification layer in the initial model to obtain the text structure relationship category of the sample image.
[0094] S404. Through a detection layer in the initial model that corresponds to the text structure relationship category of the sample image, and using a detection method corresponding to the text structure relationship category of the sample image, perform text relationship detection on the sample features to obtain the structural relationship between multiple text regions in the sample image. S405. Train the initial model according to the difference between the structural relationship between multiple text regions in the sample image and the standard structural relationship between the standard text regions to obtain a text relationship detection model.
[0095] Train the initial model according to the difference to obtain a text relationship detection model. A loss function can be constructed according to the difference.
[0096] Optionally, perform text detection on the sample features through a text detection layer in the initial model to obtain the position and confidence level of the alternative region in the sample image.
[0097] Text relationship detection is performed on the sample features by using a convolutional layer corresponding to the text structure relationship category of the sample image in the initial model, and the structural relationships between multiple candidate regions are obtained.
[0098] Through the screening layer of the initial model, according to the positions and confidences of the candidate regions in the sample image, among the structural relationships between multiple candidate regions, the structural relationships between multiple text regions are screened out. Alternatively, the initial model directly outputs the structural relationships of multiple candidate regions, which are implemented by the user or an additional module to screen out the structural relationships between multiple text regions among the structural relationships between multiple candidate regions according to the positions and confidences of the candidate regions in the sample image.
[0099] Among them, the convolutional layer can be divided into a first convolutional layer and a second convolutional layer.
[0100] Optionally, the text structure relationship category includes a non-fixed direction relationship category; the text relationship detection is performed on the sample features by using a detection layer corresponding to the text structure relationship category of the sample image in the initial model to obtain the structural relationships between multiple text regions in the sample image, including: performing text relationship detection on the sample features by using the first convolutional layer corresponding to the text structure relationship category of the sample image in the initial model to obtain the probability of whether there is a structural relationship between each text region in the sample image and each of the text regions.
[0101] Optionally, the text structure relationship category includes a fixed direction relationship category, and the structural relationships include structural relationships in the row direction and structural relationships in the column direction; the text relationship detection is performed on the sample features by using a detection layer corresponding to the text structure relationship category of the sample image in the initial model to obtain the structural relationships between multiple text regions in the sample image, including: performing text relationship detection on the sample features by using the second convolutional layer corresponding to the text structure relationship category of the sample image in the initial model to obtain the probability of whether there is a structural relationship in the row direction between each text region in the sample image and each of the text regions, and the probability of whether there is a structural relationship in the column direction between each text region in the sample image and each of the text regions.
[0102] The training sample also includes the standard position of the standard text region of the sample image.
[0103] In summary, based on the difference between the structural relationship among the text regions recognized by the initial model and the standard structural relationship, the difference between the text regions recognized by the initial model and the standard text regions, and the difference between the position of the text regions recognized by the initial model and the standard position of the standard text regions, the above cumulative difference value or the weighted sum of the above differences can be determined as the loss function to minimize or converge the loss function, thereby determining that the training of the initial model is completed.
[0104] According to the technical solution of the present disclosure, by using the text relationship detection model to recognize the structural relationship of text regions, the structural analysis of image text can be parsed, effectively improving the accuracy and practicality of the structural analysis of image text, as well as improving the efficiency of image structure extraction.
[0105] Figure 5 It is a flowchart of a text relationship detection method provided according to the technical solution of the present disclosure. As Figure 5 shown, the input of this text relationship detection method is a single picture, and it supports the structural text recognition of images of medical documents in the key-value structured form.
[0106] Specifically: Obtain the image to be detected, normalize the original picture with the given mean and variance, and constrain the pixel values to be between 0 and 1 to obtain the text image. Input the text image into the feature extraction layer in the text relationship detection model, where the feature extraction layer is a convolutional neural network. The feature extraction layer outputs the text features of the text image, that is, a three-dimensional feature map.
[0107] Input the three-dimensional feature map into the classification layer in the text relationship detection model to obtain the text structure relationship category. In the classification layer, the three-dimensional feature map is globally normalized and flattened into a one-dimensional feature vector. The one-dimensional feature vector is input into the fully connected layer, and through softmax normalization, the classification result of the text image is obtained, that is, the text structure relationship category text structure relationship category.
[0108] Input the three-dimensional feature map into the text detection layer in the text relationship detection model to obtain the position and confidence of the alternative regions. Among them, the text detection layer is a convolutional layer. Specifically, the alternative regions are in a rectangular shape, and the position of the alternative regions can be determined by the four corner points of the rectangle. The position of the alternative regions is represented by an n*m matrix, where n*m is the number of alternative regions; each element in the matrix represents a 1*8 vector, and the vector represents the position information of the element corresponding to the alternative region, specifically 4 corner points, and each corner point includes two values of x and y in the coordinate system of the text image, thus forming a 1*8 vector. The confidence is also represented by an n*m matrix, and each element in the matrix represents the confidence of whether the element is a text classification.
[0109] For the image of a general document, according to the classification result, determine that the text structure relationship category includes the non-fixed direction relationship category. Input the three-dimensional feature map into the first convolutional layer corresponding to the text structure relationship category in the text relationship detection model, perform text relationship detection on the three-dimensional feature map, and obtain the probability of whether there is a structure relationship between each alternative region and other alternative regions in the text image. The output result is represented by an n*m matrix, and each element of the matrix represents a 1*(m*n) vector, that is, the relationship between the alternative region corresponding to the current element and any other alternative region. For example, the element in the first row and the first column represents the probability of whether there is a structure relationship between the element (name) corresponding to the first row and the first column and the element (Zhang San) corresponding to the second position in the 1*(m*n) vector. A probability greater than 0.5 indicates that there is a structure relationship between the element corresponding to the first row and the first column and the element corresponding to the second position in the 1*(m*n) vector.
[0110] For the image of a table document, according to the classification result, determine that the text structure relationship category includes the fixed direction relationship category. The text structure relationship category includes the fixed direction relationship category, and the structure relationship includes the structure relationship in the row direction and the structure relationship in the column direction. Input the three-dimensional feature map into the second convolutional layer corresponding to the text structure relationship category in the text relationship detection model, perform text relationship detection on the three-dimensional feature map, and obtain the probability of whether there is a structure relationship in the row direction between each alternative region and other alternative regions in the text image, and obtain the probability of whether there is a structure relationship in the column direction between each alternative region and other alternative regions in the text image. The output results include two, one is the structure relationship in the row direction, and the other is the structure relationship in the column direction. The structure relationship in each direction can be represented by an n*m matrix, and each element of the matrix represents a 1*(m*n) vector, that is, the relationship between the alternative region corresponding to the current element and any other alternative region. For example, the element in the first row and the first column represents the probability of whether there is a row (or column) structure relationship between the element (name) corresponding to the first row and the first column and the element (Zhang San) corresponding to the second position in the 1*(m*n) vector. A probability greater than 0.5 indicates that there is a row (or column) structure relationship between the element corresponding to the first row and the first column and the element corresponding to the second position in the 1*(m*n) vector.
[0111] Actually, the structural relationship of the output is the structural relationship between alternative regions. Combining with the text detection confidence of the alternative regions, text regions are screened out, and the structural relationship between texts is determined. For the image of a general document, the prediction result of the text relationship of the general document is obtained, representing the structural relationship between any two text regions. For the image of a table document, two sets of text relationship prediction results are obtained. One set represents the structural relationship between any two text regions in the horizontal direction (row direction), and the other set represents the structural relationship between any two text regions in the vertical direction (column direction). The structured parsing results of general documents and table documents are obtained. According to the structured parsing results, structured text can be generated and stored in the database, or the structured text can be audited.
[0112] As Figure 5 shown, the structural relationship is represented by a connection relationship.
[0113] According to the technical solution of the present disclosure, a solution for structured parsing of medical documents can be provided. By innovatively using a fully convolutional network, the detection of text and the extraction of structured information are simultaneously realized. This solution can extract more accurate structured information. The method for extracting structured information based on a fully convolutional network, on the one hand, learns the positions of text regions, and on the other hand, models the connection relationships between any two text regions in both horizontal and vertical directions. Finally, the extraction of table structured information can be realized, which is a generalizable image text structured parsing solution, effectively improving the accuracy and practicality of image text structured parsing.
[0114] According to an embodiment of the present disclosure, Figure 6 is a structural diagram of a text relationship detection device in an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the situation of identifying the structural relationship between texts in an image. The device is implemented by software and / or hardware and is specifically configured in an electronic device with certain data operation capabilities.
[0115] As Figure 6 shown, a text relationship detection device 600 includes: a feature extraction module 601, an image classification module 602, and a structural relationship detection module 603. Among them,
[0116] The feature extraction module 601 is used to extract features from the text image to obtain text features;
[0117] The image classification module 602 is used to classify the text image according to the text features to obtain the text structure relationship category of the text image;
[0118] The structural relationship detection module 603 is used to perform text relationship detection on the text features by using a detection method corresponding to the text structure relationship category of the text image to obtain the structural relationship between multiple text regions in the text image.
[0119] According to the technical solution of the present disclosure, by classifying a text image and adopting an appropriate detection method according to the classification result to detect the structural relationship between text regions in the text image, more accurate structured information extraction can be achieved.
[0120] Furthermore, the structural relationship includes: structural relationships in at least one direction.
[0121] Furthermore, the text structure relationship category includes a non-fixed direction relationship category; the structure relationship detection module 603 includes: a one-way relationship detection unit, which is used to perform text relationship detection on the text features by adopting a detection method corresponding to the non-fixed direction relationship category, so as to obtain the probability of whether there is a structural relationship between each text region in the text image and each of the text regions.
[0122] Furthermore, the text structure relationship category includes a fixed direction relationship category, and the structural relationship includes a structural relationship in the row direction and a structural relationship in the column direction; the structure relationship detection module 603 includes: a table relationship detection unit, which is used to perform text relationship detection on the text features by adopting a detection method corresponding to the fixed direction relationship category, so as to obtain the probability of whether there is a structural relationship in the row direction between each text region in the text image and each of the text regions, and the probability of whether there is a structural relationship in the column direction between each text region in the text image and each of the text regions.
[0123] Furthermore, the structure relationship detection module 603 includes: an alternative region structure detection unit, which is used to perform text relationship detection on the text features by adopting a detection method corresponding to the text structure relationship category of the text image, so as to obtain the structural relationship between multiple alternative regions; a text recognition unit, which is used to perform text detection on the text features to obtain the position and confidence of the alternative regions in the text image; a text region screening unit, which is used to screen out the structural relationship between multiple text regions from the structural relationship between multiple alternative regions according to the position and confidence of each alternative region in the text image.
[0124] Furthermore, the text relationship detection device further includes: a region image acquisition module, which is used to acquire region images at the positions of each text region in the text image according to the position and confidence of each alternative region in the text image; a text recognition module, which is used to perform text recognition on the region images to obtain the text content corresponding to the region images; a structured text generation module, which is used to generate the structured text corresponding to the text image according to the structural relationship between multiple text regions in the text image and the text content corresponding to the region images.
[0125] Furthermore, the feature extraction module 601 includes a feature extraction layer in the text relationship detection model, which is used to extract features from the text image to obtain text features; the image classification module 602 includes a classification layer in the text relationship detection model, which is used to classify the text image according to the text features to obtain the text structure relationship category of the text image; the structure relationship detection module 603 includes a detection layer in the text relationship detection model corresponding to the text structure relationship category of the text image, which is used to perform text relationship detection on the text features to obtain the structure relationship between multiple text regions in the text image.
[0126] The above text relationship detection device can execute the text relationship detection method provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the text relationship detection method.
[0127] According to an embodiment of the present disclosure, Figure 7 is a structural diagram of a training device for a text relationship detection model in an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the situation of training a text relationship detection model. The device is implemented by software and / or hardware, and is specifically configured in an electronic device with certain data operation capabilities.
[0128] Such as Figure 7 A training device 700 for a text relationship detection model shown in the figure includes a training sample acquisition module 701, a feature extraction layer 702 in the initial model, a classification layer 703 in the initial model, a detection layer 704 in the initial model corresponding to the text structure relationship category of the sample image, and a model training module 705. Among them,
[0129] The training sample acquisition module 701 is used to acquire training samples, and the training samples include the standard structure relationship between the sample image and the standard text regions included in the sample image;
[0130] The feature extraction layer 702 in the initial model is used to extract features from the sample image to obtain sample features;
[0131] The classification layer 703 in the initial model is used to classify the sample image according to the sample features to obtain the text structure relationship category of the sample image;
[0132] The detection layer 704 in the initial model corresponding to the text structure relationship category of the sample image is used to perform text relationship detection on the sample features by using a detection method corresponding to the text structure relationship category of the sample image to obtain the structure relationship between multiple text regions in the sample image;
[0133] The model training module 705 is configured to train the initial model according to the difference between the structural relationships among multiple text regions in the sample image and the standard structural relationships among the standard text regions, so as to obtain a text relationship detection model.
[0134] According to the technical solution of the present disclosure, by using the text relationship detection model to identify the structural relationships of text regions, the structured parsing of image text can be analyzed, effectively improving the accuracy and practicality of the structured parsing of image text, and improving the efficiency of image structured extraction.
[0135] The above-mentioned training method and device of the text relationship detection model can execute the training method of the text relationship detection model provided in any embodiment of the present disclosure, and have the corresponding functional modules and beneficial effects for executing the training method of the text relationship detection model.
[0136] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0137] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0138] Figure 8 FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0139] As Figure 8 shown, the device 800 includes a computing unit 801 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0140] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as a keyboard, mouse, etc.; output unit 807, such as various types of displays, speakers, etc.; storage unit 808, such as a disk, optical disc, etc.; and communication unit 809, such as a network card, modem, wireless communication transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0141] Computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 801 executes the various methods and processes described above, such as a text relationship detection method or a training method for a text relationship detection model. For example, in some embodiments, a text relationship detection method or a training method for a text relationship detection model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by computing unit 801, one or more steps of the text relationship detection method or the training method for the text relationship detection model described above can be executed. Alternatively, in other embodiments, computing unit 801 can be configured to execute the text relationship detection method or the training method for the text relationship detection model in any other suitable manner (e.g., by means of firmware).
[0142] The various embodiments of the systems and techniques described above in this article can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0143] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0144] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0145] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0146] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0147] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services. The server can also be a server of a distributed system or a server combined with blockchain.
[0148] Artificial intelligence is a discipline that studies how to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), and it has both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, machine learning / deep learning technology, big data processing technology, and knowledge graph technology.
[0149] Cloud computing refers to a technical system that accesses an elastic and scalable shared physical or virtual resource pool through a network. The resources can include servers, operating systems, networks, software, applications, and storage devices, etc., and the resources can be deployed and managed in a on-demand and self-service manner. Through cloud computing technology, it can provide efficient and powerful data processing capabilities for the application and model training of technologies such as artificial intelligence and blockchain.
[0150] It should be understood that various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions provided in this disclosure can be achieved, and no limitation is made herein.
[0151] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A text relationship detection method, comprising: Performing feature extraction on a text image to obtain text features; Classifying the text image according to the text features to obtain the text structure relationship category of the text image; Adopting a detection method corresponding to the text structure relationship category of the text image to perform text relationship detection on the text features, so as to obtain the structural relationship between multiple text regions in the text image; The text structure relationship category includes a non-fixed direction relationship category; The adopting a detection method corresponding to the text structure relationship category of the text image to perform text relationship detection on the text features, so as to obtain the structural relationship between multiple text regions in the text image, includes: Adopting a detection method corresponding to the non-fixed direction relationship category to perform text relationship detection on the text features, so as to obtain the probability of whether there is a structural relationship between each text region in the text image and each of the text regions.
2. The method according to claim 1, wherein, The structural relationship includes a structural relationship in at least one direction.
3. The method according to claim 2, wherein The text structure relationship category includes a fixed direction relationship category, and the structural relationship includes a structural relationship in the row direction and a structural relationship in the column direction; The adopting a detection method corresponding to the text structure relationship category of the text image to perform text relationship detection on the text features, so as to obtain the structural relationship between multiple text regions in the text image, includes: Adopting a detection method corresponding to the fixed direction relationship category to perform text relationship detection on the text features, so as to obtain the probability of whether there is a structural relationship in the row direction between each text region in the text image and each of the text regions, and the probability of whether there is a structural relationship in the column direction between each text region in the text image and each of the text regions.
4. The method according to claim 1, wherein, The adopting a detection method corresponding to the text structure relationship category of the text image to perform text relationship detection on the text features, so as to obtain the structural relationship between multiple text regions in the text image, includes: Adopting a detection method corresponding to the text structure relationship category of the text image to perform text relationship detection on the text features, so as to obtain the structural relationship between multiple alternative regions; Performing text detection on the text features to obtain the position and confidence of the alternative regions in the text image; According to the position and confidence of each alternative region in the text image, screening out the structural relationship between multiple text regions from the structural relationship between multiple alternative regions.
5. The method according to claim 4, further comprising: According to the position and confidence of each alternative region in the text image, obtaining a region image at the position of each text region in the text image; Performing text recognition on the region image to obtain the text content corresponding to the region image; Generating a structured text corresponding to the text image according to the structural relationship between multiple text regions in the text image and the text content corresponding to the region image.
6. The method according to claim 1, wherein The performing feature extraction on a text image to obtain text features includes: Extract features from the text image through the feature extraction layer in the text relationship detection model to obtain text features; Classify the text image according to the text features to obtain the text structure relationship category of the text image, including: Classify the text image according to the text features through the classification layer in the text relationship detection model to obtain the text structure relationship category of the text image; Adopt the detection method corresponding to the text structure relationship category of the text image to perform text relationship detection on the text features to obtain the structure relationship between multiple text regions in the text image, including: Perform text relationship detection on the text features through the detection layer corresponding to the text structure relationship category of the text image to obtain the structure relationship between multiple text regions in the text image.
7. A training method for a text relationship detection model, including: Obtain training samples, where the training samples include the standard structure relationship between the sample image and the standard text regions included in the sample image; Extract features from the sample image through the feature extraction layer in the initial model to obtain sample features; Classify the sample image according to the sample features through the classification layer in the initial model to obtain the text structure relationship category of the sample image; Adopt the detection method corresponding to the text structure relationship category of the sample image to perform text relationship detection on the sample features through the detection layer corresponding to the text structure relationship category of the sample image to obtain the structure relationship between multiple text regions in the sample image; the text structure relationship category includes the non-fixed direction relationship category; Perform text relationship detection on the sample features through the first convolutional layer corresponding to the text structure relationship category of the sample image to obtain the probability of whether there is a structure relationship between each text region in the sample image and each of the text regions; Train the initial model according to the difference between the structure relationship between multiple text regions in the sample image and the standard structure relationship between the standard text regions to obtain a text relationship detection model.
8. A text relationship detection device, including: A feature extraction module for extracting features from a text image to obtain text features; An image classification module for classifying the text image according to the text features to obtain the text structure relationship category of the text image; A structure relationship detection module for performing text relationship detection on the text features by adopting a detection method corresponding to the text structure relationship category of the text image to obtain the structure relationship between multiple text regions in the text image; The text structure relationship category includes the non-fixed direction relationship category; The structure relationship detection module includes: A one-way relationship detection unit for performing text relationship detection on the text features by adopting a detection method corresponding to the non-fixed direction relationship category to obtain the probability of whether there is a structure relationship between each text region in the text image and each of the text regions.
9. The device according to claim 8, wherein, The structural relationship includes: the structural relationship in at least one direction.
10. The device according to claim 9, wherein, The category of the text structural relationship includes the fixed direction relationship category, and the structural relationship includes the structural relationship in the row direction and the structural relationship in the column direction; The structural relationship detection module includes: The table relationship detection unit is used to perform text relationship detection on the text features by using the detection method corresponding to the fixed direction relationship category, and obtain the probability of whether there is a structural relationship in the row direction between each text region in the text image and each of the text regions, and the probability of whether there is a structural relationship in the column direction between each text region in the text image and each of the text regions.
11. The device according to claim 8, wherein, The structural relationship detection module includes: The alternative region structure detection unit is used to perform text relationship detection on the text features by using the detection method corresponding to the text structural relationship category of the text image, and obtain the structural relationship between multiple alternative regions; The text recognition unit is used to perform text detection on the text features to obtain the position and confidence of the alternative regions in the text image; The text region screening unit is used to screen out the structural relationship between multiple text regions from the structural relationship between multiple alternative regions according to the position and confidence of each alternative region in the text image.
12. The device according to claim 11 further includes: The region image acquisition module is used to acquire the region images at the positions of each text region in the text image according to the position and confidence of each alternative region in the text image; The text recognition module is used to perform text recognition on the region images to obtain the text content corresponding to the region images; The structured text generation module is used to generate the structured text corresponding to the text image according to the structural relationship between multiple text regions in the text image and the text content corresponding to the region images.
13. The apparatus according to claim 8, wherein, The feature extraction module includes: The feature extraction layer in the text relationship detection model is used to perform feature extraction on the text image to obtain text features; The image classification module includes: The classification layer in the text relationship detection model is used to classify the text image according to the text features to obtain the text structural relationship category of the text image; The structural relationship detection module includes: The detection layer corresponding to the text structural relationship category of the text image in the text relationship detection model is used to perform text relationship detection on the text features to obtain the structural relationship between multiple text regions in the text image.
14. A training device for a text relationship detection model includes: The training sample acquisition module is used to acquire training samples, and the training samples include the standard structural relationship between the sample images and the standard text regions included in the sample images; The feature extraction layer in the initial model is used to perform feature extraction on the sample images to obtain sample features; The classification layer in the initial model is used to classify the sample images according to the sample features to obtain the text structural relationship category of the sample images. The detection layer corresponding to the text structure relationship category of the sample image in the initial model is used to perform text relationship detection on the sample features by using a detection method corresponding to the text structure relationship category of the sample image, so as to obtain the structural relationship between multiple text regions in the sample image; the text structure relationship category includes a non-fixed direction relationship category; The first convolutional layer corresponding to the text structure relationship category of the sample image in the initial model is used to perform text relationship detection on the sample features, so as to obtain the probability of whether there is a structural relationship between each text region in the sample image and each of the text regions; The model training module is used to train the initial model according to the difference between the structural relationship between multiple text regions in the sample image and the standard structural relationship between the standard text regions, so as to obtain a text relationship detection model.
15. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, so that the at least one processor can execute the text relationship detection method according to any one of claims 1-6, or the training method of the text relationship detection model according to claim 7.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the text relationship detection method according to any one of claims 1-6, or the training method of the text relationship detection model according to claim 7.
17. A computer program product, comprising a computer program, where the computer program, when executed by a processor, implements the text relationship detection method according to any one of claims 1-6, or the training method of the text relationship detection model according to claim 7.
Citation Information
Patent Citations
Structured model training method, text structuring method and related devices
CN109582800A
Bill image identification method, device and equipment, and storage medium
CN111709339A