Text knowledge extraction method and device, electronic equipment and readable storage medium

By using an object detection model to perform text detection and classification on agricultural books, the problem of poor template transferability in knowledge extraction of agricultural books is solved, achieving efficient and accurate text knowledge extraction, reducing the professional knowledge requirements, and making it applicable to semi-structured book recognition in various fields.

CN115964492BActive Publication Date: 2026-03-20SHANGHAI IFLYTEK HEGUANG TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-30
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing agricultural book knowledge extraction templates have poor transferability and high specialization, resulting in high knowledge extraction thresholds and high misidentification rates, making them difficult to widely disseminate and apply.

Method used

A target detection model is used to detect semi-structured text images, identify target text detection boxes, their regions, and categories. Through classification and knowledge extraction strategies, sub-text images of different text categories are automatically processed, reducing the need for professional knowledge and improving recognition accuracy.

Benefits of technology

It achieves high transferability and high accuracy in recognizing semi-structured books in different fields, lowers the knowledge extraction threshold, and improves the efficiency and accuracy of text knowledge extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115964492B_ABST
    Figure CN115964492B_ABST
Patent Text Reader

Abstract

The application provides a text knowledge extraction method and device, electronic equipment and readable storage medium, and relates to the technical field of knowledge extraction. The method comprises the following steps: inputting a to-be-recognized text image into a target detection model to obtain target region position information of a target text detection box in a text region of interest and a target text category; obtaining a subtext image corresponding to the text region of interest where the target text detection box is located based on the target region position information, and dividing the subtext image corresponding to each target text detection box into at least two types of subtext images based on the target text category; and obtaining target text knowledge corresponding to the to-be-recognized text image based on the target region position information and the target text category of the text region of interest where each type of subtext image is located and a corresponding knowledge extraction strategy, so as to solve the technical problem of how to improve the accuracy and transferability of the knowledge extraction method and reduce the threshold of text knowledge extraction in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of knowledge extraction, and in particular to a text knowledge extraction method and device, an electronic device and a readable storage medium. BACKGROUND

[0002] Agriculture, as the first industry in China, plays a vital role in economic development and social stability. At present, there are a large number of agricultural books, and the knowledge search efficiency is low, the agricultural knowledge base is relatively limited and the quality is not guaranteed, so the arrangement of the agricultural knowledge base is very important for the popularization of agricultural book knowledge and the popularization of agricultural technology.

[0003] In the prior art, optical character recognition is performed on each page of the book image in the agricultural book to convert the picture data into text data, so as to extract entity information, description information and other knowledge from the text data based on a preset knowledge extraction template. However, the rule setting of the knowledge extraction template will be affected by the content of books in different fields, so the migration of the knowledge extraction template is poor, and the complexity and professionalism of the knowledge extraction template are high, so the knowledge extraction template needs to have professional knowledge in the related field, thereby increasing the threshold of knowledge extraction, which is not conducive to the popularization of agricultural technology and the popularization of agricultural book knowledge. In addition, the optical character recognition method has the technical defect of misrecognition when applied to a specific technical field.

[0004] Therefore, how to improve the accuracy and migration of the knowledge extraction method and reduce the threshold of text knowledge extraction is a technical problem to be solved by the related field technical personnel. SUMMARY

[0005] The present application provides a text knowledge extraction method, device, electronic device and readable storage medium to solve the technical problem of how to improve the accuracy and migration of the knowledge extraction method and reduce the threshold of text knowledge extraction in the prior art.

[0006] The present application provides a text knowledge extraction method, comprising:

[0007] inputting a to-be-recognized text image into a pre-trained target detection model to obtain at least one target text detection box and target region position information and a target text category of a text region of interest where each target text detection box is located, the to-be-recognized text image being a semi-structured text image;

[0008] based on the target region position information, cutting a subtext image corresponding to the text region of interest where each target text detection box is located from the to-be-recognized text image, and based on the target text category, classifying the subtext image corresponding to the text region of interest where each target text detection box is located to obtain at least two types of subtext images;

[0009] based on the target region position information and the target text category of the text region of interest in which each of the at least two types of subtext images is located and the corresponding knowledge extraction strategy, obtaining target text knowledge corresponding to the to-be-recognized text image.

[0010] According to the text knowledge extraction method provided by the application, the at least two types of subtext images include a first subtext image and a second subtext image;

[0011] The classification of the subtext image corresponding to the text region of interest in which each of the target text detection boxes is located based on the target text category includes:

[0012] In the case where the target text category of the text region of interest in which the target text detection box is located is a first text category, the subtext image corresponding to the text region of interest in which the target text detection box is located is determined to be the first subtext image, and the first text category includes special symbols.

[0013] In the case where the target text category of the text region of interest in which the target text detection box is located is a second text category, the subtext image corresponding to the text region of interest in which the target text detection box is located is determined to be the second subtext image, and the second text category includes entities, attributes, and attribute values.

[0014] According to the text knowledge extraction method provided by the application, based on the target region position information and the target text category of the text region of interest in which each of the at least two types of subtext images is located and the corresponding knowledge extraction strategy, the target text knowledge corresponding to the to-be-recognized text image is obtained, including:

[0015] Based on the knowledge extraction strategy corresponding to each of the subtext images, first text knowledge corresponding to each of the subtext images is obtained, and the knowledge extraction strategy is determined based on the target text category of the text region of interest in which the subtext image is located.

[0016] Based on the target region position information of the text region of interest in which each of the subtext images is located, at least two subtext images in the same image subregion in the to-be-recognized text image are determined.

[0017] For each of the image subregions, based on the target region position information and the target text category of the text region of interest in which each of the subtext images in the image subregion is located and the corresponding first text knowledge, second text knowledge corresponding to the image subregion is determined.

[0018] Determine the target text knowledge corresponding to the text image to be recognized based on the second text knowledge corresponding to each of the image sub-regions in the text image to be recognized and the sub-text images in each of the image sub-regions.

[0019] According to the text knowledge extraction method provided by the application, the second text knowledge corresponding to the image sub-region is determined based on the target region position information and the target text category of the target text region where each of the sub-text images in the image sub-region is located and the corresponding first text knowledge, and the method comprises the following steps:

[0020] Determine the first category subordination relationship between each of the sub-text images in the image sub-region based on the target text category of the target text region where each of the sub-text images in the image sub-region is located, wherein the first category subordination relationship comprises the category subordination relationship between entities, attributes and attribute values.

[0021] Perform data correlation processing on the first text knowledge corresponding to each of the sub-text images in the image sub-region based on the first category subordination relationship, so as to obtain the second text knowledge corresponding to the image sub-region.

[0022] According to the text knowledge extraction method provided by the application, the target text knowledge corresponding to the text image to be recognized is determined based on the second text knowledge corresponding to each of the image sub-regions in the text image to be recognized and the sub-text images in each of the image sub-regions, and the method comprises the following steps:

[0023] For each image sub-region in the text image to be recognized, obtain the target sub-text image located at the region boundary in the image sub-region.

[0024] Based on the target region position information and the target text category of the target text region where each of the target sub-text images is located, obtain the second category subordination relationship between the target sub-text images corresponding to each of the image sub-regions, wherein the second category subordination relationship comprises the category subordination relationship between entities, attributes and attribute values.

[0025] Based on the second category subordination relationship, perform data correlation processing on the first text knowledge corresponding to each of the target sub-text images in the second text knowledge, so as to obtain the target text knowledge corresponding to the text image to be recognized.

[0026] According to the text knowledge extraction method provided by the application, the target detection model is obtained by training based on the following method:

[0027] inputting a sample text image into a pre-constructed initial detection model to obtain at least one predicted text detection box and predicted region position information data and predicted text category data of a text region of interest where each of the predicted text detection boxes is located, the sample text image being a semi-structured text image;

[0028] determining a bounding box regression loss of the initial detection model based on the predicted region position information data and actual region position information data of the text region of interest where each of the predicted text detection boxes is located;

[0029] determining a text classification loss of the initial detection model based on the predicted text category data and actual text category label of the text region of interest where each of the predicted text detection boxes is located;

[0030] determining a target loss function corresponding to the initial detection model based on the bounding box regression loss and the text classification loss, and optimizing the initial detection model based on the target loss function to obtain an optimized target detection model.

[0031] According to the text knowledge extraction method provided by the application, the sample text image is input into the pre-constructed initial detection model to obtain at least one predicted text detection box and predicted region position information data and predicted text category data of a text region of interest where each of the predicted text detection boxes is located, which comprises:

[0032] determining at least one predicted text detection box from the sample text image, and performing bounding box regression processing on each of the predicted text detection boxes to obtain first region position information of a text region of interest where each of the predicted text detection boxes is located;

[0033] based on the first region position information, obtaining local image features corresponding to the text region of interest where each of the predicted text detection boxes is located from image feature data corresponding to the sample text image, and performing pooling processing on the local image features;

[0034] based on the local image features after the pooling processing, obtaining predicted text category data and predicted region position information data of the text region of interest where each of the predicted text detection boxes is located, the predicted text category data containing a class recognition probability corresponding to at least one selectable text category.

[0035] The application further provides a text knowledge extraction device, comprising:

[0036] The target detection module is configured to input the text image to be recognized into a pre-trained target detection model to obtain at least one target text detection frame, target region position information of a text region of interest where each target text detection frame is located, and a target text category.

[0037] The image acquisition module is configured to acquire a sub-text image corresponding to the text region of interest where each target text detection frame is located from the text image to be recognized based on the target region position information, and classify the sub-text image corresponding to the text region of interest where each target text detection frame is located based on the target text category to obtain at least two types of sub-text images.

[0038] The knowledge extraction module is configured to acquire target text knowledge corresponding to the text image to be recognized based on the target region position information of the text region of interest where each type of sub-text image in the at least two types of sub-text images is located, the target text category, and a corresponding knowledge extraction strategy.

[0039] The present application also provides an electronic device, comprising an image sensor, a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the image sensor is in communication connection with the processor, and is configured to acquire a text image to be recognized and transmit the text image to be recognized to the processor.

[0040] The processor is configured to input the text image to be recognized into a pre-trained target detection model to obtain at least one target text detection frame, target region position information of a text region of interest where each target text detection frame is located, and a target text category, wherein the text image to be recognized is a semi-structured text image; acquire a sub-text image corresponding to the text region of interest where each target text detection frame is located from the text image to be recognized based on the target region position information, and classify the sub-text image corresponding to the text region of interest where each target text detection frame is located based on the target text category to obtain at least two types of sub-text images; and acquire target text knowledge corresponding to the text image to be recognized based on the target region position information of the text region of interest where each type of sub-text image in the at least two types of sub-text images is located, the target text category, and a corresponding knowledge extraction strategy.

[0041] The present application also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the text knowledge extraction method according to any one of the above.

[0042] The application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the text knowledge extraction method according to any one of the above.

[0043] The application further provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the text knowledge extraction method according to any one of the above.

[0044] The text knowledge extraction method, device, electronic equipment and readable storage medium provided by the application are characterized in that: the semi-structured text image to be recognized is input into a target detection model, so that the target detection model identifies at least one target text detection frame and target region position information and a target text category of a text region of interest where each target text detection frame is located, according to the rules of the layout format of the semi-structured text image on the image distribution. Since the recognition effect of the image recognition technology is not affected by the content of books in different fields, the application is suitable for the recognition of semi-structured books in various fields, has high transferability and reusability, and the whole recognition process is automatically executed by a computer, without the need for professional knowledge in the relevant field, thereby reducing the knowledge extraction threshold. The subtext images corresponding to the text region of interest where each target text detection frame is located are classified based on the target text category, and a knowledge extraction strategy suitable for the subtext images of different text categories is adopted, so as to improve the extraction accuracy of the text knowledge, and solve the technical problems of how to improve the accuracy, transferability and reduce the knowledge extraction threshold of the knowledge extraction method in the prior art. BRIEF DESCRIPTION OF DRAWINGS

[0045] In order to more clearly illustrate the technical solutions of the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0046] Figure 1 is one of the flowcharts of the text knowledge extraction method provided by the embodiments of the application;

[0047] Figure 2 is the second flowchart of the text knowledge extraction method provided by the embodiments of the application;

[0048] Figure 3 is a schematic diagram of the target text detection frame corresponding to the text image to be recognized in the embodiments of the application;

[0049] Figure 4 is the third flowchart of the text knowledge extraction method provided by the embodiments of the application;

[0050] Figure 5 Figure 4 is a flowchart of a text knowledge extraction method according to an embodiment of the present application;

[0051] Figure 6 Figure 5 is a flowchart of a text knowledge extraction method according to an embodiment of the present application;

[0052] Figure 7 Figure 6 is a flowchart of a text knowledge extraction method according to an embodiment of the present application;

[0053] Figure 8 Figure 7 is a flowchart of a text knowledge extraction method according to an embodiment of the present application;

[0054] Figure 9 Figure 8 is a structural diagram of a text knowledge extraction device according to an embodiment of the present application;

[0055] Figure 10 Figure 9 is a structural diagram of an electronic device according to an embodiment of the present application;

[0056] Figure 11 Figure 10 is a structural diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0057] In order to make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0058] The text knowledge extraction method provided by the present application will be described below with reference to the drawings. Figures 1-8 As shown in FIG. 1, the present application provides a text knowledge extraction method, which comprises the following steps: Figure 1

[0059] In step 101, input the text image to be recognized into a pre-trained target detection model to obtain at least one target text detection box and target region position information and a target text category of each target text detection box. The text image to be recognized is a semi-structured text image.

[0060] The text image to be recognized can be a semi-structured text image in an agricultural book, or a semi-structured text image in an economic book or a scientific and technological book. The present application does not limit the specific application field, as long as the input text image to be recognized is a semi-structured text image. The semi-structured text image represents a text image with a certain regularity in layout format.​

[0061] Furthermore, the target text detection box represents the bounding box corresponding to the region of interest in the text image to be identified. The target region location information includes the region coordinates and size information of the region of interest where the target text detection box is located. The target text category includes at least one of the following: entity, attribute, attribute value, and special symbol.

[0062] Step 102: Based on the target region location information, extract the sub-text image corresponding to the region of interest where each target text detection box is located from the text image to be identified, and classify the sub-text image corresponding to the region of interest where each target text detection box is located based on the target text category to obtain at least two types of sub-text images.

[0063] Among them, at least two types of sub-text images represent sub-text images corresponding to at least two types of text categories, and text images of different types have different target text categories.

[0064] Step 103: Based on the target region location information and target text category of the text region of interest in each of the at least two types of sub-text images, as well as the corresponding knowledge extraction strategy, obtain the target text knowledge corresponding to the text image to be identified.

[0065] Specifically, the target region location information is used to associate and integrate textual knowledge within different text categories or different regions of interest based on location information. The target text category is used to determine the logical relationships between textual knowledge in different text categories.

[0066] It should be noted that since the difficulty of recognizing text knowledge varies among different text categories, it is necessary to adopt appropriate knowledge extraction strategies for sub-text images of different text categories in order to improve the extraction efficiency and accuracy of text knowledge for different text categories.

[0067] The steps 101 to 103 are used for inputting the semi-structured text image to be recognized into the target detection model, so that the target detection model recognizes at least one target text detection box and target region position information and a target text category of each target text detection box in the text region of interest according to the layout format of the semi-structured text image on the image distribution. Since the recognition effect of the image recognition technology is not affected by the content of books in different fields, the method is suitable for the recognition of semi-structured books in various fields, has high transferability and reusability, and the whole recognition process is automatically executed by the computer without the need of professional knowledge in the related field, thereby reducing the knowledge extraction threshold. The subtext images corresponding to the text region of interest of each target text detection box are classified based on the target text category, and a knowledge extraction strategy suitable for the subtext images of different text categories is used, so as to improve the accuracy of text knowledge extraction, and solve the technical problems of how to improve the accuracy, transferability and reduce the knowledge extraction threshold of the knowledge extraction method in the prior art.

[0068] In one embodiment, the at least two types of subtext images include a first subtext image and a second subtext image. As shown in Figure 2 Step 102 includes steps 201 to 202, in which:

[0069] Step 201, in the case where the target text category of the text region of interest in the target text detection box is a first text category, the subtext image corresponding to the text region of interest in the target text detection box is determined as a first subtext image, and the first text category includes special symbols.

[0070] The first subtext image represents the subtext image corresponding to the first text category. The special symbols include physical quantity units, parameter ranges, formulas, chemical equations of drugs, etc. The first text category also includes biological structure diagrams, circuit diagrams, etc. For example, Figure 3 A biological structure diagram of a wheat powdery mildew pathogen.

[0071] Step 202, in the case where the target text category of the text region of interest in the target text detection box is a second text category, the subtext image corresponding to the text region of interest in the target text detection box is determined as a second subtext image, and the second text category includes entities, attributes and attribute values.

[0072] The second subtext image represents the subtext image corresponding to the second text category. The second text category includes entities, attributes and attribute values. For example, Figure 3 “1. Wheat powdery mildew” in the above is an entity, Figure 3 “Distribution of damage” and “Cause” in the above are attributes, Figure 3The text knowledge in the text detection box on the right side of "distribution damage" and "cause" in the text is an attribute value.

[0073] Further, the second text category further includes the relationship between entities, the relationship between attributes, and the relationship between attribute values.

[0074] It should be noted that, Figure 3 The "1. Meristematic Propagating Tissue" below the biological structure diagram of "wheat powdery mildew pathogen" in the text is not an entity, but a figure annotation of the biological structure diagram, and the knowledge extraction method in the prior art is easy to misidentify "1. Meristematic Propagating Tissue" as an entity. Therefore, compared with the prior art, the text knowledge extraction method provided by the present application has the advantages of high recognition accuracy and good knowledge extraction effect.

[0075] The above steps 201 to 202 classify the subtext image corresponding to each target text detection box in the text region of interest into a subtext image corresponding to the first text category and a subtext image corresponding to the second text category, so as to adopt a knowledge extraction strategy suitable for the subtext image of different text categories in the subsequent process, so as to eliminate the technical defects of misidentifying figure annotations as entities and misidentifying special symbols such as formulas and chemical equations of drugs in the prior art, thereby further improving the text recognition accuracy and knowledge extraction effect.

[0076] In one embodiment, as Figure 4 shown, the above step 103 includes steps 301 to 304, wherein:

[0077] Step 301: based on the knowledge extraction strategy corresponding to each type of subtext image, obtaining the first text knowledge corresponding to each type of subtext image, and the knowledge extraction strategy is determined based on the target text category of the text region of interest where the subtext image is located.

[0078] In one embodiment, for the first subtext image corresponding to the first text category, the first subtext image can be directly output as the first text knowledge, or a special symbol recognition model can be used to obtain text knowledge such as formulas and chemical equations of drugs in the first subtext image.

[0079] Further, in the case of directly outputting the first subtext image as the first text knowledge, the image display confirmation information is sent to the user terminal to make the user confirm whether to display the image corresponding to the first text knowledge, and in the case of receiving the image display confirmation request of the user, the first subtext image is sent to the user terminal.

[0080] In one embodiment, for the second sub-text image corresponding to the second text category, a second text knowledge in the second sub-text image can be directly extracted by using any one of text recognition methods, wherein the text recognition methods include but are not limited to an optical character recognition method.

[0081] In step 302, based on the target region position information of the text region of interest where each sub-text image is located, at least two sub-text images in the same image sub-region in the text image to be recognized are determined.

[0082] The target region position information includes region coordinate information and region size information of the text region of interest where the target text detection box is located. The region coordinate information includes the position coordinates of at least one point in the text region of interest. The region size information includes the side length of the bounding box of the text region of interest. For example, the length and width of a rectangular text region of interest.

[0083] In step 303, for each image sub-region, based on the target region position information of the text region of interest where each sub-text image in the image sub-region is located, the target text category and the corresponding first text knowledge, the second text knowledge corresponding to the image sub-region is determined.

[0084] The target region position information is used to associate and integrate the text knowledge in different text categories or different text regions of interest according to the position information. The target text category is used to determine the logical association relationship between the text knowledge of different text categories.

[0085] In step 304, based on the second text knowledge corresponding to each image sub-region in the text image to be recognized and the sub-text image in each image sub-region, the target text knowledge corresponding to the text image to be recognized is determined.

[0086] In one embodiment, based on the position relationship and category subordination relationship of the sub-text image in each image sub-region, the second text knowledge corresponding to each image sub-region is associated and integrated, so as to obtain the target text knowledge corresponding to the text image to be recognized.

[0087] The steps 301 to 304 eliminate the technical defects of misidentifying figure annotations as entities and incorrectly identifying special symbols such as formulas, chemical formulas of drugs, and the like in the prior art, further improve the text recognition accuracy and knowledge extraction effect, determine at least two subtext images in the same image subregion, and based on the target region position information and the target text category, associate and integrate the text knowledge in each subtext image in the same image subregion, thereby quickly and accurately extracting the second text knowledge corresponding to each image subregion; based on the positional relationship and category subordination relationship of the subtext images in each image subregion, associate and integrate the second text knowledge corresponding to each image subregion, thereby quickly and accurately extracting the target text knowledge corresponding to the text image to be recognized, to avoid technical defects such as omission of text knowledge and errors in the association relationship of text knowledge, thereby further improving the extraction accuracy and efficiency of text knowledge.

[0088] In one embodiment, as shown in Figure 5 The step 303 includes steps 401 to 402, wherein:

[0089] The step 401 determines a first category subordination relationship between each subtext image in the image subregion based on the target text category of the text region of interest where each subtext image in the image subregion is located, and the first category subordination relationship includes the category subordination relationship between entities, attributes, and attribute values.

[0090] The first category subordination relationship includes the category subordination relationship between entities and attributes, the category subordination relationship between attributes and attribute values, and the category subordination relationship between entities and entities. For example, Figure 3 The entity "1. Wheat powdery mildew" in the above example has a category subordination relationship with the attribute "distribution of damage" and the attribute "cause of disease", respectively.

[0091] The step 402 performs data association processing on the first text knowledge corresponding to each subtext image in the image subregion based on the first category subordination relationship, to obtain the second text knowledge corresponding to the image subregion.

[0092] Further, the second text knowledge corresponding to the image subregion includes at least one of the text knowledge in the form of a triple of "entity, attribute, attribute value", "entity, relationship, entity", "attribute, relationship, attribute", and "attribute value, relationship, attribute value".

[0093] In one embodiment, as shown in Figure 6 The step 304 includes steps 501 to 503, wherein:

[0094] Step 501, for each image sub-region in the to-be-identified text image, obtain a target sub-text image located at the region boundary in the image sub-region.

[0095] Wherein, the target sub-text image represents a sub-text image located on the region boundary corresponding to the image sub-region.

[0096] Step 502, based on the target region position information and the target text category of the text region of interest where each target sub-text image is located, obtain the second category subordination relationship between the target sub-text images corresponding to each image sub-region, and the second category subordination relationship includes the category subordination relationship between entities, attributes and attribute values.

[0097] Wherein, the second category subordination relationship includes the category subordination relationship between entities and attributes, the category subordination relationship between attributes and attribute values, and the category subordination relationship between entities and entities. For example, Figure 3 The entity "1. Wheat powdery mildew" in the above has a category subordination relationship with the attribute "distribution of damage" and the attribute "cause of disease", respectively.

[0098] Step 503, based on the second category subordination relationship, performing data association processing on the first text knowledge corresponding to each target sub-text image in the second text knowledge, to obtain the target text knowledge corresponding to the to-be-identified text image.

[0099] Further, the target text knowledge corresponding to the to-be-identified text image includes at least one of the text knowledge in the form of "entity, attribute, attribute value", "entity, relationship, entity", "attribute, relationship, attribute" and "attribute value, relationship, attribute value".

[0100] In one embodiment, after the above step 503, the text knowledge extraction method provided by the present application further comprises:

[0101] S101, obtain a pre-set homograph dictionary library, the homograph dictionary library includes correct word forms and error word forms corresponding to special terms or rare characters that are prone to errors. Detect whether there is an error word form corresponding to a special term or a rare character in the homograph dictionary library in the target text knowledge corresponding to the to-be-identified text image.

[0102] S102, in the case where it is determined that there is an error word form corresponding to a special term or a rare character in the homograph dictionary library in the target text knowledge, modify the error word form corresponding to the special term or the rare character to the corresponding correct word form. For example, "fermentation" in agricultural books is usually misidentified as "drive fertilizer".

[0103] The above embodiment can improve the accuracy of knowledge extraction by constructing a homograph dictionary library to correct the target text knowledge corresponding to the to-be-identified text image.

[0104] In one embodiment, as shown in FIG. 1, the target detection model is trained based on the following manner: Figure 7

[0105] Step 601, inputting a sample text image into a pre-constructed initial detection model to obtain at least one predicted text detection box, and predicted region position information data and predicted text category data of a text region of interest where each predicted text detection box is located, the sample text image being a semi-structured text image.

[0106] The predicted text detection box represents a bounding box corresponding to the text region of interest in the sample text image. The predicted region position information data includes region coordinate information and region size information of the text region of interest where the predicted text detection box is located. The predicted text category data includes a predicted category recognition probability corresponding to at least one optional text category corresponding to the text region of interest where the predicted text detection box is located. The optional text category includes at least one of an entity, a relationship, an attribute, an attribute value, and a special symbol.

[0107] Step 602, determining a bounding box regression loss of the initial detection model based on the predicted region position information data and actual region position information data of the text region of interest where each predicted text detection box is located.

[0108] In one embodiment, a first difference amount between the predicted region position information data and the actual region position information data of the text region of interest where each predicted text detection box is located is obtained, and the bounding box regression loss of the initial detection model is determined based on the first difference amount. The predicted region position information data includes a predicted region center point horizontal coordinate, a predicted region center point vertical coordinate, a predicted region width, and a predicted region length of the text region of interest where the predicted text detection box is located. The actual region position information data includes an actual region center point horizontal coordinate, an actual region center point vertical coordinate, an actual region width, and an actual region length of the text region of interest where the predicted text detection box is located.

[0109] In one embodiment, the bounding box regression loss of the initial detection model can be represented by the following formula (1) to formula (2):

[0110]

[0111]

[0112] wherein L loc (t u ,v) represents the bounding box regression loss. t u represents the predicted region position information data or the predicted bounding box regression parameter, and ​represents a predicted region center point horizontal coordinate, represents a predicted region center point vertical coordinate, represents a predicted region width, represents a predicted region length.v i represents an actual bounding box regression parameter, andv i = (v x , v y , v w , v h ),v x represents an actual region center point horizontal coordinate, v y represents an actual region center point vertical coordinate, v w represents an actual region width, v h represents an actual region length. x represents a region center point horizontal coordinate, y represents a region center point vertical coordinate, w represents a region width, and h represents a region length. x can be used to replace the value ofv , and the first difference quantity between the predicted region position information data and the actual region position information data is calculated based on a smooth L1 segment function.

[0113] In step 603, text classification loss of the initial detection model is determined based on the predicted text category data of the text region of interest where each predicted text detection box is located and the actual text category label.

[0114] In one embodiment, a second difference quantity between the predicted text category data of the text region of interest where each predicted text detection box is located and the actual text category label is obtained, and text classification loss of the initial detection model is determined based on the second difference quantity. The predicted text category data includes a predicted category recognition probability corresponding to at least one selectable text category corresponding to the text region of interest where the predicted text detection box is located. The actual text category label represents a category label corresponding to the actual text category of the text region of interest where the predicted text detection box is located.

[0115] Further, the predicted text category data can be a multi-dimensional array composed of predicted category recognition probabilities corresponding to at least one selectable text category, and the number of dimensions of the multi-dimensional array is the same as the number of selectable text categories corresponding to the text region of interest where the predicted text detection box is located. For example, in the case where the number of selectable text categories is 5, the predicted text category data can be represented by a 5-dimensional array (0.9, 0.6, 0.7, 0.5, 0.3), where 0.9, 0.6, 0.7, 0.5, and 0.3 respectively represent predicted category recognition probabilities corresponding to each selectable text category.

[0116] Further, the actual text category label can be a category label composed of a multi-bit binary number or a category label composed of a multi-dimensional array. The dimension number of the actual text category label is the same as the number of selectable text categories corresponding to the text region of interest where the predicted text bounding box is located. The value of the number in the multi-bit binary number or the multi-dimensional array is 1 or 0, and the selectable text category corresponding to the number with the value of 1 is the actual text category. For example, in the case where the number of selectable text categories is 5, the category label composed of a multi-bit binary number can be represented as 10000, or can be represented as a 5-dimensional array (0, 1, 0, 0, 0). The second number 1 corresponds to the actual text category.

[0117] In one embodiment, the text classification loss of the initial detection model can be represented by the following formula (3):

[0118] L cls (p,u)=-log p u (3)

[0119] wherein L cls (p,u) represents the text classification loss, p represents the predicted text category data, and p=(p0,p1,…,pk), p0 represents the category recognition probability corresponding to the 0th selectable text category corresponding to the predicted text bounding box, p1 represents the category recognition probability corresponding to the 1st selectable text category corresponding to the predicted text bounding box, and pk represents the category recognition probability corresponding to the kth selectable text category corresponding to the predicted text bounding box. u represents the actual text category label, for example, u is 01000000, that is, the actual text category corresponding to the predicted text bounding box is the 1st selectable text category.

[0120] In step 604, based on the boundary box regression loss and the text classification loss, a target loss function corresponding to the initial detection model is determined, and the initial detection model is optimized based on the target loss function to obtain an optimized target detection model.

[0121] In one embodiment, the model parameters of the initial detection model are optimized along the direction of the gradient descent of the target loss function to obtain the optimized target detection model. Further, the target loss function corresponding to the initial detection model can be represented by the following formula (4):

[0122] L(p,u,t u ,v)=L cls (p,u)+λ[u≥1]L loc (t u ,v) (4)

[0123] wherein L(p,u,t u ,v) represents the target loss function, L cls(p, u) represents a text classification loss, L loc (t u , v) represents a bounding box regression loss. p represents predicted text category data, u represents actual text category label. t u represents predicted region position information data or predicted bounding box regression parameters, and v represents actual bounding box regression parameters. [u≥1] represents an Iverson bracket, which is a kind of square bracket notation, and is 1 if the condition in the square bracket is met, and is 0 if the condition is not met.

[0124] The steps 601 to 604 described above determine the bounding box regression loss of the initial detection model based on the predicted region position information data and the actual region position information data, and determine the text classification loss of the initial detection model based on the predicted text category data and the actual text category label, and optimize the initial detection model based on the target loss function composed of the bounding box regression loss and the text classification loss, so as to reduce the text classification loss and the bounding box regression loss of the initial detection model in the target detection process, so that the optimized target detection model can more accurately detect the position information and text category information of the text detection frame in the to-be-recognized text image, and further improve the extraction accuracy of text knowledge.

[0125] In one embodiment, as shown in Figure 8 the step 601 described above includes steps 701 to 703, wherein:

[0126] Step 701, at least one predicted text detection frame is determined from the sample text image, and each predicted text detection frame is subjected to bounding box regression processing to obtain first region position information of the text region of interest in each predicted text detection frame.

[0127] In one embodiment, the step 701 described above specifically includes the following steps: S7011, image feature data in the sample text image is extracted based on a convolutional neural network, and the image feature data is a feature vector or a multi-dimensional feature matrix.

[0128] S7012, each feature point in the image feature data is mapped to the sample text image, each feature point is taken as a center point, and at least one preselected text detection frame is drawn in the sample text image based on at least one set of length-width ratios and at least one set of region size parameters.

[0129] The at least one set of aspect ratios includes an aspect ratio of 1:1, an aspect ratio of 1:2, and an aspect ratio of 2:1. The at least one set of region size parameters includes a region size of 32*32, a region size of 128*128, and a region size of 256*256. The number of pre-selected text detection boxes is a product of the first number of sets of aspect ratios and the second number of sets of region size parameters. For example, the first number of sets of aspect ratios is 3, and the second number of sets of region size parameters is 3, and the number of pre-selected text detection boxes is 3*3=9.

[0130] S7013, performing binary classification processing and bounding box regression processing on the at least one pre-selected text detection box to obtain binary classification categories and first region position information of each pre-selected text detection box. The pre-selected text detection box with a positive binary classification category is determined as a predicted text detection box.

[0131] The binary classification category of the pre-selected text detection box includes a positive class or a negative class. The positive class contains an object, and the negative class does not contain an object. The bounding box regression processing is used to make the position information of the text detection box closer to the position information of the real bounding box. The first region position information includes a predicted region center point horizontal coordinate, a predicted region center point vertical coordinate, a predicted region width, and a predicted region length of the interested text region where the pre-selected text detection box is located.

[0132] Step 702, based on the first region position information, local image feature corresponding to the interested text region where each predicted text detection box is located is obtained from the image feature data corresponding to the sample text image, and the local image feature is subjected to pooling processing.

[0133] In one embodiment, each predicted text detection box is mapped to the image feature data corresponding to the sample text image, so as to extract the local image feature corresponding to the interested text region where the predicted text detection box is located from the image feature data based on the first region position information of the predicted text detection box. The local image feature corresponding to the interested text region where each predicted text detection box is located is divided into a plurality of sub-image features, and the plurality of sub-image features are subjected to maximum pooling processing or average pooling processing to obtain the local image feature after the pooling processing.

[0134] Step 703, based on the local image feature after the pooling processing, the predicted text category data and the predicted region position information data of the interested text region where each predicted text detection box is located are obtained. The predicted text category data includes a class recognition probability corresponding to at least one selectable text category.

[0135] In one embodiment, the local image features after the pooling processing corresponding to each predicted text detection box are classified and regressed fine-tuned based on a fully connected network and a softmax function to obtain predicted text category data and predicted region position information data of the text region of interest where each predicted text detection box is located. The predicted region position information data includes a predicted region center point horizontal coordinate, a predicted region center point vertical coordinate, a predicted region width, and a predicted region length of the text region of interest where the predicted text detection box is located.

[0136] The text knowledge extraction device provided by the present application is described below. The text knowledge extraction device described below can be referred to in correspondence with the text knowledge extraction method described above.

[0137] As shown in Figure 9 The present application provides a text knowledge extraction device. The text knowledge extraction device 100 includes:

[0138] A target detection module 101 is configured to input a to-be-recognized text image into a pre-trained target detection model to obtain at least one target text detection box, target region position information of the text region of interest where each target text detection box is located, and a target text category. The to-be-recognized text image is a semi-structured text image.

[0139] An image acquisition module 102 is configured to acquire a subtext image corresponding to the text region of interest where each target text detection box is located from the to-be-recognized text image based on the target region position information, and classify the subtext image corresponding to the text region of interest where each target text detection box is located based on the target text category to obtain at least two types of subtext images.

[0140] A knowledge extraction module 103 is configured to acquire target text knowledge corresponding to the to-be-recognized text image based on the target region position information and the target text category of the text region of interest where each type of subtext image is located and a corresponding knowledge extraction strategy.

[0141] In one embodiment, the at least two types of subtext images include a first subtext image and a second subtext image. The image acquisition module 102 is further configured to, in a case where the target text category of the text region of interest where the target text detection box is located is a first text category, determine that the subtext image corresponding to the text region of interest where the target text detection box is located is the first subtext image, and in a case where the target text category of the text region of interest where the target text detection box is located is a second text category, determine that the subtext image corresponding to the text region of interest where the target text detection box is located is the second subtext image. The first text category includes special symbols, and the second text category includes entities, attributes, and attribute values.

[0142] In an embodiment, the knowledge extraction module 103 is further configured to: obtain, based on a knowledge extraction strategy corresponding to each sub-text image of each category, first text knowledge corresponding to each sub-text image of each category, the knowledge extraction strategy being determined based on a target text category of a text region of interest in which the sub-text image is located; determine, based on target region position information of the text region of interest in which each sub-text image is located, at least two sub-text images in the to-be-recognized text image that are in a same image sub-region; for each image sub-region, determine, based on the target region position information and the target text category of the text region of interest in which each sub-text image in the image sub-region is located and the first text knowledge corresponding to each sub-text image, second text knowledge corresponding to the image sub-region; and determine, based on the second text knowledge corresponding to each image sub-region in the to-be-recognized text image and the sub-text images in each image sub-region, target text knowledge corresponding to the to-be-recognized text image.

[0143] In an embodiment, the knowledge extraction module 103 is further configured to: determine, based on the target text category of the text region of interest in which each sub-text image in the image sub-region is located, a first category subordination relationship between the sub-text images in the image sub-region, the first category subordination relationship including a category subordination relationship between entities, attributes, and attribute values; and perform data association processing on the first text knowledge corresponding to the sub-text images in the image sub-region based on the first category subordination relationship, to obtain the second text knowledge corresponding to the image sub-region.

[0144] In an embodiment, the knowledge extraction module 103 is further configured to: for each image sub-region in the to-be-recognized text image, obtain a target sub-text image located at a region boundary in the image sub-region; obtain, based on target region position information and a target text category of a text region of interest in which each target sub-text image is located, a second category subordination relationship between the target sub-text images in each image sub-region, the second category subordination relationship including a category subordination relationship between entities, attributes, and attribute values; and perform data association processing on the first text knowledge corresponding to each target sub-text image in the second text knowledge based on the second category subordination relationship, to obtain target text knowledge corresponding to the to-be-recognized text image.

[0145] In one embodiment, the text knowledge extraction device 100 includes a model training module, used to input sample text images into a pre-built initial detection model to obtain at least one predicted text detection box, as well as predicted region location information data and predicted text category data for the text region of interest where each predicted text detection box is located. The sample text image is a semi-structured text image. Based on the predicted region location information data and actual region location information data for the text region of interest where each predicted text detection box is located, the bounding box regression loss of the initial detection model is determined. Based on the predicted text category data and actual text category label for the text region of interest where each predicted text detection box is located, the text classification loss of the initial detection model is determined. Based on the bounding box regression loss and the text classification loss, the target loss function corresponding to the initial detection model is determined, and the initial detection model is optimized based on the target loss function to obtain an optimized target detection model.

[0146] In one embodiment, the model training module is further configured to determine at least one predicted text detection box from the sample text image, and perform bounding box regression processing on each predicted text detection box to obtain the first region location information of the text region of interest where each predicted text detection box is located; based on the first region location information, obtain the local image features corresponding to the text region of interest where each predicted text detection box is located from the image feature data corresponding to the sample text image, and perform pooling processing on the local image features; based on the pooled local image features, obtain the predicted text category data and the predicted region location information data of the text region of interest where each predicted text detection box is located, wherein the predicted text category data includes the category recognition probability corresponding to at least one optional text category.

[0147] Figure 10 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 10 As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830, a communication bus 840, and an image sensor 850. The processor 810, the communication interface 820, the memory 830, and the image sensor 850 communicate with each other through the communication bus 840. The image sensor 850 is communicatively connected to the processor 810 and is used to acquire the text image to be recognized and transmit the text image to be recognized to the processor 810.

[0148] The processor 810 can invoke the logic instructions in the memory 830 for inputting the text image to be recognized into a pre-trained target detection model, obtaining at least one target text detection box, target region position information of each target text detection box in the text region of interest, and a target text category, the text image to be recognized being a semi-structured text image; based on the target region position information, the sub-text image corresponding to the text region of interest where each target text detection box is located is intercepted from the text image to be recognized, and based on the target text category, the sub-text image corresponding to the text region of interest where each target text detection box is located is classified, obtaining at least two types of sub-text images; based on the target region position information and the target text category of the text region of interest where each type of sub-text image is located and the corresponding knowledge extraction strategy, the target text knowledge corresponding to the text image to be recognized is obtained.

[0149] In addition, the logic instructions in the memory 830 described above can be implemented in the form of a software function unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method of each embodiment of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0150] In one embodiment, the at least two types of sub-text images include a first sub-text image and a second sub-text image. The processor 810 is further configured to, in a case where the target text category of the text region of interest where the target text detection box is located is a first text category, determine that the sub-text image corresponding to the text region of interest where the target text detection box is located is a first sub-text image, and the first text category includes special symbols; in a case where the target text category of the text region of interest where the target text detection box is located is a second text category, determine that the sub-text image corresponding to the text region of interest where the target text detection box is located is a second sub-text image, and the second text category includes entities, attributes, and attribute values.

[0151] In an embodiment, the processor 810 is further configured to obtain the first text knowledge corresponding to each sub-text image based on a knowledge extraction strategy corresponding to each sub-text image, the knowledge extraction strategy being determined based on a target text category of the interested text region in which the sub-text image is located; determine at least two sub-text images in the same image sub-region in the to-be-recognized text image based on the target region position information of the interested text region in which each sub-text image is located; for each image sub-region, determine the second text knowledge corresponding to the image sub-region based on the target region position information and the target text category of the interested text region in which each sub-text image in the image sub-region is located and the first text knowledge corresponding to each sub-text image; and determine the target text knowledge corresponding to the to-be-recognized text image based on the second text knowledge corresponding to each image sub-region in the to-be-recognized text image and the sub-text images in each image sub-region.

[0152] In an embodiment, the processor 810 is further configured to determine a first category subordination relationship between the sub-text images in the image sub-region based on the target text category of the interested text region in which each sub-text image in the image sub-region is located, the first category subordination relationship including a category subordination relationship between entities, attributes and attribute values; and perform data association processing on the first text knowledge corresponding to the sub-text images in the image sub-region based on the first category subordination relationship, to obtain the second text knowledge corresponding to the image sub-region.

[0153] In an embodiment, the processor 810 is further configured to, for each image sub-region in the to-be-recognized text image, obtain a target sub-text image located at a region boundary in the image sub-region; obtain a second category subordination relationship between the target sub-text images in each image sub-region based on the target region position information and the target text category of the interested text region in which each target sub-text image is located, the second category subordination relationship including a category subordination relationship between entities, attributes and attribute values; and perform data association processing on the first text knowledge corresponding to each target sub-text image in the second text knowledge based on the second category subordination relationship, to obtain the target text knowledge corresponding to the to-be-recognized text image.

[0154] In an embodiment, the processor 810 is further configured to input the sample text image into the pre-constructed initial detection model to obtain at least one predicted text detection box, and predicted region position information data and predicted text category data of a text region of interest where each predicted text detection box is located, the sample text image being a semi-structured text image; determine a bounding box regression loss of the initial detection model based on the predicted region position information data and actual region position information data of the text region of interest where each predicted text detection box is located; determine a text classification loss of the initial detection model based on the predicted text category data and actual text category label of the text region of interest where each predicted text detection box is located; determine a target loss function corresponding to the initial detection model based on the bounding box regression loss and the text classification loss, and optimize the initial detection model based on the target loss function to obtain an optimized target detection model.

[0155] In an embodiment, the processor 810 is further configured to determine at least one predicted text detection box from the sample text image, and perform bounding box regression processing on each predicted text detection box to obtain first region position information of a text region of interest where each predicted text detection box is located; obtain local image features corresponding to the text region of interest where each predicted text detection box is located from image feature data corresponding to the sample text image based on the first region position information, and perform pooling processing on the local image features; obtain predicted text category data and predicted region position information data of the text region of interest where each predicted text detection box is located based on the local image features after the pooling processing, the predicted text category data containing a category recognition probability corresponding to at least one selectable text category.

[0156] Figure 11 An example of a schematic diagram of an entity structure of an electronic device is shown in FIG. 1. Figure 11As shown, the electronic device can include a processor 910, a communications interface 920, a memory 930, and a communications bus 940, wherein the processor 910, the communications interface 920, and the memory 930 complete mutual communication through the communications bus 940. The processor 910 can invoke a logic instruction in the memory 930 to execute the text knowledge extraction method provided by each method described above, which includes: inputting a to-be-recognized text image into a pre-trained target detection model to obtain at least one target text detection box and target region position information and a target text category of each target text detection box in a text region of interest, the to-be-recognized text image being a semi-structured text image; based on the target region position information, cutting a subtext image corresponding to the text region of interest in each target text detection box from the to-be-recognized text image, and based on the target text category, classifying the subtext image corresponding to the text region of interest in each target text detection box to obtain at least two types of subtext images; based on the target region position information and the target text category of the text region of interest in each type of subtext image and the corresponding knowledge extraction strategy, obtaining target text knowledge corresponding to the to-be-recognized text image.

[0157] In addition, the logic instruction in the memory 930 described above can be implemented in the form of a software function unit and sold or used as an independent product, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of each embodiment method of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), a magnetic disk or an optical disk, and various program code storage media.

[0158] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program being stored in a computer readable storage medium, and the computer program being capable of executing the text knowledge extraction method provided by the above method when executed by a processor, the method comprising: inputting a to-be-recognized text image into a pre-trained target detection model to obtain at least one target text detection box, target region position information of a text region of interest where each target text detection box is located, and a target text category; based on the target region position information, cutting a subtext image corresponding to the text region of interest where each target text detection box is located from the to-be-recognized text image, and based on the target text category, classifying the subtext image corresponding to the text region of interest where each target text detection box is located to obtain at least two types of subtext images; and based on the target region position information and the target text category of the text region of interest where each type of subtext image is located and a corresponding knowledge extraction strategy, obtaining target text knowledge corresponding to the to-be-recognized text image.

[0159] In another aspect, the present application also provides a computer readable storage medium, which stores a computer program, and the computer program is capable of implementing the text knowledge extraction method provided by the above method when executed by a processor, the method comprising: inputting a to-be-recognized text image into a pre-trained target detection model to obtain at least one target text detection box, target region position information of a text region of interest where each target text detection box is located, and a target text category; based on the target region position information, cutting a subtext image corresponding to the text region of interest where each target text detection box is located from the to-be-recognized text image, and based on the target text category, classifying the subtext image corresponding to the text region of interest where each target text detection box is located to obtain at least two types of subtext images; and based on the target region position information and the target text category of the text region of interest where each type of subtext image is located and a corresponding knowledge extraction strategy, obtaining target text knowledge corresponding to the to-be-recognized text image.

[0160] The device embodiments described above are only schematic, wherein the units shown as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment scheme. Those skilled in the art can understand and implement without creative labor.

[0161] Those skilled in the art can clearly understand the technical solutions of the various embodiments from the above description of the embodiments, and the various embodiments can be implemented by means of software with the necessary general hardware platforms, and of course can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part of the prior art that makes a contribution, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods of each embodiment or some parts of the embodiments.

[0162] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features therein; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for extracting textual knowledge, characterized in that, include: The text image to be identified is input into a pre-trained target detection model to obtain at least one target text detection box and the target region location information and target text category of the text region of interest where each target text detection box is located. The text image to be identified is a semi-structured text image. Based on the target region location information, sub-text images corresponding to the text region of interest where each target text detection box is located are extracted from the text image to be identified, and the sub-text images corresponding to the text region of interest where each target text detection box is located are classified based on the target text category to obtain at least two types of sub-text images; Based on the knowledge extraction strategy corresponding to each type of sub-text image, the first text knowledge corresponding to each type of sub-text image is obtained. The knowledge extraction strategy is determined based on the target text category of the text region of interest where the sub-text image is located. Based on the target region location information of the region of interest where each sub-text image is located, at least two sub-text images in the text image to be identified that are located within the same sub-region of the image are determined; For each of the image sub-regions, based on the target region location information and target text category of the text region of interest where each of the sub-text images in the image sub-region is located, as well as its corresponding first text knowledge, the second text knowledge corresponding to the image sub-region is determined; For each of the image sub-regions, obtain the target sub-text image located at the region boundary within the image sub-region; Based on the target region location information and target text category of each target sub-text image in the region of interest, a second category dependency relationship between the target sub-text images corresponding to each image sub-region is obtained. The second category dependency relationship includes the category dependency relationship between entities, attributes, and attribute values. Based on the second category of subordinate relationship, data association processing is performed on the first text knowledge corresponding to each target sub-text image in the second text knowledge to obtain the target text knowledge corresponding to the text image to be identified.

2. The text knowledge extraction method according to claim 1, characterized in that, The at least two types of sub-text images include a first sub-text image and a second sub-text image; The sub-text images corresponding to the text regions of interest where each target text detection box is located are classified based on the target text category to obtain at least two types of sub-text images, including: If the target text category of the text region of interest where the target text detection box is located is a first text category, then the sub-text image corresponding to the text region of interest where the target text detection box is located is determined to be the first sub-text image, and the first text category includes special symbols; If the target text category of the text region of interest where the target text detection box is located is the second text category, the sub-text image corresponding to the text region of interest where the target text detection box is located is determined as the second sub-text image, and the second text category includes entities, attributes, and attribute values.

3. The text knowledge extraction method according to claim 1, characterized in that, The step of determining the second text knowledge corresponding to the image sub-region based on the target region location information and target text category of each sub-text image in the image sub-region, and its corresponding first text knowledge, includes: Based on the target text category of the text region of interest where each sub-text image in the image sub-region is located, a first category dependency relationship between the sub-text images in the image sub-region is determined. The first category dependency relationship includes the category dependency relationship between entities, attributes, and attribute values. Based on the first category hierarchy, the first text knowledge corresponding to each sub-text image in the image sub-region is processed by data association to obtain the second text knowledge corresponding to the image sub-region.

4. The text knowledge extraction method according to any one of claims 1-3, characterized in that, The target detection model was trained in the following manner: The sample text image is input into a pre-built initial detection model to obtain at least one predicted text detection box, as well as the predicted region location information data and predicted text category data of the text region of interest where each predicted text detection box is located. The sample text image is a semi-structured text image. Based on the predicted region location information data and the actual region location information data of the text region of interest where each predicted text detection box is located, the bounding box regression loss of the initial detection model is determined. Based on the predicted text category data and the actual text category label of the region of interest where each predicted text detection box is located, the text classification loss of the initial detection model is determined; Based on the bounding box regression loss and the text classification loss, the target loss function corresponding to the initial detection model is determined, and the initial detection model is optimized based on the target loss function to obtain the optimized target detection model.

5. The text knowledge extraction method according to claim 4, characterized in that, The step of inputting the sample text image into a pre-built initial detection model to obtain at least one predicted text detection box, as well as predicted region location information data and predicted text category data for the text region of interest where each predicted text detection box is located, includes: At least one predicted text detection box is determined from the sample text image, and bounding box regression is performed on each predicted text detection box to obtain the first region location information of the text region of interest where each predicted text detection box is located. Based on the location information of the first region, local image features corresponding to the text region of interest where each predicted text detection box is located are obtained from the image feature data corresponding to the sample text image, and the local image features are pooled. Based on the local image features after pooling, the predicted text category data and the predicted region location information data of the text region of interest where each predicted text detection box is located are obtained. The predicted text category data includes the category recognition probability corresponding to at least one optional text category.

6. A text knowledge extraction device, characterized in that, include: The target detection module is used to input the text image to be identified into a pre-trained target detection model to obtain at least one target text detection box and the target region location information and target text category of the text region of interest where each target text detection box is located. The text image to be identified is a semi-structured text image. The image acquisition module is used to extract sub-text images corresponding to the text regions of interest where each target text detection box is located from the text image to be identified based on the target region location information, and to classify the sub-text images corresponding to the text regions of interest where each target text detection box is located based on the target text category, so as to obtain at least two types of sub-text images. The knowledge extraction module is used to: acquire first text knowledge corresponding to each type of sub-text image based on a knowledge extraction strategy corresponding to each type of sub-text image, wherein the knowledge extraction strategy is determined based on the target text category of the text region of interest where the sub-text image is located; determine at least two sub-text images in the text image to be identified that are located within the same sub-region based on the target region location information of the text region of interest where each sub-text image is located; for each sub-region, determine second text knowledge corresponding to the sub-region based on the target region location information and target text category of the text region of interest where each sub-text image is located, and its corresponding first text knowledge; acquire target sub-text images located at the region boundary of the sub-region for each sub-region; and acquire second category dependency relationships between target sub-text images corresponding to each sub-region based on the target region location information and target text category of the text region of interest where each target sub-text image is located, wherein the second category dependency relationship includes category dependency relationships between entities, attributes, and attribute values. Based on the second category of subordinate relationship, data association processing is performed on the first text knowledge corresponding to each target sub-text image in the second text knowledge to obtain the target text knowledge corresponding to the text image to be identified.

7. An electronic device comprising an image sensor, a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The image sensor is communicatively connected to the processor and is used to acquire images of text to be recognized and transmit the images of text to be recognized to the processor. The processor is configured to input the text image to be recognized into a pre-trained target detection model to obtain at least one target text detection box and target region location information and target text category of the text region of interest where each target text detection box is located, wherein the text image to be recognized is a semi-structured text image; based on the target region location information, it extracts sub-text images corresponding to the text region of interest where each target text detection box is located from the text image to be recognized, and classifies the sub-text images corresponding to the text region of interest where each target text detection box is located based on the target text category to obtain at least two types of sub-text images; based on the knowledge extraction strategy corresponding to each type of sub-text image, it obtains first text knowledge corresponding to each type of sub-text image, wherein the knowledge extraction strategy is determined based on the target text category of the text region of interest where the sub-text image is located; Based on the target region location information of the region of interest where each sub-text image is located, at least two sub-text images in the text image to be identified are determined to be within the same sub-region of the image. For each sub-region of the image, based on the target region location information and target text category of the region of interest where each sub-text image is located in the sub-region of the image, as well as its corresponding first text knowledge, the second text knowledge corresponding to the sub-region of the image is determined. For each sub-region of the image, the target sub-text image located at the region boundary of the sub-region of the image is obtained. Based on the target region location information and target text category of the region of interest where each target sub-text image is located, the second category dependency relationship between the target sub-text images corresponding to each sub-region of the image is obtained. The second category dependency relationship includes the category dependency relationship between entities, attributes, and attribute values. Based on the second category of subordinate relationship, data association processing is performed on the first text knowledge corresponding to each target sub-text image in the second text knowledge to obtain the target text knowledge corresponding to the text image to be identified.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the text knowledge extraction method as described in any one of claims 1 to 5.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the text knowledge extraction method as described in any one of claims 1 to 5.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the text knowledge extraction method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Image processing method and device, medium and computing equipment

    CN109726661A

  • Text detection method and device and recognition system

    CN111027563A

  • Bill image identification method, device and equipment, and storage medium

    CN111709339A