Document knowledge extraction method and device, computer equipment and storage medium
By performing structured processing and entity recognition on biomedical PDF documents, the problem of low accuracy in knowledge extraction from biomedical PDF documents in the existing technology is solved, and efficient and accurate information extraction is achieved.
Patent Information
- Application Number
- CN202410337443.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-23
- Publication Date
- 2025-09-23
AI Technical Summary
Existing knowledge extraction schemes have low accuracy in biomedical PDF documents and cannot meet structured requirements, especially for PDF documents containing flowcharts and tables.
Hypertext Markup Language (HTML) is used to structure biomedical PDF documents, extract semantically relevant element information, and normalize it into effective model features. The semantic type of text regions is determined using a probabilistic graphical model, and knowledge is extracted in combination with entity recognition technology.
The accuracy of knowledge extraction from biomedical PDF documents is improved, achieving efficient and accurate information extraction.
Smart Images

Figure CN120688500A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of artificial intelligence, and in particular to a document knowledge extraction method, apparatus, computer equipment, and storage medium. Background Art
[0002] Clinical guidelines in biomedical data are crucial reference materials in the medical field, providing guidance and recommendations for doctors and other healthcare professionals in diagnosing, treating, and preventing diseases. With the continuous advancement of diagnostic and treatment technologies and drug development methods, new treatments are emerging, and the frequency of guideline updates has accelerated. As a result, the publicly available clinical guideline texts are numerous and complex, with significant differences between versions, making it difficult to quickly and accurately obtain relevant information. Therefore, efficiently and accurately extracting relevant information from biomedical data, including clinical guidelines, is a pressing issue.
[0003] However, existing knowledge extraction solutions generally target text, while clinical guidelines and some biomedical data are mostly stored in PDF format and contain a large number of flow charts, tables, and other content. Direct extraction using generic text methods results in very low accuracy. For example, extracting information based on rules (such as text aggregation, font size, and position) has low accuracy. While statistical model-based solutions offer higher accuracy than the aforementioned solutions, currently available statistical models are not designed with features specific to biomedical PDF documents and therefore cannot meet the structural requirements of biomedical PDFs.
[0004] It can be seen that the existing neural network models are not trained for biomedical PDF documents. Therefore, when applied to biomedical PDF documents for knowledge extraction, the accuracy is low and cannot meet the needs of this field. Summary of the Invention
[0005] The purpose of the present invention is to provide a document knowledge extraction method, apparatus, computer equipment and storage medium for improving the accuracy of knowledge extraction by structurally processing biomedical PDF documents, and even facilitating downstream task processing.
[0006] In a first aspect, the present invention provides a document knowledge extraction method, comprising: Get the biomedical PDF documents to be processed; Based on Hypertext Markup Language, the biomedical PDF document is structured to obtain the target document; Obtain the valid target document corresponding to the target document to perform entity recognition and extraction on the valid target document to obtain target knowledge information.
[0007] In some embodiments of the present invention, a biomedical PDF document is structured based on hypertext markup language to obtain a target document, including: extracting various semantically related element information in the biomedical PDF document; wherein the element information includes at least one of text content, text size, text position, line direction, and line position; normalizing the original value of the element information to a preset numerical range to convert the element information into a valid model feature; inputting the valid model feature into a trained probabilistic graphical model so that the trained probabilistic graphical model analyzes and determines the semantic type of each text area and outputs the target document.
[0008] In some embodiments of the present invention, the original value of the element information is normalized to a preset numerical range to convert the element information into a valid model feature, including: normalizing the original value of the element information to a preset numerical range to convert the element information into text shape features, text position features, text content features and graphic features; determining the text shape features, text position features, text content features and graphic features as valid model features.
[0009] In some embodiments of the present invention, the original value of the element information is normalized to a preset numerical range to convert the element information into text shape features, text position features, text content features and graphic features, including: normalizing the original value of the element information to a preset numerical range according to the preset font size item, font color item, and whether it is bold, to obtain the text shape features; and normalizing the original value of the element information to a preset numerical range according to the preset text position item, first line indent item, last line indent item, horizontal alignment item, and vertical alignment item, to obtain the text position features; and normalizing the original value of the element information to a preset numerical range according to the probability of each word appearing in the title, abstract, and text structure, to obtain the text content features; and normalizing the original value of the element information to a preset numerical range according to the preset graphic position item, whether it contains a straight line item, whether it contains a slash item, whether it contains an arrow item, whether there are text items around it, and the graphic color item, to obtain the graphic features.
[0010] In some embodiments of the present invention, before inputting effective model features into a trained probabilistic graph model so that the trained probabilistic graph model analyzes and determines the semantic type of each text region and outputs a target document, it also includes: obtaining a PDF document set, the PDF document set including multiple biomedical PDF documents with text regions labeled with semantic types; the semantic types include at least one of title, abstract, text, picture, table, and reference; dividing the PDF document set into a training set and a test set; using the training set to perform preliminary training on the initial probabilistic graph model to obtain a preliminarily trained probabilistic graph model; using the test set to test and adjust the preliminarily trained probabilistic graph model to obtain a trained probabilistic graph model.
[0011] In some embodiments of the present invention, a valid target document corresponding to a target document is obtained to perform entity recognition and extraction on the valid target document to obtain target knowledge information, including: sending the target document to a client to obtain feedback from the client on a semantic type correction result for the target document; determining a valid target document based on the semantic type correction result; and performing entity recognition and extraction on the valid target document based on entity recognition technology to obtain target knowledge information.
[0012] In some embodiments of the present invention, a valid target document is determined based on the semantic type correction result, including: if the semantic type correction result is empty, determining the target document as a valid target document; if the semantic type correction result is not empty, updating the semantic type of each region in the target document based on the semantic type correction result to obtain a valid target document.
[0013] In a second aspect, the present invention provides a document knowledge extraction device, comprising: Document acquisition module, used to obtain biomedical PDF documents to be processed; The document processing module is used to perform structured processing on the PDF document based on Hypertext Markup Language to obtain the target document; The knowledge extraction module is used to obtain the valid target document corresponding to the target document, so as to perform entity recognition and extraction on the valid target document and obtain target knowledge information.
[0014] In a third aspect, the present invention further provides a computer device, comprising: one or more processors; a memory; and one or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the above-mentioned document knowledge extraction method.
[0015] In a fourth aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, which is loaded by a processor to execute the steps in the document knowledge extraction method.
[0016] In a fifth aspect, an embodiment of the present invention provides a computer program product or computer program, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method provided in the first aspect.
[0017] The above-mentioned document knowledge extraction method, device, computer equipment and storage medium, the server obtains the biomedical PDF document to be processed and uses the hypertext markup language format as the carrier for structuring the PDF document to obtain the text, graphics and other information therein as features, thereby realizing the semantic structuring of the document, thereby facilitating efficient and accurate knowledge extraction from the structured document, and ultimately improving the knowledge extraction accuracy of the biomedical PDF document. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0019] Figure 1 Schematic diagram of a scenario of a document knowledge extraction method according to an embodiment of the present invention; Figure 2 Schematic diagram of the process of document knowledge extraction method in an embodiment of the present invention; Figure 3 Schematic diagram of the structure of the document knowledge extraction device in an embodiment of the present invention; Figure 4 Schematic diagram of the structure of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION
[0020] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0021] It should be noted that, in the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, "plurality" means two or more, unless otherwise specifically defined.
[0022] At the same time, the document knowledge extraction method provided by the embodiment of the present application can be applied to Figure 1In the document knowledge extraction system shown in FIG. , the document knowledge extraction system includes a client 102 and a server 104. The client 102 can be a device that includes both receiving and transmitting hardware, that is, a device with receiving and transmitting hardware that can perform two-way communication on a two-way communication link. Such devices may include: cellular or other communication devices, which have a single-line display or a multi-line display. The client 102 can specifically be a desktop client or a mobile client, and the client 102 can specifically be one of a mobile phone, a tablet computer, and a laptop computer. The server 104 can be an independent server, or a server network or server cluster composed of servers, including but not limited to computers, network hosts, single network servers, multiple network server sets, or cloud servers composed of multiple servers. The cloud server is composed of a large number of computers or network servers based on cloud computing (Cloud Computing). In addition, a communication connection is established between the client 102 and the server 104 through a network, and the network can specifically be any one of a wide area network, a local area network, and a metropolitan area network.
[0023] In addition, those skilled in the art will understand that Figure 1 The application environment shown in the figure is only one application scenario applicable to the present application solution and does not constitute a limitation on the application scenario of the present application solution. Other application environments may also include Figure 1 More or fewer devices are shown in . For example, Figure 1 Only one server is shown. It is understandable that the document knowledge extraction system may also include one or more other devices, which are not specifically limited here. In addition, the document knowledge extraction system may also include a memory for storing data, such as storing biomedical PDF documents.
[0024] certainly, Figure 1 The scenario diagram of the document knowledge extraction system shown is only an example. The document knowledge extraction system and scenario described in the embodiment of the present invention are intended to more clearly illustrate the technical solution of the embodiment of the present invention, and do not constitute a limitation on the technical solution provided by the embodiment of the present invention. Ordinary technicians in this field can know that with the evolution of the document knowledge extraction system and the emergence of new business scenarios, the technical solution provided by the embodiment of the present invention is also applicable to similar technical problems.
[0025] See Figure 2 , is a flow chart of a document knowledge extraction method provided by an embodiment of the present invention. This embodiment mainly applies this method to the above Figure 1 Taking the server 104 in the example, the method includes steps S201 to S203, which are specifically as follows: S201, obtaining a biomedical PDF document to be processed.
[0026] Among them, biomedical PDF documents can be obtained in a compliant manner in the embodiments of this application, including but not limited to: clinical guidelines, scientific research papers, research reports, journal articles, etc.
[0027] In a specific implementation, the server 104 can obtain the biomedical PDF document to be processed through the client 102 or other devices. The biomedical PDF document can be obtained in any of the following ways: 1. In a common network structure, the server 104 can receive the biomedical PDF document from the client 102 or other cloud devices with a network connection; 2. In a pre-set blockchain network, the server 104 can synchronously obtain the biomedical PDF document from other client nodes or server nodes. The blockchain network can be a public chain, a private chain, etc.; 3. In a pre-set tree structure, the server 104 can request the biomedical PDF document from an upper-level server or poll the biomedical PDF document from a lower-level server. Therefore, the specific method of obtaining the biomedical PDF document is not limited in the embodiment of the present invention, and the acquisition time can be periodic or random.
[0028] S202 , performing structured processing on the biomedical PDF document based on Hypertext Markup Language to obtain a target document.
[0029] Among them, Hypertext Markup Language (HTML) is a markup language used to create the structure and content of web pages. It uses a series of tags to describe how text, images, links and other media resources are presented on a web page.
[0030] In a specific implementation, in order to improve the accuracy of knowledge extraction from biomedical PDF documents, an embodiment of the present application proposes that the HTML format can be used as a semantically structured carrier to convert biomedical PDF documents into an HTML format with a semantic structure, so that it is quickly known which part of the text belongs to the title, abstract, table, etc., and then re-integrated into a structured document that is convenient for knowledge extraction as the target document. Taking a clinical guideline PDF document as an example, since the diseases that clinical guidelines mainly explore are generally written in the title, the title position should be the first choice for extracting the "disease" entity. If the attributes of the text in each area are not clarified through structuring and a full-text extraction strategy is adopted, it is possible to extract error items that cause interference. For example, the secondary diseases or related diseases extracted in the "Introduction" section are definitely not the focus of this guideline compared to the main diseases extracted in the "Title" section, and there is a deviation in knowledge extraction. Therefore, it is clear which part of the text in the biomedical PDF document belongs to the title, which part belongs to the abstract, which part belongs to the table, etc., so that subsequent entity extraction of any type will be more efficient and accurate. In addition, structuring PDF documents into a format that is easier to identify and extract entities can also improve knowledge extraction efficiency. The specific method of obtaining the target document will be explained in detail below.
[0031] In one embodiment, step S202 includes: extracting various semantically related element information in a biomedical PDF document; wherein the element information includes at least one of text content, text size, text position, line direction, and line position; normalizing the original value of the element information to a preset numerical range to convert the element information into a valid model feature; inputting the valid model feature into a trained probabilistic graph model so that the trained probabilistic graph model analyzes and determines the semantic type of each text area and outputs a target document.
[0032] Here, text content can refer to any content in the biomedical field (e.g., drug information, disease information, etc.). Text size can refer to the font size (e.g., size 4, size 5, etc.). Line direction can refer to the spatial position of a line in a PDF document, including but not limited to horizontal, vertical, or other.
[0033] The text position can be information represented by one or more "{top, left, width, height}" quadruplets after normalizing the length and width of the PDF document page to [0, 100]. Taking two preset edges (the top edge of the PDF document and the left edge of the PDF document) and the intersection of the two edges as the origin "0" as an example, "top" represents the distance between the text box and the top edge of the PDF document, "left" represents the distance between the text box and the left edge of the PDF document, "width" represents the width of the text box, and "height" represents the height of the text box. Furthermore, the representation of line position can be similar to the representation of text position, and can also be represented by the coordinates of the upper left corner and the lower right corner, which is not limited in the specific embodiments of this application. Furthermore, the normalized length and width values described above are not limited to [0, 100], and can also be [0, 1000], [0, 10], etc., which is not limited in the specific embodiments of this application.
[0034] In a specific implementation, to obtain a target document with a semantic structure, server 104 may first extract various semantically relevant element information within the PDF, including but not limited to at least one of text content, text size, text position, line direction, and line position. This element information is then converted into features usable by a probabilistic graphical model (hereinafter referred to as a CRF model) (here, the feature range may be [0, 1]), thereby obtaining effective model features. The effective model features are then input into the CRF model, which then generates the target document.
[0035] Specifically, the CRF model, short for Conditional Random Field, is commonly used in sequence labeling tasks such as named entity recognition and part-of-speech tagging. The steps for normalizing the raw values of element information to a preset range to obtain valid model features are detailed below.
[0036] In one embodiment, the step of normalizing the original value of the element information to a preset numerical range to convert the element information into a valid model feature includes: normalizing the original value of the element information to a preset numerical range to convert the element information into text shape features, text position features, text content features and graphic features; determining the text shape features, text position features, text content features and graphic features as valid model features.
[0037] In a specific implementation, server 104 may employ a minimum-maximum scaling method to normalize the raw values of the element information to a preset numerical range. Specifically, the minimum (min) and maximum (max) values in the data are first found, denoted as X_{min} and X_{max}, respectively. Then, for each raw value X_i, normalization is performed using the following formula: [X_{norm} = \frac{X_i - X_{min}}{X_{max} - X_{min}}], thereby representing each feature within the numerical range [0, 1]. This description only illustrates how to perform feature conversion; the specific content represented by each feature will be explained in detail below.
[0038] In one embodiment, the steps of normalizing the original value of element information to a preset numerical range to convert the element information into text shape features, text position features, text content features and graphic features include: normalizing the original value of the element information to a preset numerical range according to a preset font size item, a font color item and a bold item to obtain text shape features; and normalizing the original value of the element information to a preset numerical range according to a preset text position item, a first line indent item, an end line indent item, a horizontal alignment item and a vertical alignment item to obtain text position features; and normalizing the original value of the element information to a preset numerical range according to the probability of each word appearing in the title, abstract and text structure to obtain text content features; and normalizing the original value of the element information to a preset numerical range according to a preset graphic position item, whether it contains a straight line item, whether it contains a slash item, whether it contains an arrow item, whether there are text items around it and a graphic color item to obtain graphic features.
[0039] Among them, the preset numerical range can be the range of any numerical value. For example, if the preset numerical range is [0, 1], when the value is 0, it means that this feature reaches the minimum degree or does not exist in a certain aspect. When the value is 1, it means that this feature reaches the maximum degree or completely exists in a certain aspect.
[0040] In a specific implementation, the server 104 can normalize the original value of the element information to a preset value range based on the preset font size item, font color item, and whether it is bold item to obtain the text shape feature, so as to be able to distinguish the semantic category of the text in each area of the PDF document later. For example, the title of a general article will use black bold font size 1 or 2, so the text shape feature of the "title" is represented by the following information: font size item [1], font color item [1], whether it is bold item [1]. In this way, if the text shape feature that meets the conditions can be detected, it can be determined that the monitored object is the "title" of the PDF document.
[0041] Furthermore, the server 104 can also normalize the original value of the element information to a preset value range based on the preset text position item, line start indent item, line end indent item, horizontal alignment item, and vertical alignment item to obtain text position features. The purpose is also to be able to subsequently distinguish the semantic categories of text in various areas of the PDF document. For example, a paragraph of text is generally located in the middle of the page, with the first line having a line start indent and the last line having a line end indent. In this way, if a text position feature that meets this condition can be detected, it can be reversely determined that the monitored object is a "natural paragraph" of the PDF document.
[0042] Furthermore, server 104 can also normalize the original value of the element information to a preset value range based on the probability of each word appearing in the title, abstract, and body structure to obtain text content features. For example, the word "abstract" has a very high probability of appearing in the abstract. In this way, if a text content feature that meets this condition is detected, it can be reversely determined that the monitored object is the "abstract" of the PDF document.
[0043] Furthermore, server 104 may also normalize the original value of the element information to a preset numerical range based on preset graphic position items, whether it contains straight lines, whether it contains diagonal lines, whether it contains arrows, whether it contains surrounding text, and graphic color items to obtain graphic features. For example, a graphic containing only vertical and horizontal straight lines without arrows is most likely a table, while a graphic containing diagonal lines and arrows is most likely a graph. In this way, if a graphic feature that meets these conditions can be detected, it can be determined that the monitored object is a "picture" or "table" in the PDF document.
[0044] In one embodiment, before inputting effective model features into a trained probabilistic graphical model so that the trained probabilistic graphical model analyzes and determines the semantic type of each text region and outputs a target document, the method further includes: obtaining a PDF document set, the PDF document set including multiple biomedical PDF documents with text regions labeled with semantic types; the semantic types include at least one of a title, an abstract, a body, a picture, a table, and a reference; dividing the PDF document set into a training set and a test set; using the training set to perform preliminary training on the initial probabilistic graphical model to obtain a preliminarily trained probabilistic graphical model; and using the test set to test and adjust the preliminarily trained probabilistic graphical model to obtain a trained probabilistic graphical model.
[0045] The initial probability graph model can be an open-source CRF model that has not been specifically trained. The characteristics of the CRF model have been briefly described above and will not be repeated here.
[0046] In a specific implementation, before server 104 retrieves the target document, it must first obtain a trained probabilistic graphical model. To do so, it must first invoke the initial probabilistic graphical model and then obtain a set of PDF documents used to train the model. This set of PDF documents can be sampled from all publicly available clinical guidelines and could contain thousands or even more documents. After obtaining the PDF document set, it can be annotated, assigning semantic categories to regions within the PDF document.
[0047] Furthermore, after completing the collection of the PDF document set, the server 104 can divide the PDF document set into a training set and a test set according to a preset ratio (such as the preset ratio is that the training set accounts for "8" and the test set accounts for "2"), and then use the training set to train the initial overview map model, and then use the test set to test the model. After the training is completed, the trained overview map model can be obtained.
[0048] It should be noted that the stopping conditions for model training include but are not limited to: 1. The error is less than a predetermined minimum value. 2. The change in weight between two iterations is minimal. A threshold can be set, and training stops when it falls below this threshold. 3. A maximum number of iterations is set, and training stops when the maximum number of iterations is exceeded, for example, "273 cycles." 4. The recognition accuracy reaches a predetermined maximum value.
[0049] S203, obtaining a valid target document corresponding to the target document, performing entity recognition and extraction on the valid target document, and obtaining target knowledge information.
[0050] In specific implementations, efficiently and accurately extracting target knowledge from biomedical PDF documents requires more than simply structuring the biomedical PDF document to obtain the target document. The target document must also be validated to improve the reliability of the semantic categories within each region of the document. Therefore, server 104 can further obtain a valid target document corresponding to the target document through a pre-defined method. Once a reliable valid target document is obtained, entity recognition technology can be used to perform targeted entity recognition on the target knowledge to accurately extract the target knowledge information. The method for obtaining a valid target document through a pre-defined method and then obtaining the target knowledge information will be described in detail below.
[0051] In one embodiment, step S203 includes: sending the target document to the client to obtain feedback from the client on the semantic type correction result of the target document; determining the valid target document based on the semantic type correction result; and performing entity recognition and extraction on the valid target document based on entity recognition technology to obtain target knowledge information.
[0052] In a specific implementation, the server 104 can obtain a valid target document by sending the target document to the client 102, so that the staff can receive and correct the target document through the client 102, and then feedback the semantic type correction results. Then, some open source entity recognition models, such as the BERT model, the CRF model, and the BiLSTM-CRF model, can be used to perform entity recognition and extraction on the valid target document for the target object, thereby obtaining the target knowledge information.
[0053] In one embodiment, the step of determining a valid target document based on the semantic type correction result includes: if the semantic type correction result is empty, determining the target document as a valid target document; if the semantic type correction result is not empty, updating the semantic type of each region in the target document based on the semantic type correction result to obtain a valid target document.
[0054] In a specific implementation, server 104 can determine whether a target document is valid after receiving the semantic type correction result from client 102. If the semantic type correction result is null, this means that the semantic category determinations for each region of the target document sent to client 102 in the previous step were correct. No correction is required, and the target document can be directly determined as valid. Conversely, if the semantic type correction result is not null, this means that the semantic category determinations for each region of the target document sent to client 102 in the previous step were incorrect. Server 104 must correct the corresponding region category of the target document based on the semantic type correction result before the document can be considered valid.
[0055] For example, for a biomedical PDF document, if no "abstract" is recognized after structured processing, then the first paragraph of text judged as "main text" may be "abstract". At this time, its semantic type correction result is not empty, and the semantic type correction result corresponding to the error area should be "abstract" rather than "main text"; or, if a document name containing a person's name is detected after "reference", it should be "reference". If the semantic type is identified as other, it should be corrected according to the semantic type correction result.
[0056] In the document knowledge extraction method in the above embodiment, the server obtains the biomedical PDF document to be processed and uses the hypertext markup language format as the carrier for structuring the PDF document to obtain the text, graphics and other information therein as features, thereby realizing the semantic structuring of the document, thereby facilitating efficient and accurate knowledge extraction from the structured document, and ultimately improving the accuracy of knowledge extraction of biomedical PDF documents.
[0057] It should be understood that although Figure 2The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 2 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0058] In order to better implement the document knowledge extraction method provided in the embodiment of the present application, based on the document knowledge extraction method proposed in the embodiment of the present application, the embodiment of the present application also provides a document knowledge extraction device, such as Figure 3 As shown, the document knowledge extraction device 300 includes: The document acquisition module 310 is used to acquire the biomedical PDF document to be processed; The document processing module 320 is used to perform structured processing on the PDF document based on Hypertext Markup Language to obtain a target document; The knowledge extraction module 330 is used to obtain a valid target document corresponding to the target document, perform entity recognition and extraction on the valid target document, and obtain target knowledge information.
[0059] In one embodiment, the document processing module 320 is also used to extract various element information related to semantics in a biomedical PDF document; wherein the element information includes at least one of text content, text size, text position, line direction, and line position; normalize the original value of the element information to a preset numerical range to convert the element information into a valid model feature; input the valid model feature into a trained probabilistic graph model so that the trained probabilistic graph model analyzes and determines the semantic type of each text area and outputs the target document.
[0060] In one embodiment, the document processing module 320 is also used to normalize the original value of the element information to a preset numerical range to convert the element information into text shape features, text position features, text content features and graphic features; and determine the text shape features, text position features, text content features and graphic features as valid model features.
[0061] In one embodiment, the document processing module 320 is also used to normalize the original value of the element information within a preset numerical range based on the preset font size item, font color item, and whether it is bold, so as to obtain text shape features; and to normalize the original value of the element information within a preset numerical range based on the preset text position item, line start indent item, line end indent item, horizontal alignment item, and vertical alignment item, so as to obtain text position features; and to normalize the original value of the element information within a preset numerical range based on the probability of each word appearing in the title, abstract, and text structure, so as to obtain text content features; and to normalize the original value of the element information within a preset numerical range based on the preset graphic position item, whether it contains a straight line item, whether it contains a slash item, whether it contains an arrow item, whether there are text items around it, and graphic color item, so as to obtain graphic features.
[0062] In one embodiment, the document processing module 320 is also used to obtain a PDF document set, which includes a plurality of biomedical PDF documents with semantic types marked in text areas; the semantic types include at least one of a title, an abstract, a body, a picture, a table, and a reference; the PDF document set is divided into a training set and a test set; the initial probability graph model is preliminarily trained using the training set to obtain a preliminarily trained probability graph model; and the preliminarily trained probability graph model is tested and adjusted using the test set to obtain a trained probability graph model.
[0063] In one embodiment, the knowledge extraction module 330 is also used to send the target document to the client to obtain the client's feedback on the semantic type correction result of the target document; based on the semantic type correction result, determine the valid target document; based on entity recognition technology, perform entity recognition and extraction on the valid target document to obtain target knowledge information.
[0064] In one embodiment, the knowledge extraction module 330 is also used to determine that the target document is a valid target document if the semantic type correction result is empty; if the semantic type correction result is not empty, the semantic type of each region in the target document is updated based on the semantic type correction result to obtain a valid target document.
[0065] In the above embodiment, the server obtains the biomedical PDF document to be processed and uses the hypertext markup language format as the carrier for structuring the PDF document to obtain the text, graphics and other information therein as features, thereby realizing the semantic structuring of the document, thereby facilitating efficient and accurate knowledge extraction from the structured document, and ultimately improving the accuracy of knowledge extraction of biomedical PDF documents.
[0066] It should be noted that the specific limitations of the document knowledge extraction device can be found in the limitations of the document knowledge extraction method above, and will not be repeated here. The various modules in the above-mentioned document knowledge extraction device can be implemented in whole or in part by software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the electronic device in the form of hardware, or can be stored in the memory of the electronic device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0067] In some embodiments of the present application, the document knowledge extraction device 300 can be implemented in the form of a computer program. The computer program can be used in Figure 4 The computer device shown in FIG. 3 is run on the computer device shown in FIG. The memory of the computer device can store various program modules constituting the document knowledge extraction device 300, such as: Figure 3 The document acquisition module 310, document processing module 320 and knowledge extraction module 330 shown in the figure; the computer program composed of each program module enables the processor to execute the steps of the document knowledge extraction method of each embodiment of the present application described in this specification. For example, Figure 4 The computer device shown can be Figure 3 The document acquisition module 310 in the document knowledge extraction device 300 shown executes step S201. The computer device can execute step S202 through the document processing module 320. The computer device can execute step S203 through the knowledge extraction module 330. The computer device includes a processor, a memory and a network interface connected through a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external computer device through a network connection. When the computer program is executed by the processor, a document knowledge extraction method is implemented.
[0068] Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0069] In some embodiments of the present application, a computer device is provided, comprising one or more processors; a memory; and one or more application programs, wherein the one or more application programs are stored in the memory and configured to cause the processor to execute the steps of the document knowledge extraction method described above. The steps of the document knowledge extraction method herein may be the steps of the document knowledge extraction method described in the various embodiments described above.
[0070] In some embodiments of the present application, a computer-readable storage medium is provided, storing a computer program, which is loaded by a processor, causing the processor to execute the steps of the document knowledge extraction method described above. The steps of the document knowledge extraction method here can be the steps of the document knowledge extraction and analysis method described in each of the above embodiments.
[0071] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media used in the embodiments provided herein may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0072] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0073] The above is a detailed introduction to a document knowledge extraction method, device, computer equipment and storage medium provided in the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A document knowledge extraction method, characterized in that: include: Get the biomedical PDF documents to be processed; Based on Hypertext Markup Language, the biomedical PDF document is structured to obtain a target document; Obtain a valid target document corresponding to the target document, perform entity recognition and extraction on the valid target document, and obtain target knowledge information.
2. The method according to claim 1, wherein The method of performing structured processing on the biomedical PDF document based on the hypertext markup language to obtain a target document includes: Extracting various semantically relevant element information from the biomedical PDF document; wherein the element information includes at least one of text content, text size, text position, line direction, and line position; Normalizing the original value of the element information to a preset value range to convert the element information into a valid model feature; The effective model features are input into a trained probabilistic graphical model, so that the trained probabilistic graphical model analyzes and determines the semantic type of each text region and outputs the target document.
3. The method according to claim 2, wherein Normalizing the original value of the element information to a preset value range to convert the element information into a valid model feature includes: Normalizing the original value of the element information to a preset value range to convert the element information into text shape features, text position features, text content features, and graphic features; The character shape feature, the character position feature, the character content feature and the graphic feature are determined as the effective model features.
4. The method according to claim 3, wherein Normalizing the original value of the element information to a preset value range to convert the element information into text shape features, text position features, text content features, and graphic features includes: Normalizing the original value of the element information to a preset value range according to the preset font size item, font color item, and whether it is bold or not, to obtain the text shape feature; and Normalizing the original value of the element information to a preset value range according to the preset text position item, line start indent item, line end indent item, horizontal alignment item, and vertical alignment item to obtain the text position feature; and Normalizing the original value of the element information to a preset value range based on the probability of each word appearing in the title, abstract, and body structure to obtain the text content feature; and According to the preset graphic position item, whether it contains a straight line item, whether it contains a slash item, whether it contains an arrow item, whether there are text items around it, and the graphic color item, the original value of the element information is normalized to a preset numerical range to obtain the graphic feature.
5. The method according to any one of claims 2 to 4, characterized in that Before inputting the effective model features into the trained probabilistic graphical model so that the trained probabilistic graphical model analyzes and determines the semantic type of each text region and outputting the target document, the method further includes: Obtaining a PDF document set, the PDF document set comprising a plurality of biomedical PDF documents in which text regions are annotated with semantic types; the semantic types comprising at least one of a title, an abstract, a body of text, a picture, a table, and a reference; Dividing the PDF document set into a training set and a test set; Performing preliminary training on the initial probability graph model using the training set to obtain a preliminary trained probability graph model; The test set is used to test and adjust the probabilistic graphical model after the preliminary training to obtain the trained probabilistic graphical model.
6. The method according to claim 1, wherein The obtaining of a valid target document corresponding to the target document to perform entity recognition and extraction on the valid target document to obtain target knowledge information includes: Sending the target document to a client to obtain feedback from the client on a semantic type correction result of the target document; Determining the valid target document based on the semantic type correction result; Based on entity recognition technology, entity recognition and extraction are performed on the valid target document to obtain the target knowledge information.
7. The method according to claim 6, wherein The determining the valid target document based on the semantic type correction result includes: If the semantic type correction result is empty, determining that the target document is the valid target document; If the semantic type correction result is not empty, the semantic type of each region in the target document is updated based on the semantic type correction result to obtain the valid target document.
8. A document knowledge extraction device, characterized in that: include: Document acquisition module, used to obtain biomedical PDF documents to be processed; A document processing module, configured to perform structured processing on the PDF document based on Hypertext Markup Language to obtain a target document; The knowledge extraction module is used to obtain the valid target document corresponding to the target document, so as to perform entity recognition and extraction on the valid target document to obtain target knowledge information.
9. A computer device, characterized in that: The computer device comprises: one or more processors; A memory; and one or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the document knowledge extraction method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps of the document knowledge extraction method according to any one of claims 1 to 7.