File structured information extraction method, device, equipment, medium and product

By obtaining the file content type and constructing content entity relationship data for structured information extraction, the problem of accuracy and completeness in extracting structured information from multiple types of files is solved, and efficient structured information processing of text and image files is achieved.

CN120849649BActive Publication Date: 2026-02-06HANGZHOU HAOLINK INTELLIGENT TECHNOLOGY CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511358445.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2026-02-06
Estimated Expiration
2045-09-23

AI Technical Summary

Technical Problem

In existing technologies, the extraction of structured information from multiple types of documents lacks accuracy and completeness, and fails to effectively consider the correlation between texts.

Method used

By obtaining the file content type of the file to be processed, text recognition and structured content entity recognition are performed to construct content entity relationship data. The content entity relationship data is then used to extract structured information, including the processing of text content files and image content files.

Benefits of technology

It improves the accuracy and completeness of structured information extraction, expands the scope of application, and meets the needs of structured information extraction from multimodal documents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849649B_ABST
    Figure CN120849649B_ABST
Patent Text Reader

Abstract

The application discloses a kind of extraction methods, device, equipment, medium and product of file structured information, it is related to data processing technical field, comprising: determining the file content type of to-be-processed file;In the case where it is determined that file content type is image content file, to-be-processed file is carried out text recognition, and the text area coordinates of to-be-processed text contained in to-be-processed file and to-be-processed text in to-be-processed file are determined;To-be-processed text is carried out structured content entity identification, and the content entity coordinates of each structured content entity in text area coordinates respectively corresponding to structured content entity contained in to-be-processed text are determined;According to each content entity coordinates, the content entity relationship data between each structured content entity is constructed, and according to content entity relationship data, to-be-processed text is carried out structured information extraction, and the target structured information contained in to-be-processed file is obtained.The application can improve the accuracy and integrity of structured information extraction.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and in particular to a file structured information extraction method, device, equipment, medium and product. BACKGROUND

[0002] With the development of knowledge base and retrieval enhancement generation technology, efficient parsing and structured processing of multi-type files become core problems.

[0003] In the prior art, for the structured information extraction of multi-type files, the structured information is usually extracted directly according to the text contained in the file, without considering the relevance between the texts, which leads to insufficient accuracy and completeness of the extracted structured information. SUMMARY

[0004] The present application provides a file structured information extraction method, device, equipment, medium and product to solve the problem of insufficient accuracy and completeness of existing file structured information extraction.

[0005] According to an aspect of the present application, a file structured information extraction method is provided, which comprises:

[0006] Obtaining a to-be-processed file whose structured information needs to be extracted, and determining the file content type of the to-be-processed file; wherein the file content type is a text content file or an image content file;

[0007] In the case where the file content type is determined to be the image content file, performing text recognition on the to-be-processed file, and determining the to-be-processed text contained in the to-be-processed file and the text area coordinates corresponding to the to-be-processed text in the to-be-processed file according to the text recognition result;

[0008] Performing structured content entity recognition on the to-be-processed text, determining at least one structured content entity contained in the to-be-processed text, and the content entity coordinates corresponding to each structured content entity in the text area coordinates;

[0009] According to each content entity coordinate, constructing content entity relationship data between each structured content entity, and performing structured information extraction on the to-be-processed text according to the content entity relationship data, to obtain target structured information contained in the to-be-processed file.

[0010] According to another aspect of the present application, a file structured information extraction device is provided, which comprises:

[0011] The file content type determining module is configured to acquire a to-be-processed file in which structured information is to be extracted, and determine a file content type of the to-be-processed file; wherein the file content type is a text content file or an image content file.

[0012] The text recognition module is configured to, in a case where it is determined that the file content type is the image content file, perform text recognition on the to-be-processed file, and determine to-be-processed text contained in the to-be-processed file and text region coordinates corresponding to the to-be-processed text in the to-be-processed file according to a text recognition result.

[0013] The structured content entity recognition module is configured to perform structured content entity recognition on the to-be-processed text, determine at least one structured content entity contained in the to-be-processed text, and determine content entity coordinates corresponding to each of the structured content entities in the text region coordinates.

[0014] The structured information extraction module is configured to construct content entity relationship data between each of the structured content entities according to the content entity coordinates, and perform structured information extraction on the to-be-processed text according to the content entity relationship data, to obtain target structured information contained in the to-be-processed file.

[0015] According to another aspect of the present application, an electronic device is provided, which comprises:

[0016] at least one processor; and

[0017] a memory connected to the at least one processor in communication; wherein,

[0018] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the file structured information extraction method according to any one of the present application.

[0019] According to another aspect of the present application, a computer readable storage medium is provided, which stores computer instructions for enabling a processor to execute the file structured information extraction method according to any one of the present application.

[0020] According to another aspect of the present application, a computer program product is provided, which comprises a computer program for enabling a processor to execute the file structured information extraction method according to any one of the present application.

[0021] The application obtains a to-be-processed file needing to extract structured information, and determines a file content type of the to-be-processed file; wherein the file content type is a text content file or an image content file; in the case that the file content type is determined as the image content file, text recognition is performed on the to-be-processed file, and a to-be-processed text contained in the to-be-processed file and text region coordinates of the to-be-processed text in the to-be-processed file are determined according to a text recognition result; structured content entity recognition is performed on the to-be-processed text, at least one structured content entity contained in the to-be-processed text and content entity coordinates of each structured content entity in the text region coordinates are determined; content entity relationship data between each structured content entity is constructed according to the content entity coordinates, and structured information extraction is performed on the to-be-processed text according to the content entity relationship data, to obtain target structured information contained in the to-be-processed file, and the beneficial effects are as follows:

[0022] By determining the structured content entity and the content entity coordinates contained in the to-be-processed text, constructing the content entity relationship data between the structured content entities according to the content entity coordinates, and performing the structured information extraction by using the content entity relationship data, the extraction process of the structured information refers to the correlation between each structured content entity, so that the accuracy and integrity of the structured information extraction are improved.

[0023] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the application, nor is it used to limit the scope of the application. Other features of the application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0025] Figure 1 A flowchart of a file structured information extraction method provided for the first embodiment of the application;

[0026] Figure 2 A flowchart of a file structured information extraction method provided for the second embodiment of the application;

[0027] Figure 3 A structural schematic diagram of a file structured information extraction device provided for the third embodiment of the application;

[0028] Figure 4 It is a structural schematic diagram of an electronic device for implementing the file structured information extraction method of the embodiment of the application. DETAILED DESCRIPTION

[0029] In order to make the person skilled in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by the person skilled in the art without creative labor should belong to the protection scope of the present application.

[0030] It should be noted that the terms "first", "second", "third", "fourth" and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0031] Embodiment one

[0032] Figure 1 A flowchart of a file structured information extraction method provided by the first embodiment of the present application, the present embodiment can be applicable to the case of using content entity relationship data in a file to extract structured information from the file. The method can be executed by a file structured information extraction device, which can be realized in the form of hardware and / or software, such as a computer. Figure 1 As shown in the figure, the method comprises:

[0033] S101, obtaining a to-be-processed file from which structured information is to be extracted, and determining the file content type of the to-be-processed file.

[0034] The structured information refers to information with a predefined format, clear data relationship and can be directly parsed by a computer program. Its core feature is that data elements are organized in a standardized manner, facilitating machine automatic extraction and analysis. For example, the structured information can be title information, sub-title information, body information, table information, list item information and the like included in a file.

[0035] The to-be-processed file refers to an original input file from which structured information is to be extracted. The file content type refers to the technical classification of the underlying data structure and information bearing form of the to-be-processed file.

[0036] In the present embodiment, the file content type is a text content file or an image content file. The text content file refers to a file type that is stored in the form of character encoding as the core and whose semantic information can be directly read by a text parsing tool. The essence of the text content file is to represent human-readable characters through binary codes. For example, the file content type can include a Word file, a TXT file, and the like.

[0037] The image content file refers to a file type that takes visual pixels or vector graphics as the core data carrier and needs to be parsed by computer vision technology to obtain the semantic information. The essence of the image content file is to record color and coordinate information through binary encoding. For example, the image content file can include a scanned file, a screenshot file, a picture-type PDF file, and the like.

[0038] In one embodiment, when a user or a business system uploads a to-be-processed file that needs to extract structured information to a cloud storage service, a computer automatically triggers an event listening mechanism to determine that the cloud storage service currently stores the latest to-be-processed file, and then the computer obtains the to-be-processed file that needs to extract structured information from the cloud storage service according to the communication connection between the computer and the cloud storage service.

[0039] In another embodiment, a listening service for file storage and file modification is deployed in the computer to monitor a specified directory of a storage device of the computer in real time. When a new to-be-processed file that needs to extract structured information is stored in the specified directory or an existing to-be-processed file that needs to extract structured information is modified, the computer automatically triggers a file acquisition operation to obtain the to-be-processed file that needs to extract structured information.

[0040] After obtaining the to-be-processed file that needs to extract structured information, in one embodiment, the computer reads the binary header characteristic code of the to-be-processed file, and performs comparison in a header characteristic code comparison library through the identified header characteristic code to determine the file content type of the to-be-processed file. For example, the header characteristic code of the text content file usually starts with a printable character code; the header characteristic code of the image content file has a clear identifier, such as the header characteristic code of a JPEG file has a clear identifier “FF D8 FF”, the header characteristic code of a PNG file has a clear identifier “89 50 4E 47”, and the like.

[0041] After obtaining the to-be-processed file that needs to extract structured information, in another embodiment, the file content type of the to-be-processed file is determined according to the statistical characteristics and structural rules of the file content of the to-be-processed file.

[0042] A. It is determined whether the character distribution of the to-be-processed file conforms to the natural language rule or whether there is a continuous semantic paragraph separator. If yes, it is determined that the file content type of the to-be-processed file is a text content file.

[0043] B. verifying whether the binary data of the to-be-processed file contains a regular pixel matrix or a compression algorithm feature, and if so, determining that the file content type of the to-be-processed file is an image content file.

[0044] S102. In a case where it is determined that the file content type is an image content file, performing text recognition on the to-be-processed file, and determining, according to a text recognition result, to-be-processed text contained in the to-be-processed file and a text region coordinate corresponding to the to-be-processed text in the to-be-processed file.

[0045] In an embodiment, in a case where it is determined that the file content type is an image content file, the to-be-processed file is preprocessed, including but not limited to at least one of image denoising, image rectification, and image resolution reconstruction. The image denoising includes but is not limited to removing image noise such as spots and shadows by using an adaptive binarization algorithm; the image rectification includes but is not limited to correcting a tilted image by using a Hough transform, such as making the image tilt angle ≤ 15°; and the image resolution reconstruction includes but is not limited to performing super-resolution reconstruction on a low-resolution image, such as an image with a resolution ≤ 200 dpi, such as using an ESRGAN model to increase the resolution to above 300 dpi.

[0046] In an embodiment, in a case where it is determined that the file content type is an image content file, the to-be-processed file is preprocessed, including but not limited to at least one of image denoising, image rectification, and image resolution reconstruction. The image denoising includes but is not limited to removing image noise such as spots and shadows by using an adaptive binarization algorithm; the image rectification includes but is not limited to correcting a tilted image by using a Hough transform, such as making the image tilt angle ≤ 15°; and the image resolution reconstruction includes but is not limited to performing super-resolution reconstruction on a low-resolution image, such as an image with a resolution ≤ 200 dpi, such as using an ESRGAN model to increase the resolution to above 300 dpi.

[0047] Further, for the preprocessed to-be-processed file, a multi-modal OCR (Optical Character Recognition) engine is used to perform text recognition, wherein the multi-modal OCR engine fuses CTPN (Connectionist Text Proposal Network) text detection and CRNN (Convolutional Recurrent Neural Network) character recognition, and outputs to-be-processed text contained in the to-be-processed file and a text region coordinate corresponding to the to-be-processed text in the to-be-processed file. The output format includes but is not limited to “text content + text region coordinate”, such as “text content XXX, x: 100-500, y: 300-600”.

[0048] S103, performing structured content entity recognition on the to-be-processed text to determine at least one structured content entity contained in the to-be-processed text, and content entity coordinates corresponding to each structured content entity in text region coordinates.

[0049] The structured content entity recognition refers to a process of automatically recognizing predefined structured content entities from unstructured to-be-processed text and determining the coordinate positions of the structured content entities in the text.

[0050] The structured content entity refers to a logical unit with clear semantic boundaries and functional attributes in the to-be-processed text, including but not limited to title, sub-title, text paragraph, table, list item, and the like. The content entity coordinates refer to spatial positioning information of different structured content entities in the to-be-processed text, that is, coordinates corresponding to the structured content entities in the text region coordinates. The structured content entities are associated with physical positions through digital mapping of the coordinates.

[0051] In an embodiment, the structured content entity recognition is performed on the to-be-processed text by a named entity recognition model to determine at least one structured content entity contained in the to-be-processed text, and content entity coordinates corresponding to each structured content entity in text region coordinates.

[0052] For example, a “title entity (‘3.1 device parameters’, coordinates x: 50-200, y: 100-150)”, a “table entity (coordinates x: 100-500, y: 300-600)”, a “text entity (‘device parameter description’, coordinates x: 50-300, y: 200-250)”, and the like are recognized.

[0053] S104, constructing content entity relationship data between the structured content entities according to the content entity coordinates, and performing structured information extraction on the to-be-processed text according to the content entity relationship data to obtain target structured information contained in the to-be-processed file.

[0054] The content entity relationship data refers to structured data describing logical associations and spatial positions between different structured content entities in the document structuring process. The core function of the content entity relationship data is to convert discrete structured content entities into a knowledge network with context associations. For example, the content entity relationship data can be represented in the form of a triple: (structured content entity A, relationship, structured content entity B), such as (title entity “3.1 device parameters”, contains, text entity “device parameter description”); for example, (title 1, contains, sub-title 1.1); for example, (table 2, subordinate, text paragraph 3).

[0055] Structured information extraction refers to the technical process of identifying structured information from unstructured files to be processed, analyzing the logical relationship thereof, and outputting machine-readable structured data.

[0056] In an embodiment, coordinate distances between each structured content entity are determined according to respective content entity coordinates, each structured content entity is associated according to the coordinate distances, structured content entities with similar coordinate distances are determined from each structured content entity, entity relationship construction is performed on each structured content entity with similar coordinate distances, and content entity relationship data is generated. Further, a large language model is used to extract structured information from the text to be processed according to the content entity relationship data, and the target structured information contained in the file to be processed is output.

[0057] For example, the output example of the target structured information includes "| parameter name | value | note |\n| voltage | 220V | N / A |", and it can be understood that the above example is only an example of the form and content of the target structured information, and does not limit the form and content of the target structured information.

[0058] Optionally, after generating the content entity relationship data, the method further includes generating a knowledge graph according to the entity relationship data by using an attribute graph model, wherein the structured content entity is a node in the knowledge graph, and the relationship between each structured content entity is an edge relationship. The node attributes include at least one of entity type, content entity text, content entity coordinate, and confidence; and the edge relationship attributes include edge relationship types such as containing relationship or subordinate relationship, coordinate distance, etc.

[0059] The embodiment of the application obtains a file to be processed for extracting structured information, and determines the file content type of the file to be processed. In the case where the file content type is determined to be an image content file, the file to be processed is subjected to text recognition, and the text to be processed contained in the file to be processed and the text area coordinate corresponding to the text to be processed in the file to be processed are determined according to the text recognition result. The text to be processed is subjected to structured content entity recognition, at least one structured content entity contained in the text to be processed and respective content entity coordinates corresponding to each structured content entity in the text area coordinate are determined, content entity relationship data between each structured content entity is constructed according to the respective content entity coordinates, and structured information extraction is performed on the text to be processed according to the content entity relationship data, so as to obtain target structured information contained in the file to be processed, which has the following beneficial effects:

[0060] In the first aspect, by determining the structured content entity and the content entity coordinates contained in the to-be-processed text, the content entity relationship data between the structured content entities is constructed according to the content entity coordinates, and the structured information extraction is performed by using the content entity relationship data, so that the extraction process of the structured information refers to the correlation between the structured content entities, thereby improving the accuracy and integrity of the structured information extraction.

[0061] In the second aspect, the effect of structured information extraction on image content files is realized, the application range of structured information extraction is expanded, and the business needs of users for structured information extraction on multi-modal files are met.

[0062] Optionally, before the structured information extraction on the to-be-processed text according to the content entity relationship data is performed to obtain the target structured information contained in the to-be-processed file, the method further includes:

[0063] In the case that the file content type is a text content file, a preset structured content hierarchical label in the to-be-processed file is obtained; and the content entity relationship data between the structured content entities is constructed according to the structured content hierarchical label.

[0064] The structured content hierarchical label refers to a semantic mark used to identify different hierarchical logical structures in the to-be-processed file, and the core function thereof is to construct a structured correlation network of content entities through the hierarchical relationship of the label. For example, the structured content hierarchical label <w:t>For marking up body text; also, for example, structured content hierarchy tags <w:tbl>For marking tables and the like.

[0065] In one embodiment, in case that the file content type is a text content file, the file to be processed is parsed to obtain a preset structured content hierarchy label in the file to be processed. Further, according to the preset structured content hierarchy label, content entity relationship data between each structured content entity is constructed.

[0066] For example, assume that the structured content hierarchy label <w:t>Structured content hierarchy tag <w:tbl>and structured content hierarchy tags <w:t>For marking body 1, structured content hierarchy tags <w:tbl>For marking table 1, it is determined that the text 1 contains the table 1.

[0067] By acquiring the preset structured content hierarchical label in the to-be-processed file when the file content type is a text content file, and constructing the content entity relationship data between the structured content entities according to the structured content hierarchical label, the beneficial effects are as follows:

[0068] In the first aspect, the structured content entities in the to-be-processed text can be automatically identified through the preset structured content hierarchical label, the problem of ambiguous structure in traditional text processing is solved, and the effect of accurately identifying the content structure is achieved.

[0069] In the second aspect, the parent-child, parallel and other relationship data of the structured content entities are automatically generated based on the structured content hierarchical label, the entity relationship data in the form of a knowledge graph is formed, and the effect of automatically constructing the entity relationship data is achieved.

[0070] Optionally, after obtaining the target structured information contained in the to-be-processed file, the method further includes:

[0071] The target structured information is associated and stored with the content entity relationship data in a target knowledge base.

[0072] The associated storage refers to the binding storage of the target structured information and the content entity relationship data through a specific technical means, so that a traceable semantic network is formed between the data. The target knowledge base refers to a special database system that is structured and designed for the associated storage of specific structured information and content entity relationship data.

[0073] In an embodiment, after obtaining the target structured information contained in the to-be-processed file, the target structured information is associated and stored with the content entity relationship data in a target knowledge base by using a data association storage technology.

[0074] By associating and storing the target structured information with the content entity relationship data in the target knowledge base, the beneficial effects are as follows: the structured information and the content entity relationship data can be called simultaneously during subsequent retrieval.

[0075] Embodiment two

[0076] Figure 2 A flowchart of a file structured information extraction method provided for the second embodiment of the present application, the present embodiment further optimizes and extends the above-mentioned embodiments, and can be combined with the above-mentioned optional embodiments. As shown in the figure, the method includes: Figure 2

[0077] S201, acquiring a to-be-processed file that needs to extract structured information, and determining a file content type of the to-be-processed file.

[0078] ​Optionally, after determining the file content type of the to-be-processed file, the method further comprises:

[0079] In a case where the file content type is determined as the text content file, text extraction is performed on the to-be-processed file, and the to-be-processed text contained in the to-be-processed file and the text region coordinates corresponding to the to-be-processed text in the to-be-processed file are determined.

[0080] The text extraction refers to a process of recognizing and extracting editable character sequences, i.e., the to-be-processed text, from the text content file, while recording the spatial position information of the to-be-processed text in the to-be-processed file.

[0081] In an embodiment, if the file content type of the to-be-processed file is determined as the text content file, a text extraction algorithm is used to directly perform text extraction on the to-be-processed file, and the to-be-processed text contained in the to-be-processed file and the text region coordinates corresponding to the to-be-processed text in the to-be-processed file are determined.

[0082] By performing text extraction on the to-be-processed file in a case where the file content type is determined as the text content file, the to-be-processed text contained in the to-be-processed file and the text region coordinates corresponding to the to-be-processed text in the to-be-processed file are determined, which has the beneficial effect of accurately positioning the physical position of the to-be-processed text in the to-be-processed file through the text region coordinates, and realizing dual analysis of content and structure.

[0083] S202, in a case where the file content type is determined as the image content file, text recognition is performed on the to-be-processed file, and the to-be-processed text contained in the to-be-processed file and the text region coordinates corresponding to the to-be-processed text in the to-be-processed file are determined according to the text recognition result.

[0084] S203, structured content entity recognition is performed on the to-be-processed text, at least one structured content entity contained in the to-be-processed text and content entity coordinates corresponding to each structured content entity in the text region coordinates are determined.

[0085] S204, coordinate distances between the structured content entities are determined according to the content entity coordinates, distance weight information between the structured content entities is determined according to the coordinate distances, and similar structured content entities are determined from the structured content entities according to the distance weight information.

[0086] The coordinate distance refers to a difference in spatial positions of the structured content entities in a page coordinate system calculated by a mathematical method, and the core role is to quantify the physical layout correlation between the structured content entities.

[0087] The distance weight information refers to a numerical indicator quantifying the correlation strength or influence degree between each structured content entity based on the coordinate distance between the structured content entities. The core logic is that the closer the coordinate distance between each structured content entity, the higher the weight between them, and the farther the distance, the weight decays. The similar structured content entity refers to the structured content entity with adjacent physical position in the layout space of the text to be processed.

[0088] In an embodiment, the coordinate distance is calculated according to the coordinates of each content entity, and the coordinate distance between each structured content entity is determined according to the calculation result. Further, the coordinate distance is matched in the mapping relationship between the preset candidate coordinate distance and the candidate distance weight information, and the distance weight information between each structured content entity is determined according to the matching result. The distance weight information is optionally set in the interval of 0-1.

[0089] According to the distance weight information and the distance weight threshold, each structured content entity is screened, and each structured content entity with distance weight information greater than the distance weight threshold is selected as the similar structured content entity.

[0090] S205, entity relationship construction is performed on each similar structured content entity, and content entity relationship data is generated.

[0091] By determining the coordinate distance between each structured content entity according to the coordinates of each content entity, determining the distance weight information between each structured content entity according to the coordinate distance, and determining the similar structured content entity from each structured content entity according to the distance weight information, and performing entity relationship construction on each similar structured content entity to generate content entity relationship data, the beneficial effects are:

[0092] First, based on the coordinates of the content entities, the coordinate distance between each structured content entity is automatically calculated, replacing manual calculation and improving efficiency.

[0093] Second, according to the distance decay principle, the distance weight information is automatically generated to quantify the correlation strength between each structured content entity.

[0094] Third, the scattered content entity coordinates are converted into entity relationship data with weights, realizing the conversion from unstructured data to structured data.

[0095] S206, the preset structure information extraction instruction is obtained from the instruction database, and the structure information extraction instruction, the text to be processed and the content entity relationship data are input into the large language model.

[0096] The instruction database refers to a special data warehouse for storing and managing preset structured operation instructions, which is used to guide the large language model to perform specific information extraction tasks. The core function is to realize efficient and accurate data structured extraction through standardized instruction templates.

[0097] The structured information extraction instruction refers to a preset command template in the database, which guides the large language model to accurately extract structured data from unstructured text and output standardized formats. The large language model refers to a super-large artificial intelligence model trained by massive text, with deep semantic understanding and generation capabilities, and the parameter scale usually reaches tens of billions to hundreds of billions. In this embodiment, the core function of the large language model is to realize structured information extraction through instruction analysis, text understanding, and structured output.

[0098] In one embodiment, the preset structured information extraction instruction is called from the prompt type instruction database, and the structured information extraction instruction, the to-be-processed text, and the content entity relationship data are collectively input into the large language model as input text.

[0099] For example, the form of the structured information extraction instruction includes but is not limited to: core task: "based on the entity relationship in the content entity relationship data, classify the to-be-processed text into title, body, table, and list, and extract the corresponding content"; constraint condition: "the title must contain ≥3 characters and ≤20 characters; the table must extract the table header and data row completely, and the missing cells are marked as 'N / A'".

[0100] It can be understood that the embodiment only exemplifies the form of the structured information extraction instruction, and does not specifically limit the form of the structured information extraction instruction.

[0101] S207, structured information extraction is performed from the to-be-processed text by the large language model according to the structured information extraction instruction and using the content entity relationship data, to obtain target structured information.

[0102] By obtaining the preset structured information extraction instruction from the instruction database, and inputting the structured information extraction instruction, the to-be-processed text, and the content entity relationship data into the large language model, structured information extraction is performed from the to-be-processed text by the large language model according to the structured information extraction instruction and using the content entity relationship data, to obtain target structured information, which has the beneficial effects of:

[0103] First, the preset structured information extraction instruction provides a standardized analysis framework for the large language model, avoiding the generalization problem of traditional regular expressions or rule engines, and improving the accuracy of structured information extraction.

[0104] In a second aspect, the large language model uses content entity relationship data as a semantic anchor point, can process nested entities and ambiguous expressions, dynamically adapt to complex semantics, and further ensure the accuracy of structured information extraction.

[0105] In a third aspect, the structured information extraction instruction forces the large language model to output by field, avoids the random divergence problem of generative models, and ensures the stability of the target structured information output by the model.

[0106] Optionally, the large language model uses content entity relationship data to extract structured information from the to-be-processed text according to the structured information extraction instruction, and obtains target structured information, including:

[0107] S2071, the large language model uses content entity relationship data to extract structured content text corresponding to each structured content entity from the to-be-processed text according to the structured information extraction instruction.

[0108] Among them, the structured content text refers to the information unit extracted from the unstructured original text by the large language model and conforming to the pre-defined logical framework.

[0109] S2072, the semantic similarity of each structured content text with content entity relationship is calculated respectively, and the semantic similarity between each structured content text with content entity relationship is determined.

[0110] Among them, the semantic similarity calculation refers to measuring the closeness of two pieces of structured content text in the meaning level through a quantitative method, especially focusing on the semantic association of entity relationship. The core goal is to go beyond the surface lexical difference and capture the deep logical consistency.

[0111] Semantic similarity refers to measuring the closeness of two pieces of structured content text in the meaning level rather than in the literal form. It focuses on whether the core concepts, logical relationships or emotional intentions expressed by the text are consistent, rather than simply overlapping words.

[0112] In an embodiment, the large language model calculates the semantic similarity of each structured content text with content entity relationship respectively, and calculates the semantic similarity between each structured content text with content entity relationship.

[0113] S2073, according to the semantic similarity, determine the semantic similar content text from each structured content text with content entity relationship, and perform text association on each semantic similar content text, to obtain target structured information.

[0114] The semantic similar content text refers to a structured content text with a highly similar concept, entity relationship or core intent expressed in the text. The text association refers to a process of constructing a structured information network by logical connection, feature integration or relationship mapping after identifying the semantic similar content text based on semantic similarity.

[0115] In an embodiment, the structured content texts with the highest semantic similarity and content entity relationship are determined as the semantic similar content texts by screening and sorting based on semantic similarity through the large language model, and the target structured information is obtained by text association of each semantic similar content text.

[0116] According to the structure information extraction instruction of the large language model, the structured content text corresponding to each structured content entity is extracted from the to-be-processed text using the content entity relationship data. The semantic similarity between each structured content text with content entity relationship is determined by calculating the semantic similarity of each structured content text with content entity relationship. The semantic similar content text is determined from each structured content text with content entity relationship according to the semantic similarity, and the target structured information is obtained by text association of each semantic similar content text. The beneficial effects are as follows:

[0117] In the first aspect, if only the content entity relationship data is used for structured information extraction, it may lead to structured information extraction errors, for example, assuming that the coordinate distance between title 1 and the text 1 in the previous text, and the text 2 in the following text is relatively close. If only the content entity relationship data is used for structured information extraction, it will lead to the error of extracting the structured information of title 1 and text 1.

[0118] However, the embodiment combines semantic similarity for structured information extraction, which can effectively avoid the above problems and improve the accuracy and reliability of structured information extraction.

[0119] In the second aspect, the text association of the structured content text with semantic similarity eliminates information redundancy.

[0120] Optionally, after obtaining the target structured information, the method further includes:

[0121] A1, the information identifier and the structure information extraction instruction of the target structured information are associated and stored, and a storage association relationship between the information identifier and the structure information extraction instruction is generated.

[0122] The information identifier refers to a unique identification symbol or code allocated for the target structured information, which is used for accurate positioning and distinguishing different structured information. The storage association relationship refers to the mapping relationship between the information identifier and the structure information extraction instruction.

[0123] In an embodiment, after obtaining the target structured information, the information identifier of the target structured information and the structure information extraction instruction are stored in association in the instruction database, and the record field includes but is not limited to instruction text, call time, entity relationship data summary, output format type, etc.

[0124] B1, determine whether the target structured information is abnormal, and if it is determined that the target structured information is abnormal, obtain the structure information extraction instruction according to the association between the information identifier and the storage.

[0125] Wherein, the target structured information is abnormal refers to the extracted target structured information violates the pre-defined rule template, business logic constraint or data structure specification, and its determination needs to be based on the pre-set abnormality detection rule.

[0126] In an embodiment, based on the pre-set abnormality detection rule, it is determined whether the target structured information is abnormal, and if it is determined that the target structured information is abnormal, the structure information extraction instruction associated with the information identifier is obtained according to the matching of the association between the information identifier and the storage.

[0127] Optionally, determining whether the target structured information is abnormal comprises:

[0128] B11, determine the character similarity between the title structured text included in the target structured information and the title text included in the content entity relationship data.

[0129] Wherein, the title structured text refers to the standardized title text obtained by performing structured information extraction on the to-be-processed text. The title text refers to the natural language title in the to-be-processed text recorded in the content entity relationship data. The character similarity refers to a measurement method for quantifying the similarity by comparing the surface character composition and sequence similarity of the text string.

[0130] In an embodiment, the character similarity between the title structured text included in the target structured information and the title text included in the content entity relationship data is determined by including an edit distance algorithm.

[0131] B12, calculate the table text extraction rate according to the table structured text included in the target structured information and the table text included in the to-be-processed text.

[0132] The table structured text refers to the standardized table text obtained by performing structured information extraction on the to-be-processed text. The table text refers to the original table content in the to-be-processed text without processing. The table text extraction rate refers to the number of table structured texts successfully recognized and extracted from the to-be-processed text in the structured information extraction process, and accounts for the proportion of the total amount of table texts included in the to-be-processed text. It is a core index for measuring the accuracy and completeness of table structured information extraction.

[0133] B13, calculating a list item text extraction rate according to the list item structured text included in the target structured information and the list item text included in the to-be-processed text.

[0134] The list item structured text refers to the standardized list item text obtained by performing structured information extraction on the to-be-processed text. The list item text refers to the original list item content in the to-be-processed text without processing. The list item text extraction rate refers to the number of list item structured texts successfully recognized and extracted from the to-be-processed text in the structured information extraction process, and accounts for the proportion of the total amount of list item texts included in the to-be-processed text. It is a core index for measuring the accuracy and completeness of list item structured information extraction.

[0135] B14, determining a text similarity between the target structured information and the to-be-processed text.

[0136] The text similarity refers to the closeness between the target structured information and the to-be-processed text in semantic content, structural features or expression intention by a quantitative method. The core is to evaluate the relevance or repeatability between texts.

[0137] In an embodiment, the BLEU value between the target structured information and the to-be-processed text is calculated by BLEU (Bilingual Evaluation Understudy) to serve as the text similarity.

[0138] B15, determining that the target structured information is abnormal in a case that the character similarity is less than a first threshold value, the table text extraction rate is less than a second threshold value, the list item text extraction rate is less than a third threshold value, or the text similarity is less than a fourth threshold value.

[0139] The first threshold value, the second threshold value, the third threshold value and the fourth threshold value are set according to actual business requirements. For example, the first threshold value can be set to 0.9, the second threshold value can be set to 0.95, the third threshold value can be set to 0.95, and the fourth threshold value can be set to 0.85.

[0140] The character similarity between the title structured text included in the target structured information and the title text included in the content entity relationship data is determined, the table text extraction rate is calculated according to the table structured text included in the target structured information and the table text included in the to-be-processed text, the list item text extraction rate is calculated according to the list item structured text included in the target structured information and the list item text included in the to-be-processed text, the text similarity between the target structured information and the to-be-processed text is determined, and in a case where the character similarity is less than a first threshold value, the table text extraction rate is less than a second threshold value, the list item text extraction rate is less than a third threshold value, or the text similarity is less than a fourth threshold value, it is determined that the target structured information is abnormal, and the beneficial effects are that:

[0141] The character similarity between the title structured text and the title text is determined, the problem of title misplacement is solved, the table text extraction rate is calculated according to the table structured text and the table text, the list item text extraction rate is calculated according to the list item structured text and the list item text, the effect of quantifying the content completeness of table / list item structured information extraction is achieved, the text similarity between the target structured information and the to-be-processed text is determined, so that the problem of semantic tampering can be captured in time, and the target structured information is detected abnormally from three dimensions of the title layer, the data layer and the semantic layer, so that the comprehensiveness and reliability of the target structured information abnormality detection are ensured.

[0142] C1, fine-tuning the structure information extraction instruction to obtain a fine-tuned instruction corresponding to the structure information extraction instruction, and updating the fine-tuned instruction to the instruction database.

[0143] In an embodiment, an entity relationship abnormality existing in the content entity relationship data is determined, and an instruction defect of the structure information extraction instruction is located according to the entity relationship abnormality, and the structure information extraction instruction is fine-tuned according to a defect type of the instruction defect to obtain a fine-tuned instruction corresponding to the structure information extraction instruction.

[0144] For example, the structure information extraction instruction is supplemented with rules, such as "when the distance weight between the title entity and the body entity is 0.6, it is determined as a secondary title"; for another example, the structure information extraction instruction is constrained and strengthened, such as "table extraction needs to check the consistency of the number of table headers and data rows, and if the consistency is not met, mark 'row number mismatching'".

[0145] Further, the fine-tuned instruction is updated to the instruction database to replace the structure information extraction instruction before fine-tuning.

[0146] By storing the information identifier of the target structured information and the structure information extraction instruction in association, and generating a storage association relationship between the information identifier and the structure information extraction instruction, it is determined whether the target structured information is abnormal, and in the case where it is determined that the target structured information is abnormal, the structure information extraction instruction is obtained according to the information identifier and the storage association relationship; the structure information extraction instruction is fine-tuned to obtain a fine-tuning instruction corresponding to the structure information extraction instruction, and the fine-tuning instruction is updated to the instruction database, and the beneficial effects are that the structure information extraction instruction is bound with the information identifier, the accurate mapping of "index abnormality -> relationship abnormality -> instruction defect" is realized, and the problem that the traditional evaluation cannot locate the root cause is solved.

[0147] Embodiment three

[0148] Figure 3 A structural schematic diagram of a file structured information extraction device provided for the embodiment three of the application can be applicable to the case of extracting structured information from a file by using content entity relationship data in the file, as shown in the figure, the device comprises: Figure 3

[0149] A file content type determination module 31 is configured to acquire a to-be-processed file whose structured information is to be extracted, and determine a file content type of the to-be-processed file; wherein the file content type is a text content file or an image content file.

[0150] A text recognition module 32 is configured to, in the case where the file content type is determined to be the image content file, perform text recognition on the to-be-processed file, and determine a to-be-processed text contained in the to-be-processed file and text region coordinates corresponding to the to-be-processed text in the to-be-processed file according to a text recognition result.

[0151] A structured content entity recognition module 33 is configured to perform structured content entity recognition on the to-be-processed text, determine at least one structured content entity contained in the to-be-processed text, and determine content entity coordinates corresponding to each of the structured content entities in the text region coordinates.

[0152] A structured information extraction module 34 is configured to construct content entity relationship data between each of the structured content entities according to each of the content entity coordinates, and perform structured information extraction on the to-be-processed text according to the content entity relationship data, to obtain target structured information contained in the to-be-processed file.

[0153] Optionally, the structured information extraction module 34 is specifically configured to:

[0154] determine a coordinate distance between each of the structured content entities according to each of the content entity coordinates; ​

[0155] determine distance weight information between each of the structured content entities according to the coordinate distance, and determine similar structured content entities from each of the structured content entities according to the distance weight information;

[0156] perform entity relationship construction on each of the similar structured content entities, and generate the content entity relationship data.

[0157] Optionally, the apparatus further comprises a structured content hierarchical label parsing module, specifically configured to:

[0158] in a case where the file content type is the text content file, acquire a preset structured content hierarchical label in the to-be-processed file;

[0159] construct content entity relationship data between each of the structured content entities according to the structured content hierarchical label.

[0160] Optionally, the structured information extraction module 34 is specifically further configured to:

[0161] acquire a preset structured information extraction instruction from an instruction database, and input the structured information extraction instruction, the to-be-processed text and the content entity relationship data into a large language model;

[0162] perform structured information extraction from the to-be-processed text by the large language model according to the structured information extraction instruction and by using the content entity relationship data, to obtain the target structured information.

[0163] Optionally, the structured information extraction module 34 is specifically further configured to:

[0164] extract, by the large language model according to the structured information extraction instruction and by using the content entity relationship data, structured content texts respectively corresponding to each of the structured content entities from the to-be-processed text;

[0165] perform semantic similarity calculation on each of the structured content texts having content entity relationship respectively, to determine semantic similarity between each of the structured content texts having content entity relationship;

[0166] determine semantic similar content texts from each of the structured content texts having content entity relationship according to the semantic similarity, and perform text association on each of the semantic similar content texts, to obtain the target structured information.

[0167] Optionally, the apparatus further comprises a text extraction module, specifically configured to:

[0168] In a case where the file content type is determined as the text content file, text extraction is performed on the to-be-processed file, to-be-processed text contained in the to-be-processed file is determined, and a text region coordinate corresponding to the to-be-processed text in the to-be-processed file is determined.

[0169] Optionally, the apparatus further comprises an instruction fine-tuning module, specifically configured to:

[0170] store the information identifier of the target structured information and the structured information extraction instruction in association, and generate a storage association relationship between the information identifier and the structured information extraction instruction;

[0171] determine whether the target structured information is abnormal, and in a case where it is determined that the target structured information is abnormal, acquire the structured information extraction instruction according to the information identifier and the storage association relationship;

[0172] fine-tune the structured information extraction instruction to obtain a fine-tuned instruction corresponding to the structured information extraction instruction, and update the fine-tuned instruction to the instruction database.

[0173] Optionally, the instruction fine-tuning module is specifically further configured to:

[0174] determine a character similarity between title structured text included in the target structured information and title text included in the content entity relationship data;

[0175] calculate a table text extraction rate according to table structured text included in the target structured information and table text included in the to-be-processed text;

[0176] calculate a list item text extraction rate according to list item structured text included in the target structured information and list item text included in the to-be-processed text;

[0177] determine a text similarity between the target structured information and the to-be-processed text;

[0178] in a case where the character similarity is less than a first threshold value, the table text extraction rate is less than a second threshold value, the list item text extraction rate is less than a third threshold value, or the text similarity is less than a fourth threshold value, determine that the target structured information is abnormal.

[0179] Optionally, the apparatus further comprises an association storage module, specifically configured to:

[0180] store the target structured information and the content entity relationship data in association in a target knowledge base.

[0181] The file structured information extraction device provided by the embodiment of the present application can execute the file structured information extraction method provided by any embodiment of the present application, and has the function modules and beneficial effects corresponding to the execution method.

[0182] According to embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0183] Embodiment four

[0184] Figure 4 A structural diagram of an electronic device 40 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present application described and / or claimed in this document.

[0185] As shown in Figure 4 The electronic device 40 includes at least one processor 41, and a memory, such as a read-only memory (ROM) 42, a random access memory (RAM) 43, etc., connected to the at least one processor 41 in communication, wherein the memory stores a computer program executable by the at least one processor. The processor 41 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 42 or the computer program loaded from the storage unit 48 into the random access memory (RAM) 43. In the RAM 43, various programs and data required for the operation of the electronic device 40 can also be stored. The processor 41, the ROM 42, and the RAM 43 are connected to each other through a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.

[0186] A plurality of components in the electronic device 40 are connected to the I / O interface 45, including: an input unit 46, such as a keyboard, a mouse, etc.; an output unit 47, such as various types of displays, speakers, etc.; a storage unit 48, such as a magnetic disk, an optical disk, etc.; and a communication unit 49, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 49 allows the electronic device 40 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunications networks.

[0187] The processor 41 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, and the like. The processor 41 performs various methods and processes described above, such as the extraction method of file structured information.

[0188] In some embodiments, the extraction method of file structured information can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 48. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 40 via the ROM 42 and / or the communication unit 49. When the computer program is loaded onto the RAM 43 and executed by the processor 41, one or more steps of the extraction method of file structured information described above can be performed. Alternatively, in other embodiments, the processor 41 can be configured to perform the extraction method of file structured information by any other suitable means, such as by means of firmware.

[0189] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0190] Computer programs used to implement the methods of the application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed, implements the functions / acts specified in the flowcharts and / or block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, and partially on a machine or a remote machine or a server.

[0191] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0192] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0193] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0194] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business expansion in traditional physical host and virtual private service.

[0195] It should be understood that the various forms of flow shown above can be used to reorder, add or delete steps. For example, each step described in the present application can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions of the present application can be achieved, which is not limited herein.

[0196] The above detailed description does not constitute a limitation on the scope of protection of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application shall be included in the scope of protection of the present application.< / w:tbl> < / w:t> < / w:tbl> < / w:t> < / w:tbl> < / w:t>

Claims

1. A method of extracting structured information from a file, characterized by, The method includes: Obtain the file to be processed from which structured information needs to be extracted, and determine the file content type of the file to be processed; wherein, the file content type is a text content file or an image content file; If the file content type is determined to be the image content file, text recognition is performed on the file to be processed, and the text to be processed contained in the file to be processed is determined according to the text recognition result, as well as the text region coordinates corresponding to the text to be processed in the file to be processed. The text to be processed is subjected to structured content entity recognition to determine at least one structured content entity contained in the text to be processed, and the content entity coordinates corresponding to each structured content entity in the text region coordinates respectively; Based on the coordinates of each content entity, content entity relationship data between each structured content entity is constructed, and the structured information of the text to be processed is extracted based on the content entity relationship data to obtain the target structured information contained in the file to be processed. The step of extracting structured information from the text to be processed based on the content entity relationship data to obtain the target structured information contained in the file to be processed includes: The system retrieves preset structural information extraction instructions from the instruction database and inputs the structural information extraction instructions, the text to be processed, and the content entity relationship data into the large language model. Based on the structural information, the large language model extracts instructions and uses the content entity relationship data to extract the structured content text corresponding to each structured content entity from the text to be processed. Semantic similarity is calculated for each structured content text with content entity relationships to determine the semantic similarity between the structured content texts with content entity relationships. Based on the semantic similarity, semantically similar content texts are determined from each of the structured content texts with content entity relationships, and text association is performed on each of the semantically similar content texts to obtain the target structured information.

2. The method of claim 1, wherein, The step of constructing content entity relationship data between the structured content entities based on the coordinates of each content entity includes: The coordinate distance between each structured content entity is determined based on the coordinates of each content entity. Based on the coordinate distance, distance weight information between each structured content entity is determined, and similar structured content entities are determined from each structured content entity based on the distance weight information. Entity relationships are constructed for each of the similar structured content entities to generate the content entity relationship data.

3. The method of claim 1, wherein, Before extracting structured information from the text to be processed based on the content entity relationship data to obtain the target structured information contained in the file to be processed, the method further includes: When the file content type is a text content file, obtain the preset structured content hierarchy tags in the file to be processed; Based on the structured content hierarchy tags, content entity relationship data between each structured content entity is constructed.

4. The method according to claim 1, characterized in that, After determining the file content type of the file to be processed, the process further includes: If the file content type is determined to be a text content file, text extraction is performed on the file to be processed to determine the text to be processed contained in the file to be processed, and the coordinates of the text region corresponding to the text to be processed in the file to be processed.

5. The method according to claim 1, characterized in that, After obtaining the target structured information, the process further includes: The information identifier of the target structured information and the structured information extraction instruction are associated and stored together, and a storage association relationship is generated between the information identifier and the structured information extraction instruction; Determine whether the target structured information is abnormal, and if it is determined that the target structured information is abnormal, obtain the structured information extraction instruction based on the information identifier and the storage association relationship; The structural information extraction instruction is fine-tuned to obtain the fine-tuning instruction corresponding to the structural information extraction instruction, and the fine-tuning instruction is updated to the instruction database.

6. The method according to claim 5, characterized in that, Determining whether the target structured information is abnormal includes: Determine the character similarity between the title structured text included in the target structured information and the title text included in the content entity relationship data; The table text extraction rate is calculated based on the table structured text included in the target structured information and the table text included in the text to be processed. The list item text extraction rate is calculated based on the list item structured text included in the target structured information and the list item text included in the text to be processed. Determine the text similarity between the target structured information and the text to be processed; If the character similarity is less than a first threshold, the table text extraction rate is less than a second threshold, the list item text extraction rate is less than a third threshold, or the text similarity is less than a fourth threshold, then the target structured information is determined to be abnormal.

7. The method according to claim 1, characterized in that, After obtaining the target structured information contained in the file to be processed, the process further includes: The target structured information is associated with the content entity relationship data and stored in the target knowledge base.

8. A device for extracting structured information from documents, characterized in that, The device includes: The file content type determination module is used to obtain the file to be processed from which structured information needs to be extracted, and to determine the file content type of the file to be processed; wherein, the file content type is a text content file or an image content file; The text recognition module is used to perform text recognition on the file to be processed when the file content type is determined to be the image content file, and to determine the text to be processed contained in the file to be processed based on the text recognition result, as well as the text region coordinates corresponding to the text to be processed in the file to be processed. The structured content entity recognition module is used to perform structured content entity recognition on the text to be processed, determine at least one structured content entity contained in the text to be processed, and the content entity coordinates corresponding to each structured content entity in the text region coordinates. The structured information extraction module is used to construct content entity relationship data between the structured content entities based on the coordinates of each content entity, and to extract structured information from the text to be processed based on the content entity relationship data to obtain the target structured information contained in the file to be processed. The structured information extraction module is specifically used for: The system retrieves preset structural information extraction instructions from the instruction database and inputs the structural information extraction instructions, the text to be processed, and the content entity relationship data into the large language model. Based on the structural information, the large language model extracts instructions and uses the content entity relationship data to extract the structured content text corresponding to each structured content entity from the text to be processed. Semantic similarity is calculated for each structured content text with content entity relationships to determine the semantic similarity between the structured content texts with content entity relationships. Based on the semantic similarity, semantically similar content texts are determined from each of the structured content texts with content entity relationships, and text association is performed on each of the semantically similar content texts to obtain the target structured information.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the method for extracting structured information from documents according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a processor to perform the method for extracting structured information from a file as described in any one of claims 1-7.

11. A computer program product comprising a computer program that, when executed by a processor, implements a method for extracting structured information from a file according to any one of claims 1-7.

Citation Information

Patent Citations

  • Bill information identification method, apparatus and device, and storage medium

    CN120014649A

  • Text data processing and medical data structuring method

    CN120126646A