Annotation Method and System for Unstructured Data Documents

By combining knowledge graphs and labeling tools, visual annotation and automated processing of unstructured data are achieved, the problem of inefficiency in the existing technology is solved, the efficiency and quality of data labeling are improved, and the full life cycle management is supported.

CN115599908BActive Publication Date: 2025-07-04金现代信息产业股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211371394.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-03
Publication Date
2025-07-04
Estimated Expiration
2042-11-03

AI Technical Summary

Technical Problem

In the prior art, unstructured data labeling is inefficient, relies on manual operations and lacks visual display, and cannot be directly used for knowledge graph applications.

Method used

By combining knowledge graphs and labeling tools, we provide visual labeling rules construction and data review process, support corpus mode and graph entry mode, and realize automated data annotation and online preview.

Benefits of technology

It improves the efficiency and quality of data labeling, simplifies the operation process, reduces redundant data, supports full life cycle management, and improves the direct application of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115599908B_ABST
    Figure CN115599908B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for annotating unstructured data documents; the method includes: constructing annotation rules; uploading the document to be annotated and auditing the document to be annotated; creating an annotation task; performing data annotation, auditing the data annotation result, and judging whether the audit passes. If it is judged that the task mode is the corpus mode or the image input mode, if it is the corpus mode, the annotation result is directly generated as a corpus; if it is the image input mode, alignment operation is performed on the annotation result, and the result after the alignment operation is subjected to image input processing. The present invention realizes the visualization of the annotated data through the combination of a knowledge graph and an annotation tool, and after the data annotation is completed, the annotated data can be previewed online.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of document annotation, and particularly to a method and system for unstructured data document annotation. Background Art

[0002] The statements in this section merely mention the background art related to the present invention and do not necessarily constitute prior art.

[0003] With the gradual maturity of knowledge graph technology, more and more systems have begun to integrate the application of knowledge graphs. Applying knowledge graphs requires extracting entities, relationships, and attributes from a large amount of data to form a knowledge network. One important data source is unstructured data, which has increased the demand for data annotation. Currently, most data annotation relies on the experience of annotators for manual annotation, resulting in low efficiency. The annotated data is not visually displayed and cannot be directly used. Summary of the Invention

[0004] To solve the deficiencies of the prior art, the present invention provides a method and system for unstructured data document annotation. The present invention realizes the visualization of annotated data by combining a knowledge graph and an annotation tool, and after the data annotation is completed, the annotated data can be previewed online.

[0005] In a first aspect, the present invention provides a method for unstructured data document annotation;

[0006] The method for unstructured data document annotation includes:

[0007] (1) Construct annotation rules; upload the document to be annotated, review the document to be annotated; create an annotation task;

[0008] (2) Perform data annotation, review the data annotation result, and proceed to (3);

[0009] (3) Determine whether the review is passed. If yes, proceed to (4); if not, return to (2);

[0010] (4) Determine whether the task mode is the corpus mode or the graph entry mode. If it is the corpus mode, directly generate the corpus from the annotation result; if it is the graph entry mode, perform an alignment operation on the annotation result and perform graph entry processing on the result after the alignment operation.

[0011] In a second aspect, the present invention provides a system for unstructured data document annotation;

[0012] The system for unstructured data document annotation includes:

[0013] A rule construction module, which is configured to: construct annotation rules; upload the document to be annotated, review the document to be annotated; create an annotation task;

[0014] A data annotation module, which is configured to: perform data annotation, review the data annotation results, and enter the review judgment module;

[0015] A review judgment module, which is configured to: judge whether the review is passed. If so, enter the mode judgment module; if not, return to the data annotation module;

[0016] A mode judgment module, which is configured to: judge whether the task mode is a corpus mode or an image entry mode. If it is the corpus mode, directly generate a corpus from the annotation results; if it is the image entry mode, perform an alignment operation on the annotation results and perform image entry processing on the results after the alignment operation.

[0017] In a third aspect, the present invention further provides an electronic device, including:

[0018] A memory for non-temporarily storing computer-readable instructions; and

[0019] A processor for running the computer-readable instructions,

[0020] wherein, when the computer-readable instructions are run by the processor, the method described in the first aspect above is executed.

[0021] In a fourth aspect, the present invention further provides a storage medium for non-temporarily storing computer-readable instructions, wherein when the non-temporary computer-readable instructions are executed by a computer, the instructions for executing the method described in the first aspect are executed.

[0022] In a fifth aspect, the present invention further provides a computer program product, including a computer program, and the computer program is used to implement the method described in the first aspect above when running on one or more processors.

[0023] Compared with the prior art, the beneficial effects of the present invention are:

[0024] (1) Improve data annotation efficiency: This solution provides visual annotation rule construction, which is simple to operate and use. It provides visual data annotation, and at the same time supports the annotation of instances, relationships, and attributes. Users can annotate more types of data within one task. It also supports real-time preview of the knowledge graph to help users timely discover problems with the annotated data. For the annotated data, annotation corpus or image entry is automatically generated, saving user operations without manual conversion by the user.

[0025] (2) Improve data annotation quality: This solution provides an annotation alignment function. By calculating the similarity of the annotated instances, it intelligently recommends duplicate instances, greatly reducing the entry of redundant data into the image, and thus reducing the data cleaning work after image entry.

[0026] (3)Full - life - cycle management of labeled data: The system manages the entire life cycle of data annotation, from the creation of annotation rules to the management of files to be annotated, the creation of annotation tasks, and finally to data annotation and entry into the map. This reduces the need for users to switch between multiple systems, thereby reducing the workload of users. Brief Description of the Drawings

[0027] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0028] Figure 1 It is a flowchart of the method for the first embodiment. Detailed Description of the Embodiments

[0029] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0030] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "comprising" and "having" and any variations thereof are intended to cover non - exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0031] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0032] All data acquisition in this embodiment is a legal application of data on the basis of compliance with laws, regulations, and user consent.

[0033] Term Explanation: Knowledge Graph: A knowledge graph is a large - scale knowledge network that describes concepts, entities, and their relationships in the objective world in a structured form.

[0034] The First Embodiment

[0035] This embodiment provides a method for annotating unstructured data documents;

[0036] As Figure 1 shown, the method for annotating unstructured data documents includes:

[0037] S101: Build annotation rules; upload the document to be annotated, and review the document to be annotated; create an annotation task;

[0038] S102: Perform data annotation, review the data annotation results, and proceed to S103;

[0039] S103: Determine whether the review passes. If it does, proceed to S104; if not, return to S102;

[0040] S104: Determine whether the task mode is the corpus mode or the image entry mode. If it is the corpus mode, directly generate a corpus from the annotation results; if it is the image entry mode, perform an alignment operation on the annotation results and perform image entry processing on the results after the alignment operation.

[0041] Further, S101: The building of the annotation rules specifically includes:

[0042] S101-1: Add a group to the grouping tree, and set the group name and group path;

[0043] S101-2: Set the entities under each group, and set the entity name, entity identifier, entity path, and attributes of the entity;

[0044] S101-3: Set the relationship between two entities; the relationship between the two entities includes: relationship name and relationship attributes.

[0045] Further, S101: The review of the document to be annotated specifically includes:

[0046] Review whether the format of the document to be annotated is doc format, docx format, txt format, or editable pdf format. If it is, review whether the size of the document to be annotated is less than the set value, and the set value is 5M; if the review shows that the size of the document to be annotated is less than the set value, the document to be annotated can be used for the annotation task, otherwise the review fails.

[0047] Further, S101: The creation of the annotation task, wherein the annotation task includes: a corpus generation task and a knowledge graph generation task.

[0048] Further, S102: The performance of data annotation, and the specific process includes:

[0049] Annotate the entities, relationships between entities, and attributes of entities in the document;

[0050] Allowing graph display during the annotation process;

[0051] During the annotation process, store the relationships between entities in a list;

[0052] During the annotation process, if an entity query instruction is received, the position of the entity in the document is output.

[0053] Further, S102: Review the data annotation results, specifically including:

[0054] According to the entities, attributes, and relationships designed in the ontology construction, review whether the entity category of the annotated instance is correct, whether the attributes of the annotated instance are correct, and whether the relationships between the annotated instances are correct. Mark the data that fails the review with errors and remind the annotator to re-annotate;

[0055] During the review process, if a data modification instruction is received, the data is allowed to be modified.

[0056] Further, S103: Judge whether the review passes. The judgment criterion is

[0057] All the annotation results in the same annotation task are completely correct, and the review passes; otherwise, the review fails.

[0058] Further, S104: Judge whether the task mode is the corpus mode or the graph input mode. If it is the corpus mode, directly generate the corpus from the annotation results. The process of generating the corpus includes:

[0059] Retrieve the annotation results from the database and convert them into a txt text in json format containing instances, relationships, and attributes.

[0060] Further, when generating the corpus from the annotation results, entity corpus, relationship corpus, and attribute corpus are generated according to different uses.

[0061] Exemplarily, entity corpus, relationship corpus, and attribute corpus are generated according to different deep learning models for training. For example, when training a named entity recognition model, entity corpus is generated. When training a relationship recognition model, relationship corpus is generated.

[0062] Further, S104: If it is the graph input mode, perform an alignment operation on the annotation results and perform graph input processing on the results after the alignment operation. The specific process of the alignment operation includes:

[0063] For any two entities of the same entity type in the annotated data, calculate the text similarity between the two entities;

[0064] Align the entities with a text similarity higher than the set threshold.

[0065] Exemplarily, for the text similarity algorithm, the edit distance calculation is selected.

[0066] Further, S104: If it is the in-graph mode, perform an alignment operation on the annotation result, and perform in-graph processing on the result after the alignment operation. The specific process of the alignment operation includes:

[0067] S104-1: Receive an alignment instruction, and display the names, attributes, and relationships of at least two entities to be aligned;

[0068] S104-2: Receive the entity selected by the user from the two entities to be aligned, and save the selected entity;

[0069] S104-3: According to the user's selection, merge or perform the first overwrite on the attributes of the entity; the merge means saving all the attributes of the two entities to be aligned, and the first overwrite means only retaining the attributes of the entity selected by the user;

[0070] S104-4: According to the user's selection, perform merge and duplicate removal or the second overwrite on the relationships of the entity; the merge and duplicate removal means merging the relationships of the two entities to be aligned and removing duplicate relationships; the second overwrite means only retaining the entity relationships selected by the user;

[0071] S104-5: Perform data preview on the names, attributes, and relationships of the aligned entities; save the names, attributes, and relationships of the aligned entities.

[0072] For example, in the annotation data, there are two entities, "A Company" and "A Co., Ltd.", and these two entities actually refer to the same entity. Display a list of entities with similar names in the same entity type in the annotation task.

[0073] Further, S104: If it is the in-graph mode, perform an alignment operation on the annotation result, and perform in-graph processing on the result after the alignment operation. The specific process of the in-graph processing includes:

[0074] After the annotation data is aligned, perform an online graph display on the annotation data, and manually confirm whether the instance types and relationships between the aligned annotation data instances are accurate. After confirmation, directly store the annotation data in the graph database.

[0075] Online preview and graph display of the annotation data. For the data during the annotation process, online preview can be performed and it can be displayed according to the graph. At the same time, the number of annotated instances can be counted, the list of annotated instances and relationships can be displayed, clicking on the annotated instance jumps to the annotation position, and querying instances and relationships by name is supported. The completed annotation data supports online preview of the graph or online preview of the corpus according to the task mode.

[0076] Embodiment 2

[0077] This embodiment provides an unstructured data document annotation system;

[0078] An unstructured data document annotation system, comprising:

[0079] A rule construction module, which is configured to: construct annotation rules; upload a document to be annotated, and review the document to be annotated; create an annotation task;

[0080] A data annotation module, which is configured to: perform data annotation, review the data annotation result, and enter an audit judgment module;

[0081] An audit judgment module, which is configured to: judge whether the audit is passed. If so, enter a mode judgment module; if not, return to the data annotation module;

[0082] A mode judgment module, which is configured to: judge whether the task mode is a corpus mode or an image input mode. If it is the corpus mode, directly generate a corpus from the annotation result; if it is the image input mode, perform an alignment operation on the annotation result, and perform image input processing on the result after the alignment operation.

[0083] It should be noted here that the above-mentioned rule construction module, data annotation module, audit judgment module, and mode judgment module correspond to steps S101 to S104 in Embodiment 1. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1 above. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer executable instructions.

[0084] In the above embodiments, the descriptions of each embodiment have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0085] The proposed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the above-mentioned module division is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0086] Embodiment 3

[0087] This embodiment also provides an electronic device, comprising: one or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, and the above one or more computer programs are stored in the memory. When the electronic device runs, the processor executes the one or more computer programs stored in the memory, so that the electronic device executes the method described in Embodiment 1 above.

[0088] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0089] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0090] In the implementation process, each step of the above method may be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software.

[0091] The method in Embodiment 1 can be directly embodied as being executed and completed by the hardware processor, or executed and completed by the combination of the hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0092] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but this implementation should not be considered to exceed the scope of the present invention.

[0093] Embodiment 4

[0094] This embodiment also provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by the processor, the method described in Embodiment 1 is completed.

[0095] The above is only the preferred embodiment of the present invention and is not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for annotating unstructured data documents, characterized by including: (1) Constructing annotation rules; uploading the document to be annotated, reviewing the document to be annotated; creating an annotation task; (2) Performing data annotation, reviewing the data annotation results, and proceeding to (3); (3) Judging whether the review passes. If it does, proceed to (4); if not, return to (2); (4) Judging whether the task mode is the corpus mode or the image entry mode. If it is the corpus mode, directly generate a corpus from the annotation results. The process of generating the corpus includes: Retrieving the annotation results from the database and converting them into a txt text in json format containing instances, relationships, and attributes; generating entity corpus, relationship corpus, and attribute corpus according to different uses when generating the corpus from the annotation results; If it is the image entry mode, perform an alignment operation on the annotation results and perform image entry processing on the results after the alignment operation. The specific process of the alignment operation includes: calculating the text similarity between any two entities of the same entity type in the annotation data; performing an alignment operation on the entities with a text similarity higher than the set threshold; Among them, the specific process of the alignment operation further includes: Receiving an alignment instruction and displaying the names, attributes, and relationships of at least two entities to be aligned; Receiving the entity selected by the user from the two entities to be aligned and saving the selected entity; Merging the attributes of the entities according to the user's selection. The merging means saving all the attributes of the two entities to be aligned; or, according to the user's selection, performing the first overwrite on the attributes of the entities. The first overwrite means only retaining the attributes of the entity selected by the user; Merging and de-duplicating the relationships of the entities according to the user's selection. The merging and de-duplicating means merging the relationships of the two entities to be aligned and removing duplicate relationships; or, according to the user's selection, performing the second overwrite on the relationships of the entities. The second overwrite means only retaining the entity relationships selected by the user; Performing data preview on the names, attributes, and relationships of the aligned entities; saving the names, attributes, and relationships of the aligned entities.

2. The method for annotating an unstructured data document according to claim 1, wherein The construction of the annotation rules specifically includes: Adding a group to the grouping tree, setting the group name and group path; Setting the entities under each group, setting the entity name, entity identifier, entity path, and attributes of the entity; Setting the relationship between two entities; the relationship between the two entities includes: relationship name and relationship attribute.

3. The method for annotating an unstructured data document according to claim 1, wherein, The review of the document to be annotated specifically includes: reviewing whether the format of the document to be annotated is doc format, docx format, txt format, or editable pdf format. If so, reviewing whether the size of the document to be annotated is less than the set value; if the size of the document to be annotated is less than the set value, the document to be annotated can be used for the annotation task, otherwise the review fails.

4. The method for annotating an unstructured data document according to claim 1, characterized in that, The process of performing data annotation specifically includes: Annotating the entities, relationships between entities, and attributes of the entity in the document; Allowing the display of the knowledge graph during the annotation process; During the annotation process, storing the relationships between entities in a list; During the annotation process, if an entity query instruction is received, the position of the entity in the document is output.

5. The method for annotating an unstructured data document according to claim 1, characterized in that, Review the data annotation results, specifically including: According to the entities, attributes, and relationships designed in the ontology construction, review whether the entity category of the annotated instance is correct, review whether the attributes of the annotated instance are correct, and review whether the relationships between the annotated instances are correct. Mark the data that fails the review with errors and remind the annotator to re-annotate; During the review process, if a data modification instruction is received, the data is allowed to be modified.

6. An unstructured data document annotation system, characterized by including: A rule construction module, which is configured to: construct annotation rules; upload the document to be annotated and review the document to be annotated; create an annotation task; A data annotation module, which is configured to: perform data annotation, review the data annotation results, and enter the review judgment module; A review judgment module, which is configured to: judge whether the review passes. If it does, enter the mode judgment module; if not, return to the data annotation module; A mode judgment module, which is configured to: judge whether the task mode is the corpus mode or the into-graph mode. If it is the corpus mode, directly generate a corpus from the annotation results. The process of generating the corpus includes: Retrieve the annotation results from the database and convert them into a txt text in json format containing instances, relationships, and attributes; generating the corpus from the annotation results, generating entity corpus, relationship corpus, and attribute corpus according to different uses; If it is the into-graph mode, perform an alignment operation on the annotation results and perform into-graph processing on the results after the alignment operation. The specific process of the alignment operation includes: calculating the text similarity between any two entities of the same entity type in the annotated data; performing an alignment operation on the entities with a text similarity higher than the set threshold; Among them, the specific process of the alignment operation also includes: Receiving an alignment instruction and displaying the names, attributes, and relationships of at least two entities to be aligned; Receiving the entity selected by the user from the two entities to be aligned and saving the selected entity; According to the user's selection, merge the attributes of the entities. The merge means saving all the attributes of the two entities to be aligned; or according to the user's selection, perform the first overwrite on the attributes of the entities. The first overwrite means only retaining the attributes of the entity selected by the user; According to the user's selection, merge and deduplicate the relationships of the entities. The merge and deduplication means merging the relationships of the two entities to be aligned and removing duplicate relationships; or according to the user's selection, perform the second overwrite on the relationships of the entities. The second overwrite means only retaining the entity relationships selected by the user; Perform data preview on the names, attributes, and relationships of the aligned entities; save the names, attributes, and relationships of the aligned entities.

7. An electronic device, characterized by including: A memory for non-temporarily storing computer-readable instructions; And A processor for running the computer-readable instructions, Wherein, when the computer-readable instructions are run by the processor, the method described in any one of claims 1-5 above is executed.

8. A storage medium, characterized in that, Non-transitorily store computer-readable instructions, wherein, when the non-transitory computer-readable instructions are executed by a computer, the instructions perform the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Relation extraction and knowledge graph construction method based on deep learning model

    CN110598000A

  • Method for generating knowledge graph and electronic equipment

    CN114417012A