Interactive annotation method for event extraction from semi-structured text

Through an interactive labeling method, table visualization and multiple labeling mechanisms are used to solve the problem of efficiently extracting spatio-temporal event information from semi-structured text, and an efficient and accurate labeling and inspection process is realized, ensuring the reliability of data analysis.

WO2025123406A1PCT designated stage expired Publication Date: 2025-06-19PEKING UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/140826
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-13
Filing Date
2023-12-22
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

There is a lack of effective annotation or inspection tools in the prior art, making it difficult to efficiently extract information with spatiotemporal events from semi-structured text, especially in situations where high reliability of results is required.

Method used

An interactive labeling method is adopted to improve the efficiency and accuracy of users in the annotation, inspection and proofreading process through table visualization, multi-detail hierarchical display, content alignment design and multiple labeling mechanisms. Specific measures include adding spaces to the original text, copying phrases, using wiring strategies, maintaining place concept trees and time dictionaries, and special labeling components for periodic and time-dependent events.

Benefits of technology

Through this method, users can efficiently complete events containing time and space information from text descriptions and check the annotation results, which improves the efficiency and accuracy of the annotation process, reduces the possibility of user errors, and ensures the reliability of subsequent data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023140826_19062025_PF_FP_ABST
    Figure CN2023140826_19062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to an interactive annotation method for event extraction from semi-structured text. Extracted events are annotated in a table format; by adding spaces into original text, and providing assistance by means of copying and connecting using lines, the annotation result is organized into a table and aligned with the original text for display; for special events such as a fuzzy event, a periodic event and a time-dependent event, corresponding annotation mechanisms are respectively used to improve the annotation efficiency; the visualization of the extracted event annotation table provides display of three levels of different details so as to meet different requirements. The use of the method disclosed in the present invention can help users to annotate and extract events containing spatiotemporal information from text descriptions, or check annotation results. By visualizing annotation results and allowing users to provide interactive feedback for the annotation results, the efficiency of annotating and extracting spatiotemporal events from text descriptions is improved and the possibility of user errors is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

An interactive annotation method for event extraction from semi-structured text Technical Field

[0001] The present invention belongs to the field of visualization, and in particular relates to an interactive annotation method for extracting events from semi-structured text. Background Art

[0002] Extracting structured data from text descriptions is a prerequisite for using many computer programs to analyze data, but the flexibility and arbitrariness of natural language brings many difficulties to the process of extracting structured data.

[0003] While the development of deep learning and natural language processing technologies has spurred the development of automated methods for converting text into data, these learning-based methods require large amounts of high-quality labeled data to train the models. This high-quality data relies on manual annotation, review, and modification. Furthermore, in many critical situations, manual review and correction of results are necessary to ensure their high reliability (for example, extracting patient trajectories from epidemiological patient trajectory survey reports).

[0004] There are currently no good annotation or checking tools for the task of annotating semi-structured text, especially for extracting annotations with spatiotemporal events.

[0005] Summary of the Invention

[0006] In response to the defects existing in the prior art, the purpose of the present invention is to provide an interactive annotation method for event extraction from semi-structured text. Through carefully designed table visualization, multi-level detail display, content alignment design, and a variety of annotation mechanisms for special event descriptions, the efficiency and accuracy of user annotation, inspection, and proofreading are improved. During the annotation process, the uncertainty of the spatiotemporal information of events in the text description can be sorted and retained, which facilitates the processing of uncertainty in subsequent data analysis and lays a solid foundation for efficient analysis of data features.

[0007] To achieve the above objectives, the technical solution adopted by the present invention is: an interactive annotation method for event extraction from semi-structured text, which annotates the extracted events in a tabular form, organizes the annotation results into a table and displays them aligned with the original text.

[0008] Furthermore, the annotation results are displayed as a table, and the original text description is inserted between the tables to align with the annotation results.

[0009] Furthermore, based on the state compression dynamic programming method, spaces are added to the original text so that the content of the text after the spaces are added is aligned to the corresponding annotation result table.

[0010] Furthermore, when there is still a small offset in the annotation table content after adding spaces to the original text, a specific phrase in the original text is copied and inserted into the corresponding annotation table position based on the state compression dynamic programming method.

[0011] Furthermore, when a small amount of offset still exists after adding spaces and copying phrases in the original text, a connecting line is used to connect the annotated table content and the corresponding phrases in the text that cannot be completely aligned.

[0012] Furthermore, a line is used to connect the table item and the original phrase, thereby guiding the user to quickly locate the association between the table item and the original phrase.

[0013] Furthermore, for spatial information ambiguity events, a hierarchical location concept tree is maintained and the spatial information filled in the annotation table is mapped to a node on the location concept tree to clarify the spatial information of the spatial information ambiguity events in the original text.

[0014] Furthermore, for events with ambiguous time information, by maintaining a dictionary, the ambiguous time concept is delayed until the downstream task is processed and then mapped to a specific space or time, and interpreted in accordance with the needs of the downstream task.

[0015] Furthermore, for periodic events, a periodic labeling component is used to label the events by specifying the time range and period interval.

[0016] Furthermore, for time-dependent events, a time-dependent event annotation component is used to express the time of an event as a relative relationship with other events using temporal relationship description terms.

[0017] The beneficial technical effect of the present invention is that the interactive annotation method for event extraction from semi-structured text disclosed in the present invention can help users complete the annotation and extraction of events containing spatiotemporal information from text descriptions, or check the annotation results. By visualizing the annotation results and allowing users to provide interactive feedback on the annotation results, the efficiency of the annotation and extraction process of spatiotemporal events from text descriptions is improved, and the possibility of user error is reduced. The uncertainty of the spatiotemporal information of events in text descriptions is sorted and retained during the annotation process, facilitating the processing of uncertainty in subsequent data analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] FIG1 illustrates an interactive annotation method for extracting events from semi-structured text according to the first embodiment of the present invention. Taking epidemiological survey record data of a case as an example, the original survey record text is converted into an event table that records spatiotemporal information.

[0019] FIG2 is a schematic diagram of a process of aligning table items and text in an interactive annotation method for event extraction from semi-structured text according to the first embodiment of the present invention;

[0020] FIG3 is a diagram of a main visual interface generated by an interactive annotation method for event extraction from semi-structured text according to the first embodiment of the present invention;

[0021] FIG4 is a schematic diagram of a process of using different annotation mechanisms to cope with different special events in an interactive annotation method for event extraction from semi-structured text according to the first embodiment of the present invention;

[0022] FIG5 is a diagram of a second-layer display interface generated by an interactive annotation method for event extraction from semi-structured text according to the first embodiment of the present invention;

[0023] FIG6 is a diagram of a third-layer display interface generated by the interactive annotation method for event extraction from semi-structured text according to the first embodiment of the present invention. DETAILED DESCRIPTION

[0024] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0025] Example 1

[0026] An embodiment of the present invention provides an interactive annotation method for extracting events from semi-structured text. The method annotates the extracted events in a table format, organizes the annotation results into a table, and displays the table aligned with the original text.

[0027] The visualization is presented as a table. The original text description is inserted between the tables and aligned with the annotation results, allowing users to map the table content to the text source and check whether the content is correct in context.

[0028] As shown in FIG1 , an interactive annotation method for extracting events from semi-structured text provided by an embodiment of the present invention is used. Taking epidemiological survey record data of cases as an example, the original survey record text is converted into an event table that records spatiotemporal information.

[0029] To align the annotated table with the original text, we employed an alignment strategy combining multiple mechanisms. First, by adding spaces to the text, we were able to align the text content as closely as possible to the corresponding table. However, because the order of event components extracted from the table can differ significantly from the organization of natural language narratives, simply adding spaces to the text might not satisfy the alignment requirements. Therefore, we added a copy strategy to change the order of phrases, and finally, a connection strategy was used as a fallback, using visual encoding to assist users in associating unaligned table items and text.

[0030] The copy strategy is used to change the order of phrases so that more phrases can be aligned with the table at the same time. The copy strategy copies a phrase in a sentence and reinserts it elsewhere in the sentence. For example, the sentence to be annotated is "It was already 3 o'clock when he left home for the hotel," and the items to be filled in the annotation table are "time," "starting point," and "end point." At this time, if we only add spaces in the sentence, we will not be able to align "3 o'clock" with "time," "home" with "starting point," and "hotel" with "end point" at the same time. Therefore, we can copy 3 o'clock and the sentence will become "It was already 3 o'clock when he (3 o'clock) left home for the hotel." This gives us the opportunity to align the corresponding phrases in the sentence with the table by inserting spaces later.

[0031] The connection strategy is to connect the table items with the corresponding phrases after copying and inserting spaces fail, so as to guide users to quickly locate the association between the table items and the phrases.

[0032] As shown in Figure 2, during implementation, a state-compression-based dynamic programming method is used. First, the table items and text are aligned as much as possible through the two strategies of inserting spaces and copying, and then lines are added between the table items and text that are not fully aligned. In the state-compression-based dynamic programming method, phrases are tried to be aligned in the order of table items from front to back. The state is set to f[i][j][k]=x, which means that we have considered the first i table items, the last table item aligned by non-copying means is the jth item, and the number of aligned items in the first i table items is k. In all cases where these conditions are met, the remaining space left from the jth item to the ith item that can be flexibly arranged is optimal (maximum) and the remaining space is x. If x is a negative value, the state is illegal (does not exist). Three cases are considered when transitioning to a state:

[0033] 1) The current (i-th item) is directly aligned with the corresponding phrase, provided that the phrase between j and i can be accommodated in the x space and transferred to f[i+1][i][k+1]=0.

[0034] 2) The current (i-th item) is aligned directly with the corresponding copy of the phrase, unconditionally. Transfer to f[i+1][j][k+1] = x.

[0035] 3) The current (i-th item) is not aligned, so it is unconditional. Move to f[i+1][j][j]=x+|width of table item i|.

[0036] Finally, the maximum number of items k that can be aligned simultaneously is the largest k that can be obtained by tracing the transition records of dynamic programming.

[0037] Figure 3 shows the main visual interface generated by the interactive annotation method for event extraction from semi-structured text, as described in an embodiment of the present invention. In this interface, the original text and the annotation result table are aligned and displayed. Both the table and text are enhanced with color coding to distinguish the different components of the event. The text and corresponding content in the table are aligned, allowing users to quickly navigate from the table to the context and verify the accuracy of the table content. To ensure that the text can be read in normal order during alignment, a copy and line connection mechanism is implemented.

[0038] In the description text, in addition to conventional phrases that can be directly extracted, there are some information that is not directly represented. In an embodiment of the present invention, as shown in Figure 4, different annotation mechanisms are used to deal with different special events, including fuzzy events, periodic events, and time-dependent events.

[0039] In text descriptions, the spatiotemporal information of events is often vaguely described in the form of concepts. For example, spatial information is described using place names (such as a certain province, a certain city, or even some non-standard place names), which have varying degrees of ambiguity. To more accurately preserve and organize the ambiguity of event spatial information in the original text, we maintain a hierarchical location concept tree for events with ambiguous spatial information and map the spatial information entered in the table to a node in the location concept tree. This clarifies the spatial information of ambiguous events in the text.

[0040] Time information, such as words like morning, afternoon, and evening, has multiple uncertain interpretations. For events with ambiguous time information, by maintaining a dictionary, the ambiguous time concept can be delayed until downstream tasks are processed and mapped to a specific space or time, and interpreted in accordance with the needs of downstream tasks.

[0041] In a text description, a sentence may describe a periodic event. For periodic events, a periodic annotation component is used to annotate them by specifying the time range and period interval of the event.

[0042] In text descriptions, the time of an event may not be explicitly expressed, but may depend on the time of other events, such as "after doing thing A, I did thing B." For time-dependent events, a time-dependent event annotation component is used to express the time of an event as a relative relationship with other events using temporal relationship description terms such as "before," "after," and "complement."

[0043] The visualization of the annotation extraction event table supports users to efficiently check the annotated results. In order to improve the inspection efficiency, the visualization of the annotation extraction event table provides three levels of display with different levels of detail to cope with large amounts of data. The first level displays all the detailed information, and the visualization shows the annotation result table and the original text at the same time. This display level also serves as the main interface for the annotation process. However, during the inspection, the user may want to view the annotation results as a whole first. At this time, you can switch to the second or third level display interface, as shown in Figure 5. In the second level of detail display, the original text description is omitted and only the result table is displayed, as shown in Figure 6. In the third level of detail display, it only shows whether each grid in the table is filled in.

[0044] It can be seen from the above embodiments that the present invention discloses an interactive annotation method for event extraction from semi-structured text, which annotates the extracted events in a table form, organizes the annotation results into a table and displays them aligned with the original text by adding spaces to the original text and assisting with copying and connecting. For special events such as fuzzy events, periodic events, and time-dependent events, corresponding annotation mechanisms are used to improve the annotation efficiency. The visualization of the annotation extraction event table provides three levels of display with different levels of detail to meet different needs. The method disclosed in the present invention can help users complete the annotation and extraction of events containing spatiotemporal information from text descriptions, or check the annotation results. By visualizing the annotation results and allowing users to interactively feedback on the annotation results, the efficiency of the annotation and extraction process of spatiotemporal events from text descriptions is improved, and the possibility of user errors is reduced.

[0045] The method described in the present invention is not limited to the embodiments described in the specific implementation manner. Those skilled in the art may derive other implementation manners based on the technical solution of the present invention, which also fall within the scope of the technical innovation of the present invention.

Claims

1. An interactive annotation method for event extraction from semi-structured text. The method annotates the extracted events in a tabular form, organizes the annotation results into a table, and displays them aligned with the original text.

2. The interactive annotation method for event extraction from semi-structured text according to claim 1, characterized in that: The overall annotation results are presented in the form of a table, and the original text descriptions are inserted between the tables and aligned with the annotation results.

3. The interactive annotation method for event extraction from semi-structured text according to claim 2, characterized in that: Based on the state compression dynamic programming method, by adding spaces to the original text, the content in the text after adding spaces is aligned to the corresponding annotation result table.

4. The interactive annotation method for event extraction from semi-structured text according to claim 3, characterized in that: When there are still a small number of offsets in the content of the annotation table after adding spaces to the original text, based on the state compression dynamic programming method, a copy of a specific phrase in the original text is inserted into the corresponding annotation table position.

5. The interactive annotation method for event extraction from semi-structured text according to claim 4, characterized in that: When there are still a small number of offsets after adding spaces and copying phrases in the original text, lines are used to connect the content of the annotation table that cannot be fully aligned and the corresponding phrases in the text.

6. The interactive annotation method for event extraction from semi-structured text according to claim 5, characterized in that: Lines are used to connect the table items and the original phrases, so as to guide users to quickly locate the association between the table items and the original phrases.

7. The interactive annotation method for event extraction from semi-structured text according to claim 6, characterized in that: For spatial information ambiguity events, by maintaining a hierarchical location concept tree and corresponding the spatial information filled in the annotation table to a node on the location concept tree, the spatial information of the spatial information ambiguity events in the original text is clarified.

8. The interactive annotation method for event extraction from semi-structured text according to claim 7, characterized in that: For time information ambiguity events, by maintaining a dictionary, the fuzzy time concept is delayed until before the downstream task processing and then corresponding to the specific space or time, and the paraphrase is carried out in combination with the requirements of the downstream task.

9. The interactive annotation method for event extraction from semi-structured text according to claim 8, characterized in that: For periodic events, a periodic annotation component is adopted, and the annotation is carried out by specifying the time range and the period interval of the event.

10. The interactive annotation method for event extraction from semi-structured text according to claim 8, characterized in that: For time-dependent events, a time-dependent event annotation component is adopted, and the time of the event is expressed as a relative relationship with other events using time-relationship description terms.

Citation Information

Patent Citations

  • Structural analysis method and system for judicial documents

    CN111145052A

  • Vertical domain knowledge graph construction method and system

    CN113177124A

  • Event extraction method and system for multi-modal financial document

    CN114881015A

  • Method and system for processing investment and research report in futures field

    CN115358201A

  • System and method of textual information analytics

    US20060277465A1