A sequence annotation method for unstructured railway knowledge entities
By introducing a sequence labeling method of unstructured railway knowledge entities in the railway field, the problem of lack of labeling corpus is solved, and rapid and accurate labeling and corpus construction is achieved, laying the foundation for railway natural language processing and smart railway knowledge graph.
Patent Information
- Application Number
- CN202210693657.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-19
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-06-19
AI Technical Summary
The lack of professional unstructured railway entity labeling corpus in the railway field makes it difficult to conduct efficient training of entity recognition in natural language processing technology, affecting the construction of smart railway knowledge maps.
Provide a sequence labeling method for unstructured railway knowledge entities, including unstructured data import, pattern selection design, custom label design, database construction and entity search, and accurately label railway entities through BIO or BIOES sequence labeling methods, and build a professional railway vocabulary library and corpus text library.
It realizes the rapid and accurate acquisition of multi-modal sequence annotations of unstructured railway knowledge entities, provides a large number of labeling corpuses, and provides a foundation for natural language processing research in the railway field and the construction of smart railway knowledge graphs.
Smart Images

Figure CN115309892B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent railway knowledge graphs, aiming to provide accurate and large amounts of unstructured annotation data for natural language processing technology in the railway field, collect corpus texts for railway entity annotation, and lay a foundation for constructing an intelligent railway knowledge graph. Specifically, it is a sequence annotation method for unstructured railway knowledge entities. Background Art
[0002] With the development of science and technology, the era of railway intelligence is coming. Intelligent railway projects in China have also attracted a great deal of attention, such as the intelligent Beijing-Zhangjiakou Railway. The National Railway Intelligent Transportation System Engineering and Technology Research Center has also been established in China, and the overall framework of the railway intelligent transportation system has been proposed. At the same time, artificial intelligence technology has also developed rapidly in recent years. Artificial intelligence includes a computing layer, a perception layer, and a cognition layer. Artificial intelligence is a complete imitation of humans and has the ability to understand and think. The knowledge graph is the basis of the cognition layer. Graph technology has achieved many accomplishments in industries such as medicine and finance, but the rail transit industry is still in its infancy. There are still many problems worth exploring and many problems that need to be solved in the research on constructing an intelligent railway knowledge graph in the railway field. Entity recognition in natural language processing technology is a necessary link in constructing a knowledge graph. To achieve railway entity recognition, a large number of railway entity annotation corpus training sets are required; so far, no website or merchant has provided professional unstructured railway entity annotation corpus, and obtaining unstructured railway entity annotation corpus is still a time-consuming and laborious project. Summary of the Invention
[0003] Aiming at the above problems, the purpose of the present invention is to provide a method that can accurately and quickly obtain unstructured railway entity annotation corpus, provide underlying materials for natural language processing research in the railway field, and lay a foundation for constructing an intelligent railway knowledge graph.
[0004] The technical solution of the present invention is as follows: A sequence annotation method for unstructured railway knowledge entities, including the following steps,
[0005] Step 1, Import of unstructured data;
[0006] Form an import function for TXT / JSON files and an interface for file import and operation guide buttons according to the file format of unstructured railway knowledge entities;
[0007] Step 2, Design of mode selection;
[0008] According to different sequence annotation methods for unstructured railway knowledge entities, the sequence annotation method adopts BIO or BIOES;
[0009] Step 3, Design of custom tags;
[0010] According to the annotation requirements for unstructured railway knowledge entity tags, a custom tag function is designed based on the BIO and BIOES sequence annotation methods;
[0011] Step 4: Construct a database and a railway professional vocabulary library;
[0012] Allocate multiple blank databases to add the unstructured railway knowledge entities to be annotated, ensuring the connection between the unstructured railway knowledge entities and the sequence annotation function, enabling the sequence annotation function to accurately retrieve and call the added unstructured railway knowledge entities, use MySQL to construct a background railway professional vocabulary library, and at the same time add the unstructured railway knowledge entities in the database to the background railway professional vocabulary library;
[0013] Step 5: Entity retrieval and construction of a railway corpus text library;
[0014] Use the BIO and BIOES sequence annotation functions in Step 2 or the custom sequence annotation function in Step 3 to accurately retrieve and add unstructured railway knowledge entity tags to complete sequence annotation; at the same time, use MySQL to construct a background railway corpus text library, and save the imported pure text containing railway knowledge entities to the background railway corpus text library according to the classification of unstructured railway knowledge entities after each use for reverse acquisition of pure text later;
[0015] Step 6: Design for importing fields and statistical display of railway annotation entities;
[0016] Design field and annotation entity statistical functions. When performing sequence annotation, use the statistical functions to count the total number of text fields, the types of unstructured railway knowledge entities to be annotated and their specific numbers, and display them in the form of a bar chart on the statistical display page;
[0017] Step 7: Design for importing and exporting files in multiple formats:
[0018] For unstructured data, import and export in accordance with TXT and JSON files.
[0019] Furthermore, the BIO sequence annotation method specifically is to label each element in the unstructured railway knowledge entity as B, I, or O, where B is the start, I is the middle, and O is others.
[0020] Furthermore, the custom tag is to expand B and I to B-X and I-X to obtain the custom tag. B-X indicates that the segment where the element is located belongs to the X type and the element is at the beginning of the segment where it is located, I-X indicates that the segment where the element is located belongs to the X type and the element is in the middle position of the segment where it is located, and O indicates not belonging to any type.
[0021] Further, the BIOES sequence annotation method is specifically as follows: B - Begin represents the start, I - Intermediate represents the middle character, E - End represents the end, S - Single represents a single character, and O - Other represents other for marking irrelevant characters.
[0022] Further, the BIOES custom tags are defined by adding extended attributes after B, I, O, E, and S to complete the custom tags.
[0023] Further, it allows users to custom - add unstructured railway knowledge entities to be annotated, allows users to customize tags in addition to using the two common annotation methods of BIO and BIOES, and allows users to import and export in TXT or JSON format; the professional vocabulary library of railway entities in the background collects and annotates keywords and is regularly added to and deleted by professionals to provide railway entity references for users; it saves the unannotated corpus imported by users according to railway annotation entity classification, and can reversely obtain the pure corpus containing railway entities through keywords.
[0024] The beneficial effects of the present invention are as follows: The present invention can perform large - scale multi - mode sequence annotation of unstructured railway knowledge entities quickly and accurately, and can trace back to relevant pure texts through railway entities, providing training corpus for natural language processing research in the railway field and laying a foundation for constructing an intelligent railway knowledge graph. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 The flow and schematic diagram of the technical solution of the present invention.
[0026] Figure 2 The sequence annotation diagram of unstructured railway knowledge entities in the technical solution of the present invention.
[0027] Figure 3 The pure text back - tracing diagram of unstructured railways in the technical solution of the present invention.
[0028] Figure 4 The schematic diagram of the main interface functions in the technical solution of the present invention.
[0029] Figure 5 The schematic diagram of tag customization in the technical solution of the present invention.
[0030] Specific usage
[0031] The following further describes the present invention in detail with reference to the drawings and specific embodiments.
[0032] Embodiment 1
[0033] The present invention realizes one - key sequence labeling of unstructured railway entities, providing underlying labeled corpus for natural language processing research in the railway field; it allows users to customize and add unstructured railway knowledge entities to be labeled, allows users to use not only the two common labeling methods of BIO and BIOES but also customize tags, and allows users to import and export in TXT or JSON format; the professional vocabulary library of railway entities in the background includes labeled keywords and is regularly added to and deleted by professionals to provide railway entity references for users; it saves the unlabeled corpus imported by users according to the classification of railway labeled entities, and can reverse - obtain the pure corpus containing railway entities through keywords.
[0034] The unstructured data of the present invention means that the data in a computer information system is divided into structured data and unstructured data. Unstructured data has irregular or incomplete data structures, no predefined data models, and is inconvenient to be represented by a database two - dimensional logical table. It includes all formats of office documents, texts, etc. In this embodiment, it is temporarily only for text documents in TXT and JSON formats.
[0035] The knowledge entities of the present invention are specific things. Knowledge entities can be people, places, organizations, concepts, etc. The railway knowledge entities in the present invention refer to professional vocabulary or equipment vocabulary in the railway field, such as: switch machine, balise, station interlocking, signal machine, etc.
[0036] The document data of the present invention without labeling operations, here it is an unlabeled TXT or JSON document, that is, a pure text without any labels.
[0037] The tags of the present invention are keywords with strong category relevance, which can easily describe and classify content for easy retrieval and sharing. For example, the most typical tags are personal names, place names, equipment names, etc. Entity tags refer to keywords for knowledge entities that classify knowledge entities, such as on - vehicle equipment, ground equipment, etc.
[0038] The sequence labeling (Sequence labeling) of the present invention is a basic problem often encountered when solving NLP problems. In sequence labeling, if a label is assigned to each knowledge entity in a sequence, generally speaking, a sequence refers to a sentence, and a knowledge entity refers to a word in the sentence.
[0039] Here, for BIO sequence labeling, each element is labeled as "B", "I", or "O", where "B" represents the beginning, "I" represents the middle, and "O" represents others. By expanding "B" and "I" into "B-X" and "I-X", custom labels are obtained. "B-X" indicates that the segment where this element is located belongs to type X and this element is at the beginning of this segment, "I-X" indicates that the segment where this element is located belongs to type X and this element is in the middle of this segment, and "O" indicates not belonging to any type. For example, if X is represented as a noun phrase (Noun Phrase, NP), then the three tags of BIO are:
[0040] (1)B-NP: The beginning of a noun phrase
[0041] (2)I-NP: The middle of a noun phrase
[0042] (3)O: Not a noun phrase
[0043] Therefore, a passage can be divided into the following results;
[0044] For example: Pierre Vinken, 61 years old, will join IBM’s board as anonexecutive director Nov.29.
[0045] [NP Pierre Vinken], [NP 61 years] old, will join [NP IBM]’s [NP board] as [NP a nonexecutive director] [NP Nov.29].
[0046] Pierre_B-NP Vinken_I-NP, _O 61_B-NP years_I-NP old_O, _O will_O join_O IBM_B-NP ’s_O board_B-NP as_O a_B-NP nonexecutive_I-NP director_I-NP Nov._B-NP 29_I-NP._O
[0047] BIOES is similar to BIO:
[0048] B, that is, Begin, represents the beginning
[0049] I, that is, Intermediate, represents the middle
[0050] E, that is, End, represents the end
[0051] S, which stands for Single, represents a single character
[0052] O, which stands for Other, represents others and is used to mark irrelevant characters
[0053] Similar to BIO, BIOES can also be extended to generate custom tags
[0054] For the labeled corpus, the data document generated after sequence labeling is still in the TXT or JSON format in the present invention. However, compared with the pure text or corpus, a labeling sequence is added after each word or character. For example, a document contains a sentence: Teacher Brown gave a speech yesterday. It includes a knowledge entity: Brown, and the label "person name" is labeled into the phrase "Brown". Among them, the document with the content "Teacher Brown gave a speech yesterday." is a pure text or corpus, while the content of the TXT or JSON document generated by labeling is "yesterday_O day_O cloth_B-person name lang_I-person name old_O teacher_O publish_O table_O le_O yi_O ci_O speech_O 。_O", and this document is called a labeled corpus
[0055] The process and principle of the labeling method are as Figure 1 shown, including constructing a TXT / JSON format file import function and interface; then designing a mode selection interface and guiding buttons; then designing BIO and BIOES labeling functions and interfaces according to common labeling methods and constructing a custom tag function mode and related interfaces and buttons; constructing a railway knowledge entity addition function and interface according to usage requirements; constructing a railway professional vocabulary database; completing sequence labeling through entity retrieval and classification and constructing a railway pure corpus database; counting and displaying the number of fields and entities; designing the export of TXT / JSON labeling sequence files. Through the above steps, by establishing various functional functions, interfaces and databases, the automatic labeling of unstructured railway entities and the collection and retrieval of pure text are realized
[0056] Users can import the TXT / JSON file to be labeled, select BIO / BIOES or custom tags, add the railway knowledge entities to be labeled, and can view the statistical display of text fields and entities with one key to obtain the labeling sequence file required for entity recognition. In addition, the pure text database also classifies and includes pure corpora containing entities in real time, and users can reverse obtain a large number of pure TXT / JSON texts containing the required entities through railway entity keywords
[0057] The specific steps are as follows
[0058] Step 1: Design for importing unstructured data
[0059] According to the common file formats of unstructured railway knowledge entities, an import function for TXT / JSON files and related interface and button designs are designed.
[0060] As Figure 4 shown, click "Browse" on the main interface to access the disk files in the computer, select the TXT or JSON file that needs to be annotated. The TXT or JSON file here is the pure text containing unstructured railway knowledge entities, and import it into the annotation tool.
[0061] For example, import a TXT document here with the file content "The switch machine is an important device in railway transportation".
[0062] Step 2: Mode selection design.
[0063] According to different sequence annotation methods for unstructured data entities, the two most common BIO / BIOES annotation methods are provided here for users to choose. The sequence annotation method of this application is the sequence annotation mode adopted, which is implemented through BIO / BIOES annotation functions.
[0064] As Figure 4 shown, you can click to select the common BIO or BIOES mode in sequence annotation as needed for text annotation. BIO annotation labels each element as "B-X", "I-X", or "O". Among them, "B-X" means that the segment where this element is located belongs to type X and this element is at the beginning of this segment, "I-X" means that the segment where this element is located belongs to type X and this element is in the middle position of this segment, and "O" means that it does not belong to any type. For BIOES annotation, B - Begin means start, I - Intermediate is the middle character, E - End means end, S - Single means single character, and O - Other means other for marking irrelevant characters.
[0065] Step 3: Custom label design.
[0066] According to the annotation requirements for entity labels, users are allowed to define relevant labels independently and select the custom labels as the annotation mode.
[0067] In the entity recognition task, the labels are not fixed. You can use the custom button on the main interface, such as Figure 5 , to define the labels as more detailed named entity recognition labels and some suffixes can be added after the defined labels, such as: B-Person, B-Location, etc. These can all be selected according to the actual task, but all basic custom rules should conform to BIO or BIOES.
[0068] For BIO sequence labeling, expand the "B" and "I" in it to "B-X" and "I-X". "B-X" means that the segment where this element is located belongs to type X and this element is at the beginning of this segment. "I-X" means that the segment where this element is located belongs to type X and this element is in the middle position of this segment. "O" means it does not belong to any type. For example, if X is represented as a noun phrase (Noun Phrase, NP), then the three tags of BIO are:
[0069] (1) B-NP: The beginning of a noun phrase
[0070] (2) I-NP: The middle of a noun phrase
[0071] (3) O: Not a noun phrase
[0072] Therefore, a passage can be divided into the following results;
[0073] Specifically, such as:
[0074] Zhang San is twenty-six years old this year and is an outstanding lawyer.
[0075] After annotation:
[0076] Zhang [B-NP] San [I-NP] Jin [B-NP] Nian [I-NP] Er [B-NP] Shi [I-NP] Liu [I-NP] Sui [I-NP], [O] Shi [O] Yi [O] Ming [O] Chu [O] Se [O] De [O] Lv [B-NP] Shi [I-NP]. [O]
[0077] The principle of BIOES custom tags is similar to that of BIO. Custom tags are completed by adding extended attribute definitions after B, I, O, E, and S.
[0078] For example, if the content to be marked here is "switch machine" in the TXT document of step 1, then define: B-S as the start of railway key equipment words, I-S as the middle part of railway key equipment words, O as others, and this tag will be used in step 5.
[0079] Step 4: Build a user database and a railway professional vocabulary database.
[0080] Set aside multiple blank user databases for users to add the railway knowledge entities to be marked, and ensure the connection between these entities and functions, so that the sequence labeling function can accurately retrieve and call the railway knowledge entities added by users. At the same time, add the railway knowledge entities in the user database to the background railway professional vocabulary database.
[0081] Such as Figure 4As shown in the figure, users can add key railway knowledge entities in this annotation task by themselves through the add button on the main interface. These entities will be temporarily stored in the blank user database in the background for the search keywords of this annotation task, such as RBC, station interlocking, etc. At the same time, MySQL builds a railway professional vocabulary database, which will classify and collect annotation entities and be regularly revised, added and deleted by professionals. The database also serves as a professional vocabulary prompt. Here, the railway professional vocabulary database will include the term "switch machine".
[0082] Step 5: Entity retrieval and railway corpus text library construction.
[0083] The specific location of the unstructured railway knowledge entity is obtained through the designed entity retrieval, and the corresponding tags are accurately added to complete the sequence annotation. At the same time, after each user use, the clean text containing the railway knowledge entity imported by the user is saved to the background railway corpus text library according to the keyword classification, so as to obtain the clean text in the later reverse. The clean text here is the unlabeled TXT or JSON document imported in step 1.
[0084] like Figure 4 As shown in the figure, after completing the above 4 steps, click the start button to perform one-click entity annotation. The background program function realizes entity classification and position determination by searching the text keywords, gives corresponding labels in turn, and builds the corresponding sequence annotation file. At the same time, a pure corpus database is constructed, and the corresponding fields are classified and collected into the database according to the keywords. Through the corpus backtracking button, you can search for all the segments containing keyword entities in the database by keywords, and export them in TXT or JSON format files.
[0085] In this step, the pure corpus database will include the TXT document before annotation, that is, the TXT document imported in step 1, and then generate a new annotated document through annotation, whose content is: "Switch [BS] track [IS] machine [IS] is [O] an [O] important [O] equipment [O] in [O] railway [O] transportation [O]. [O] ".
[0086] Step 6: Import fields and railway annotation entity statistics display design:
[0087] When the user imports the document and adds the required annotated railway knowledge entities, a statistical function will be used to count the total text fields and the types and specific numbers of annotated railway knowledge entities during sequence annotation in the background, and display them on the statistical display page in the form of a bar chart.
[0088] When performing sequence annotation, call the designed field and annotation entity statistical functions to count the word length of the imported file, and define different entity counting flags according to the added entities. When each label is assigned, the corresponding flag is counted. After the annotation is completed, the word count and entity count of the text passage are reflected in the form of a chart. As shown in Table 1, here the number of fields in the TXT document with the content "The switch machine is an important device in railway transportation." will be counted, which is 15, and the number of unstructured railway knowledge entities "switch machine" will be shown as 1.
[0089]
[0090] Step 7: Design for multi-format import and export of files.
[0091] This method designs the import and export of TXT and JSON files for unstructured data.
[0092] The TXT document exported here is the marked document, that is, the TXT document with the content "The switch machine is an important device in railway transportation."
Claims
1. A sequence annotation method for unstructured railway knowledge entities, characterized in that: It includes the following steps: Step 1, unstructured data import; Form an import function for TXT / JSON files according to the file format of unstructured railway knowledge entities, and import the files into the unstructured railway knowledge entity annotation tool. The unstructured railway knowledge entity annotation tool exists in the form of an interface. The operation guide buttons in the interface include file import, annotation mode selection, annotation entity, and corpus backtracking; Step 2, mode selection design; Design the sequence annotation mode for unstructured railway knowledge entities. The sequence annotation mode adopts BIO or BIOES; The BIO sequence annotation mode is specifically as follows: each element in the unstructured railway knowledge entity is annotated as B, I, or O, where B is the beginning, I is the middle, and O is others; Step 3, custom label design; According to the annotation requirements for the labels of unstructured railway knowledge entities, design the sequence annotation mode of the labels based on the BIO and BIOES sequence annotation modes; The custom label is to expand B and I to B-X and I-X to obtain the custom label. B-X means that the segment where the element is located belongs to the X type and the element is at the beginning of the segment where it is located. I-X means that the segment where the element is located belongs to the X type and the element is in the middle position of the segment where it is located. O means not belonging to any type; Step 4, construct a database and a railway professional vocabulary library; Set aside multiple blank databases to add the unstructured railway knowledge entities to be annotated. Use MySQL to construct a background railway professional vocabulary library, and add the unstructured railway knowledge entities in the database to the background railway professional vocabulary library; Step 5, entity retrieval and construction of a railway corpus text library; Use the BIO and BIOES sequence annotation modes in Step 2 or the custom sequence annotation mode in Step 3 to accurately retrieve and add the labels of unstructured railway knowledge entities to complete the sequence annotation; at the same time, use MySQL to construct a background railway corpus text library. After each use, save the imported pure text containing railway knowledge entities to the background railway corpus text library according to the classification of unstructured railway knowledge entities for later reverse acquisition of pure text; Step 6, import field and railway annotation entity statistics display design; Design a field and annotation entity statistics function. When performing sequence annotation, use the statistics function to count the total number of text fields, the types of unstructured railway knowledge entities to be annotated, and their specific numbers, and display them in the form of a bar chart on the statistics display page; Step 7, file multi-format export design; For unstructured data railway knowledge entities, export them according to TXT and JSON files.
2. The sequence annotation method for unstructured railway knowledge entities according to claim 1, characterized in that: The BIOES sequence annotation mode is specifically as follows: B - Begin represents the start, I - Intermediate is the middle character, E - End represents the end, S - Single represents a single character, and O - Other is used to mark irrelevant characters.
3. The sequence annotation method for unstructured railway knowledge entities according to claim 1, characterized in that: BIOES custom tags are defined by adding extended attributes after B, I, O, E, and S to complete custom tags.
Citation Information
Patent Citations
Geological intelligent question and answer-oriented data automatic sequence labeling recognition method
CN111930909A
Knowledge extraction method, apparatus, electronic device, and storage medium
WO2021212682A1