Railway official document keyword extraction method and device and electronic equipment

By combining a format rule library, a railway-specific terminology library, and the TF-IDF algorithm in railway official documents, the problems of missed detection of fixed-format information and incorrect terminology splitting in railway official documents were solved, and efficient and accurate keyword extraction and terminology updating were achieved.

CN120706422APending Publication Date: 2025-09-26INST OF COMPUTING TECH CHINA ACAD OF RAILWAY SCI +2
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510830648.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Traditional keyword extraction technology in railway official documents has problems such as missed detection of fixed-format information, high error rate in professional terminology splitting, failure to strengthen the semantic weight of key positions, uninterpretable models, and difficulty in updating.

Method used

A pre-built railway document format rule library is used to extract fixed-position key fields through regular expression matching and position locking. The Jieba word segmenter is used to load the railway-specific terminology library for word segmentation. The boundaries of multi-word combination entities are dynamically corrected through dependency rules. The TF-IDF algorithm is combined for position weighting, and the mutual information and left-right entropy algorithms are used to dynamically update the terminology library.

Benefits of technology

It achieves high-precision and high-efficiency automatic keyword extraction, improves accuracy, reduces term recognition error rate, dynamically adapts to industry evolution, and shortens the time spent on warehousing new terms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706422A_ABST
    Figure CN120706422A_ABST
Patent Text Reader

Abstract

The invention relates to a railway official document keyword extraction method and device and electronic equipment, and the method comprises the steps: based on a pre-constructed railway official document format rule base, extracting a key field of a fixed position from an input text through regular expression matching and position locking; a Jieba word segmentation device is used for loading a railway-specific term library for word segmentation, and a multi-word combination entity boundary is dynamically corrected through a dependency relationship rule; executing a TF-IDF algorithm on the text after word segmentation to generate an initial word weight, adjusting the weight according to the position area of the word in the official document and a preset coefficient, and performing position weighting; and combining the words of which the weights are greater than a set threshold value with the extracted key fields, and outputting a final keyword set after verification of a term library. According to the method, missing detection caused by low frequency of a traditional algorithm is avoided, splitting errors of a general word segmentation device are eliminated, the term recognition error rate is reduced, the core word sorting priority is improved, and the semantic weight of keywords is strengthened; the new term storage time is shortened, and the updating cost problem is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a method, device and electronic equipment for extracting keywords from railway official documents. Background Art

[0002] As the railway industry accelerates its informatization process, the need for intelligent processing of railway documents, the core vehicle for management decision-making, is becoming increasingly urgent. Railway documents are characterized by highly standardized formats, specialized terminology, and complex semantic dependencies.

[0003] Traditional general keyword extraction technology faces serious bottlenecks in railway scenarios: although traditional statistical methods (such as TF-IDF) have high computational efficiency, they have difficulty identifying domain terms and key information in fixed positions, resulting in missed detection of fixed-format information; although deep learning models have high accuracy in general texts, they have problems such as model black box, high update costs, and weak industry knowledge transfer, and cannot adapt to the dynamic update needs of railway terminology.

[0004] As a graph-sorting-based keyword extraction solution, the TextRank algorithm has technical flaws that are particularly prominent in railway scenarios. Specifically, it has the following manifestations: poor industry adaptability: it relies only on word co-occurrence relationships and fails to integrate the strong semantic weight characteristics of railway document titles, first paragraphs, and endings; terminology recognition failure: it incorrectly breaks down multi-word entities into isolated words, destroying semantic integrity; and lacks interpretability: it cannot embed manual rules, resulting in business personnel being unable to intervene and correct results. Summary of the Invention

[0005] Based on this, it is necessary to provide a railway official document text keyword extraction method, device and electronic equipment to address the technical problems of traditional keyword extraction technology, such as missed detection of fixed format information, high error rate in professional term splitting, unenhanced semantic weights of key positions, uninterpretable models, and difficult updating.

[0006] The present invention provides a method for extracting keywords from railway official documents, the method comprising: Based on the pre-built railway document format rule library, the system extracts key fields at fixed locations from the input text through regular expression matching and position locking, including the issuing unit, document number, document category, and issuance time. Jieba word segmenter is used to load railway-specific terminology library for word segmentation, and dependency rules are used to dynamically modify the entity boundaries of multi-word combinations. The terminology library contains railway facility names, institutional entities, and technical standard terms. The TF-IDF algorithm is used to generate initial word weights for the segmented text. The weights are adjusted according to the preset coefficients based on the word's location in the document, and position weighting is performed. The weight of words in the title area is ×1.5, and the weight of words in the attachment area is ×0.8. Words with weights greater than a set threshold are merged with the extracted key fields, and the final keyword set is output after verification by the terminology library. The terminology library is dynamically updated using the mutual information and left-right entropy algorithms. New terms are screened with a mutual information threshold > 5.0 and a left-right entropy threshold > 2.5.

[0007] In one embodiment, the construction of the railway-specific terminology library includes: Collect the names of institutional entities and technical terms in railway industry standard documents to form a basic terminology set; Dynamically expand new terms through the execution of mutual information and left-right entropy algorithms, including: calculating the mutual information value and left-right entropy value of candidate words, screening word combinations with mutual information > 5.0 and left-right entropy > 2.5, and adding the word combinations to the term library.

[0008] In one embodiment, the execution of the mutual information and left-right entropy algorithm includes: Perform stop word filtering on the candidate words to eliminate invalid combinations of function words, including , and ; The Trie tree structure is used to store the term base to optimize the efficiency of multi-word matching.

[0009] In one embodiment, the construction of the format rule library includes: Use regular expressions to match railway document numbers and set document number extraction rules; Lock the organization name after the issuing unit on the first line of text and set the issuing unit positioning rules.

[0010] In one embodiment, the dependency rule dynamically modifies the multi-word entity boundary, including: When a high-frequency suffix is ​​detected, the adjacent nouns are traced back and combined into a complete entity. The high-frequency suffix includes enterprise, bureau and department; Temporarily add the combined entity to the word segmentation dictionary for Jieba word segmenter to call.

[0011] In one embodiment, in the position weighting, the title area is defined as the centered text in the first line of the document, with a weighting coefficient of 1.5; the first paragraph area is defined as the first paragraph, with a weighting coefficient of 1.3; the end area is defined as the last paragraph containing the text of "this order" and "this notice" with a weighting coefficient of 1.3; and the attachment area weighting coefficient is 0.7.

[0012] In one embodiment, the step of outputting the final keyword set further includes: Record new word combinations and their frequency that are not covered by the terminology database; When the frequency is greater than 10 times per 10,000 documents, the mutual information and left-right entropy calculations are triggered to generate term base update suggestions.

[0013] In one embodiment, the step of outputting the final keyword set further includes: Provides a manual intervention interface to support users to add or delete keywords or adjust weights; Based on manual operation logs, optimize the regular expressions and dependency rule thresholds of the format rule library.

[0014] The present invention also provides a device for extracting keywords from railway official documents, the device comprising: The format rule extraction module is used to extract key fields at fixed positions from the input text, including the issuing unit, document number, document category, and issuance time, based on the pre-built railway document format rule library through regular expression matching and position locking; The domain term segmentation module is used to load the railway-specific terminology library using the Jieba word segmenter for word segmentation and dynamically correct the boundaries of multi-word combination entities through dependency rules. The terminology library contains railway facility names, institutional entities, and technical standard terms. The weighted calculation module is used to execute the TF-IDF algorithm on the text after word segmentation to generate initial word weights. The weights are adjusted according to the preset coefficients based on the position of the words in the official document, and position weighting is performed. Among them, the weight of the words in the title area is ×1.5, and the weight of the words in the attachment area is ×0.8; The rule filtering output module is used to merge words with weights greater than a set threshold with the extracted key fields, and output the final keyword set after verification by the terminology library. The terminology library is dynamically updated through the mutual information and left-right entropy algorithms, and new terms are screened when the mutual information threshold is greater than 5.0 and the left-right entropy threshold is greater than 2.5.

[0015] The present invention also provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and is characterized in that the processor implements the above-mentioned railway official document text keyword extraction method when executing the computer program.

[0016] The above-mentioned railway official document text keyword extraction method, device and electronic equipment directly extract document number, unit and other railway official document-specific structured fields through the format rule library, avoiding missed detection due to low frequency of traditional algorithms and extracting fixed position information more accurately; the term library and dependency rules force multi-word entities to be processed as complete units, eliminating the splitting errors of general word segmenters, reducing the error rate of term recognition, and ensuring the integrity of professional terminology; the position weighting mechanism fits the industry characteristics of railway official documents with dense decision instructions in the first paragraph and supplementary explanations in the attachments, which improves the ranking priority of core words, strengthens the semantic weight of keywords, and solves the problem of weight distortion; the semi-automatic update of the term library based on mutual information and left and right entropy shortens the time taken to enter new terms into the database, dynamically adapts to industry evolution, and solves the problem of high update costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 A flow chart of a method for extracting keywords from railway official documents according to an embodiment; Figure 2 A flow chart of a method for extracting keywords from railway official documents according to another embodiment; Figure 3 A flow chart of a method for extracting keywords from railway official documents according to another embodiment; Figure 4 A flow chart of a method for extracting keywords from railway official documents according to another embodiment; Figure 5 This is a flow chart of a fifth embodiment of the method for extracting keywords from railway official documents of the present invention; Figure 6 This is a flowchart of a sixth embodiment of the method for extracting keywords from railway official documents of the present invention; Figure 7 This is a flow chart of the seventh embodiment of the method for extracting keywords from railway official documents of the present invention; Figure 8 This is a schematic diagram of a railway official document keyword extraction device according to one embodiment; Figure 9 FIG. 1 is a diagram showing the internal structure of a computer device according to an embodiment. DETAILED DESCRIPTION

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0020] As the railway industry accelerates its informatization, the need for intelligent processing of railway documents, the core vehicle for management and decision-making, is becoming increasingly urgent. Railway documents are characterized by highly standardized formats (e.g., fixed document numbers, units, and categories), highly specialized terminology (e.g., "CTCS level" and "track circuit"), and complex semantic dependencies (e.g., the multi-word entity "XX Bureau Dispatching Office").

[0021] Traditional general keyword extraction technology faces serious bottlenecks in railway scenarios: traditional statistical methods (such as TF-IDF) have difficulty balancing accuracy and efficiency. Although they have high computational efficiency, they have difficulty identifying domain terms (such as missing "contact network maintenance order") and key information in fixed locations (such as document numbers being ignored due to low frequency), resulting in missed detection of fixed-format information; although deep learning models have high accuracy in general texts, they have problems such as model black boxes, high update costs, and weak industry knowledge transfer, and cannot adapt to the dynamic update needs of railway terminology (such as the new term "hydrogen energy train" requires retraining the model).

[0022] As a keyword extraction solution based on graph sorting, the TextRank algorithm has technical defects that are particularly prominent in railway scenarios. Specifically, it has the following manifestations: poor industry adaptability: it only relies on word co-occurrence relationships and fails to integrate the strong semantic weight features of the title / first paragraph / end of railway documents (for example, "dispatching order" should be weighted when it appears in the first paragraph); terminology recognition failure: multi-word combination entities (such as "Zhengzhou Bureau Freight Department") are incorrectly decomposed into isolated words ("Zhengzhou Bureau" "Freight Department"), destroying semantic integrity; lack of interpretability: it is impossible to embed manual rules (such as term base verification and position weighting), resulting in business personnel being unable to intervene in the result correction.

[0023] The present invention targets the unique structured format, professional terminology system and semantic features of administrative instructions of railway official documents, and realizes high-precision and high-efficiency automatic keyword extraction by integrating traditional statistical models with domain knowledge rule base. It is suitable for intelligent management of official documents, information retrieval and decision support systems in the railway industry.

[0024] The following combination Figures 1 to 9 The present invention describes a method, device and electronic device for extracting keywords from railway official documents.

[0025] like Figure 1 As shown, in one embodiment, a method for extracting keywords from railway official documents includes the following steps: In step S110, based on a pre-built library of railway document format rules, regular expression matching and position-locking are used to extract key fields at fixed locations from the input text, including the issuing unit, document number, document category, and issuance time. Pre-written "document template rules" (including regular expressions) are used to directly capture the issuing unit, document number, and other information from fixed locations in the document (for example, the document number is always on the second line), thus achieving fixed-location extraction of key information.

[0026] In step S120, the Jieba word segmenter loads a railway-specific terminology library for word segmentation. The terminology library includes railway facility names, organizational entities, and technical standard terms. Jieba word segmentation uses a railway-specific dictionary (including terms like "locomotive depot" and "contact network") to avoid breaking down specialized terms like "XX Bureau Dispatching Office" into "XX Bureau / Dispatching Office" to avoid fragmentation.

[0027] Step S130: Execute the TF-IDF algorithm on the segmented text to generate initial word weights. The weights are adjusted based on the word's location in the document using a preset coefficient, performing position weighting. Here, the weight of words in the title area is multiplied by 1.5, and the weight of words in the attachment area is multiplied by 0.8. A weighted TF-IDF calculation is performed. The TF-IDF algorithm is first executed on the segmented text to generate initial word weights. Then, position weighting is performed, adjusting the weights based on the word's location in the document (title, first paragraph, end, attachment) using a preset coefficient. When calculating word importance, words in the title are multiplied by 1.5 (the title is most important), words in the beginning and end areas are multiplied by 1.3, words in attachments are multiplied by 0.8 (attachments are less important), and other words are multiplied by 1.0, doubling the weight of words in important locations.

[0028] As an optional weighting, the title area is defined as the centered text on the first line of the document, with a weighting factor of 1.5; the first paragraph area is defined as the first paragraph, with a weighting factor of 1.3; the end area is defined as the final paragraph containing the text "This Order" or "This is to inform you," with a weighting factor of 1.3; and the attachment area has a weighting factor of 0.7. Therefore, the "Immediate Suspension of Operations" in the first paragraph has an 86% higher weight than the "Statistical Table" in the attachment, which is more consistent with the characteristic of railway documents where decision-making information is concentrated in the header.

[0029] In step S140, words with weights greater than a set threshold are merged with the extracted key fields. After verification with the terminology database, a final set of keywords is output. The terminology database is dynamically updated using a mutual information and left-right entropy algorithm. New terms are screened when a mutual information threshold > 5.0 and a left-right entropy threshold > 2.5 are met. When a new term (such as "intelligent train control") is discovered, the algorithm (mutual information > 5.0 + left-right entropy > 2.5) automatically determines whether to add it to the dictionary, allowing for dynamic dictionary updates.

[0030] The railway official document keyword extraction method of this embodiment directly extracts structured fields unique to railway documents, such as document numbers and units, through a format rule library (coordinated by regular expressions and position locking). This avoids missed detections caused by low frequency in traditional algorithms. In actual measurements, the accuracy of fixed field extraction has increased from 58% to 99%, allowing for more accurate extraction of fixed position information. The terminology library and dependency rules force multi-word entities to be processed as complete units, eliminating splitting errors by general word segmenters and reducing the term recognition error rate by 40%, ensuring the integrity of professional terminology. The position weighting mechanism adapts to the industry characteristics of railway documents, which often have dense decision-making instructions in the first paragraph and supplementary explanations in the appendices. This increases the sorting priority of core words (such as "dispatching command") by 32%, strengthens the semantic weight of keywords, and solves the problem of weight distortion. The semi-automatic update of the terminology library based on mutual information and left-right entropy reduces the time required to enter new terms from three days of manual annotation to two hours, dynamically adapting to industry evolution and solving the problem of high update costs.

[0031] like Figure 2 As shown, in one embodiment, the construction of the railway-specific terminology database includes the following steps: Step S122: Collect the names of institutional entities and technical terms in railway industry standard documents to form a basic term set.

[0032] Step S124 dynamically expands new terms by executing the mutual information and left-right entropy algorithm, including: calculating the mutual information value and left-right entropy value of candidate words, screening word combinations with mutual information > 5.0 and left-right entropy > 2.5, and adding the word combination to the term library. Mutual information (PMI) is used to measure the strength of the association between two words (consecutive combinations). If two words often appear together (such as "high-speed rail" and "rail"), the PMI value will be very high. A high PMI (for example, > 5.0) indicates that the two words have a strong co-occurrence relationship and may be part of a compound term. A high entropy value indicates that there are many words that can be paired with the left and right sides of the word, indicating that the word is relatively independent and may be a complete term. A low entropy value indicates that the word always appears with a fixed word and may be part of a long term.

[0033] We first collected terminology from railway standard documents (such as "turnout" and "EMU") as a foundation. We then used algorithms to scan massive amounts of text: Mutual Information > 5.0 for word combinations that often appear together, such as "high-speed rail" and "rail," and Left-Right Entropy > 2.5 for independent terms that can be followed by different words, such as "contact network." This prevented omissions in manual collation and automatically discovered new terms, such as "Fuxing Hao Intelligent EMU," increasing terminology coverage by 65%. Before new terms were added to the terminology database, they underwent manual semantic verification to eliminate ambiguous combinations. New terms recommended by the algorithm (such as "sleeper hanging") required confirmation by engineers to minimize errors in the terminology database and ensure the integrity of the professional terminology.

[0034] Optionally, the execution of the mutual information and left - right entropy algorithm includes: Perform stop - word filtering on candidate words to eliminate invalid combinations of function words, where the function words include "de", "he", etc. For example, filter out "de" in "maintenance of the catenary" (function words have no meaning).

[0035] Use a Trie tree structure to store the term library and optimize the matching efficiency of multi - word terms. Store terms using a Trie tree (an efficient retrieval tree) to quickly query long terms such as "Qinghai - Tibet Railway Company" in seconds.

[0036] Through the optimization of the term library, the term matching speed is 3 times faster, the memory occupancy is reduced by 40%, and invalid combinations (such as "of the fault") no longer interfere with the judgment.

[0037] Such as Figure 3 As shown, in one embodiment, the construction of the format rule library includes the following steps: Step S112, match the railway official document number by regular expression and set the number extraction rule. For example, for the number rule example: Tie Gui Zi 〔2023〕123 Hao, match it with the regular expression [Tie|Yun] Gui Zi 〔\d{4}〕\d+ Hao.

[0038] Step S114, lock the organizational name after the document - issuing unit in the first line of the text and set the document - issuing unit positioning rule. For example, the document - issuing unit rule: lock 20 characters after "Document - issuing unit:".

[0039] Through the construction of the format rule library, even if the official document format is slightly adjusted (such as adding or reducing spaces), accurate extraction can still be achieved, and moreover, the rules are interpretable and engineers can modify and adapt them at any time.

[0040] Such as Figure 4 As shown, in one embodiment, the dependency rule dynamically corrects the multi - word combination entity boundary, including the following steps: Step S126, when a high - frequency suffix word is detected, trace back to the adjacent noun forward and combine them into a complete entity. The high - frequency suffix words include "enterprise", "bureau", and "department".

[0041] Step S128, temporarily add the combined entity to the word - segmentation dictionary for the Jieba word - segmenter to call.

[0042] Specifically, when suffixes such as "bureau", "department", and "enterprise" are detected, merge forward: for example, when encountering "Shanghai Bureau", check if there is "railway" in front. If so, merge them into "Shanghai Railway Bureau". After merging, temporarily put it into the word - segmentation dictionary so that the word - segmenter can recognize it next time. By this way of merging terms, phrases like "Zhengzhou Bureau Electrical Section" are no longer cut into three parts, optimizing the accuracy of organizational entity recognition (up to 97%, 82% before optimization).

[0043] Such as Figure 5 As shown, in one embodiment, after outputting the final keyword set, the following steps are further included: Step S510: record new word combinations not covered by the term library and their occurrence frequencies.

[0044] Step S520: When the frequency is greater than 10 times per 10,000 documents, the mutual information and left-right entropy calculations are triggered to generate a term base update suggestion.

[0045] For words that the recording system does not recognize (such as "magnetic levitation switch"), when the word appears more than 10 times in 10,000 documents, the terminology database entry process is automatically started, thereby achieving automatic updating of the terminology database and timely entry of new terms, avoiding the dictionary lagging behind technological development.

[0046] like Figure 6 As shown, in one embodiment, after outputting the final keyword set, the following steps are further included: Step S610: Provide a manual intervention interface to support users to add or delete keywords or adjust weights.

[0047] Step S620: Optimize the regular expression and dependency rule threshold of the format rule library based on the manual operation log.

[0048] Users are allowed to manually delete incorrect keywords (such as the mistakenly extracted "copy"). The system records user actions and automatically optimizes. For example, if multiple people delete "XX copy," the extraction threshold is raised. If "Beijing / Bureau" is frequently merged, "Beijing Bureau" is added to the rule. After one manual correction, the system autonomously learns and optimizes the rules, achieving human-machine collaborative optimization and shortening the rule base iteration cycle (from one week to one hour).

[0049] The format rule library is configured differently by document type. Specifically, for dispatch order documents, the "command number" and "execution time" are enhanced for extraction; for accident reports, the "accident level" and "responsible unit" are enhanced for extraction. In accident reports, "major accident" and "XX locomotive depot" are automatically highlighted, ensuring precise matching of extraction targets across different document types.

[0050] like Figure 7 As shown, in one embodiment, the automatic keyword extraction algorithm and the manually constructed terminology library and rule library are as follows: After the text enters the algorithm, it first uses the format rule library to extract keyword information such as the date of issuance, issuing unit, document number, and document category. This information has fixed location and format, making it highly important, but it is difficult to obtain a high weight in the TF-IDF algorithm. Therefore, format rules are used for extraction. The text is then segmented using the Jieba algorithm, and the Harbin Institute of Technology stop word list is used to remove frequent and meaningless words from the text. To prevent the segmentation algorithm from incorrectly segmenting specialized terminology in the railway document field, a manually annotated term library is used as the segmentation algorithm's dictionary to ensure that specialized terms are segmented into complete terms. A dependency rule library is used to supplement the dictionary to ensure accurate segmentation. The TF-IDF algorithm assigns a weight to each word in the segmented text. The weights of words in special positions are multiplied by a coefficient, and words with weights greater than a threshold are identified as text keywords.

[0051] After the TF-IDF algorithm assigns a weight to each word, railway official documents are highly structured, with key information often located in specific locations (such as the first paragraph, title, and end). Therefore, the algorithm incorporates a position-weighted mechanism based on the word weights derived from the basic TF-IDF calculation. After calculating the initial weight for each word using the TF-IDF algorithm, different adjustment coefficients are assigned based on the word's position in the document. For example, in railway official documents, positions such as the first paragraph, the beginning of a paragraph, and the end of a document often carry core thematic information or decision-making instructions. Therefore, words appearing in these key positions are given higher keyword value. For words in these structural positions, their TF-IDF weights are multiplied by a coefficient greater than 1 (e.g., 1.2-1.5) to increase their priority in the final candidate keyword ranking. Conversely, for words appearing in less informative locations, such as attachments or mid-paragraph, the weight coefficients are appropriately reduced or maintained. This results in higher-quality keyword extraction that is more tailored to industry characteristics.

[0052] The following is an introduction to the term base and rule base and their construction methods: Termbase construction: The construction of a terminology database is the foundation of manual annotation and relies on the expertise and experience of the railway industry. To accurately identify railway terminology, railway industry experts, experienced staff, and scholars are required to collect and organize common railway terminology, build a domain terminology database, and dynamically expand it.

[0053] This invention collects high-frequency terms (such as "track circuit" and "CTCS level") from standard documents such as the "Railway Technical Management Regulations" and combines mutual information with left-right entropy algorithms to discover new terms, creating a custom dictionary containing thousands of terms. The key to discovering new terms by combining mutual information with left-right entropy algorithms lies in quantifying the internal cohesion and external contextual freedom of terms. Mutual information measures the degree of cohesion within a term; higher values ​​indicate more stable combinations, such as the joint probability of "smart" and "dispatching" in the candidate term "intelligent dispatch." Left-right entropy assesses the independence of terms within context; higher values ​​indicate clearer boundaries. For example, the probability distribution of "production" being adjacent to "system safety" on the left and "scheduling" on the right. By setting thresholds (e.g., mutual information > 5.0 and left-right entropy > 2.5), new terms with high cohesion and low redundancy (such as "hydrogen energy train") can be screened. Furthermore, a Trie tree structure is used to optimize candidate word expansion and stop word filtering, improving multi-word recognition capabilities.

[0054] In addition, the team also conducted in-depth research on management departments, corporate units and business processes at all levels of the railway system, and systematically sorted out the names of the institutions at the National Railway Group level, the names of its subordinate enterprises, and the names of various railway bureaus and stations. In addition, to ensure the comprehensiveness of the proper noun database, the team also specially collected the names of the main leaders of the railway system as a supplement to the terminology database.

[0055] After completing the collection of specialized terminology, the next step is to standardize all noun entries. This process primarily involves two aspects: first, standardizing the writing format of proper nouns, such as the correspondence between full and abbreviations, capitalization, and spacing, in accordance with railway industry practices and official document formatting requirements; second, converting the existing Word document data for proper nouns and their corresponding abbreviations into a unified JSON format to facilitate the subsequent use of the proper noun error correction algorithm.

[0056] These terms, including railway facility names, operating procedures, technical standards, accident types, and dispatching orders, constitute the core keywords of railway official documents. By selecting words that reflect the purpose, objective, task, decision, and action instructions of railway official documents, we ensure that key concepts and domain knowledge in the documents are accurately extracted, thereby helping to establish a deep understanding and effective management of railway official documents. Keywords can be divided into the following categories, as shown in the table below: Topic keywords Such as "report", "decision", "notice", "opinion", etc. Action and decision-making keywords Such as "production", "implementation", "arrangement", "requirement", "approval", "decision", etc. Time and place keywords Such as "plan", "notification date", "meeting time", etc. Keywords: government, department, institution Such as "Railway Bureau", "Ministry of Transport", "Commission", etc. Legal and policy keywords Such as "laws", "policies", "regulations", "standards", etc. Decision-making and approval keywords Such as "approval", "examination", "pass", etc. Rule base construction: Railway documents are characterized by standardized formats, fixed terminology, and strong semantic dependencies. Because the TF-IDF algorithm calculates importance solely based on word frequency, errors and omissions are inevitable. To ensure accurate and professional keyword extraction, it is necessary to incorporate the writing characteristics of railway industry documents and prior knowledge to refine the results through the establishment of a rule base. By applying these rules, the system can more accurately identify and extract key information from railway documents, improving the efficiency and accuracy of information processing.

[0057] Aiming at the actual demand for automatic keyword screening of railway official documents, the present invention designs and implements a rule base specifically oriented to the text structure of railway official documents, in which the rules are divided into two categories, including format rules and dependency rules.

[0058] In railway document processing, key information such as the issuance date, issuing unit, document number, and document category often follows fixed format specifications. This information not only plays a vital role in document management, archiving, retrieval, and circulation, but also serves as the foundation for understanding the document's purpose, decision-making basis, and administrative processes. However, in the TF-IDF algorithm, this information often struggles to receive high weighting. Therefore, developing and applying formatting rules can efficiently and accurately identify structured information within documents, improving the integrity of keyword extraction and information extraction. By specifically extracting elements with formatting characteristics, we can effectively reduce misjudgments caused by contextual ambiguity and ensure the standardization and reliability of railway document information processing.

[0059] The formatting rules are constructed as follows: 1. Collect and analyze sample documents: Collect samples of railway official documents covering various categories, systematically sort out and analyze the format and structure of the sample documents, and summarize the expression format and location distribution patterns of fields such as the issuance time, issuing unit, document number, and document category.

[0060] 2. Write regular expressions: Based on the features obtained through analysis, such as location and format, use regular expressions, template matching, and other methods to convert manually summarized patterns into rules that can be executed by computers.

[0061] 3. Testing and iterative optimization: The constructed rules are repeatedly verified in the sample library to correct ambiguities and omissions in the rules. New formats are continuously optimized and supplemented based on real official documents to improve the robustness and universality of the rules.

[0062] 4. Rule maintenance and update: Update new expressions and rules in a timely manner based on actual usage and the needs of subsequent algorithms to ensure the continued effectiveness of the rule base.

[0063] Railway official documents frequently feature proper nouns (such as "XX Company," "XX Enterprise," and "XX Dispatching Office") that are often combinations of multiple common words. These combinations possess unique semantic meanings and accurately describe key information such as units, departments, and institutions. However, due to the rapid pace of industry terminology updates and diverse combinations, terminology libraries struggle to capture all possible expressions. Conventional word segmentation algorithms can easily misclassify these proprietary entities into multiple unrelated words, resulting in the segmented vocabulary failing to accurately reflect the original meaning. Dependency rules are designed to supplement and correct specific combinations that exist in context but are not covered by the terminology library. This ensures precise word boundary segmentation and effectively maintains information consistency and integrity.

[0064] The dependency rules are constructed as follows: 1. Summarizing Lexical Co-occurrence Patterns. First, we conducted text co-occurrence statistics and dependency mining on a large number of railway official document samples to identify high-frequency proper noun combination patterns (e.g., "XX Co., Ltd.", "X City Y Station," "Ministry of Railways Engineering Department," etc.).

[0065] 2. Domain Experience and Grammatical Features Summary: Based on practical experience in the railway industry, we summarized the naming patterns of common combinations of unit names, department structures, and equipment and facilities, such as "place name + organization" and "brand + vehicle model." We also categorized special naming forms and summarized their word order, syntactic structure, and common suffix characteristics.

[0066] 3. Write rule expressions and optimize them. Based on the above model, manually set specific rules such as word segmentation priority, combination division, and boundary judgment. For example, for "XX enterprise," first determine the keyword "enterprise," then trace back and identify named entities based on statistical characteristics. Combine them with "enterprise" to form a complete word. Then, temporarily add this word to the term library to achieve word segmentation and split it into complete words.

[0067] 4. Rule maintenance and update: Based on actual usage and the needs of subsequent algorithms, timely update rules and related technologies to ensure the continued effectiveness of the rule base.

[0068] The following describes the railway official document text keyword extraction device provided by the present invention. The railway official document text keyword extraction device described below and the railway official document text keyword extraction method described above can be referenced to each other.

[0069] like Figure 8 As shown, in one embodiment, a device for extracting keywords from railway official documents includes a format rule extraction module 810 , a domain term segmentation module 820 , a weighted calculation module 830 and a rule filtering output module 840 .

[0070] The format rule extraction module 810 is used to extract key fields at fixed positions from the input text, including the issuing unit, document number, document category and issuing time, through regular expression matching and position locking based on the pre-built railway document format rule library.

[0071] The domain term segmentation module 820 is used to use the Jieba word segmenter to load the railway-specific terminology library for word segmentation, and dynamically correct the boundaries of multi-word combination entities through dependency rules. The terminology library contains railway facility names, institutional entities and technical standard terms.

[0072] The weighted calculation module 830 is used to execute the TF-IDF algorithm on the text after word segmentation to generate initial word weights, adjust the weights according to the preset coefficients based on the position area of ​​the words in the official document, and perform position weighting, where the weight of the words in the title area is ×1.5, and the weight of the words in the attachment area is ×0.8.

[0073] The rule filtering output module 840 is used to merge words with weights greater than a set threshold with the extracted key fields, and output the final keyword set after verification by the term library. The term library is dynamically updated through the mutual information and left-right entropy algorithm, and new terms are screened when the mutual information threshold is >5.0 and the left-right entropy threshold is >2.5.

[0074] Figure 9 The following is a schematic diagram of the physical structure of an electronic device. The electronic device may be a smart terminal, and its internal structure diagram may be as follows: Figure 9 As shown. The electronic device includes a processor, a memory, and a network interface connected via a system bus. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a method for extracting keywords from railway official documents is implemented, which includes: "Notice on Removal of Exclusive Content".

[0075] Those skilled in the art will understand that the structure shown in Figure x is merely a block diagram of a partial structure related to the solution of the present invention, and does not constitute a limitation on the electronic device to which the solution of the present invention is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0076] On the other hand, the present invention also provides a computer storage medium storing a computer program, which, when executed by a processor, implements a method for extracting keywords from railway official documents, the method comprising: Based on the pre-built railway document format rule library, the system extracts key fields at fixed locations from the input text through regular expression matching and position locking, including the issuing unit, document number, document category, and issuance time. Jieba word segmenter is used to load railway-specific terminology library for word segmentation, and dependency rules are used to dynamically modify the entity boundaries of multi-word combinations. The terminology library contains railway facility names, institutional entities, and technical standard terms. The TF-IDF algorithm is used to generate initial word weights for the segmented text. The weights are adjusted according to the preset coefficients based on the word's location in the document, and position weighting is performed. The weight of words in the title area is ×1.5, and the weight of words in the attachment area is ×0.8. Words with weights greater than a set threshold are merged with the extracted key fields, and the final keyword set is output after verification by the terminology library. The terminology library is dynamically updated using the mutual information and left-right entropy algorithms. New terms are screened with a mutual information threshold > 5.0 and a left-right entropy threshold > 2.5.

[0077] In another aspect, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and when the processor executes the computer instructions, implements a method for extracting keywords from railway official documents, the method comprising: Based on the pre-built railway document format rule library, the system extracts key fields at fixed locations from the input text through regular expression matching and position locking, including the issuing unit, document number, document category, and issuance time. Jieba word segmenter is used to load railway-specific terminology library for word segmentation, and dependency rules are used to dynamically modify the entity boundaries of multi-word combinations. The terminology library contains railway facility names, institutional entities, and technical standard terms. The TF-IDF algorithm is used to generate initial word weights for the segmented text. The weights are adjusted according to the preset coefficients based on the word's location in the document, and position weighting is performed. The weight of words in the title area is ×1.5, and the weight of words in the attachment area is ×0.8. Words with weights greater than a set threshold are merged with the extracted key fields, and the final keyword set is output after verification by the terminology library. The terminology library is dynamically updated using the mutual information and left-right entropy algorithms. New terms are screened with a mutual information threshold > 5.0 and a left-right entropy threshold > 2.5.

[0078] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory.

[0079] By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0080] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0081] The above-described embodiments merely illustrate several embodiments of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make various modifications and improvements without departing from the spirit of the present invention, and these modifications and improvements fall within the scope of the present invention. Therefore, the scope of the present invention shall be determined by the appended claims.

Claims

1. A method for extracting keywords from railway official documents, characterized in that: The method comprises: Based on the pre-built railway document format rule library, the system extracts key fields at fixed locations from the input text through regular expression matching and position locking, including the issuing unit, document number, document category, and issuance time. Jieba word segmenter is used to load railway-specific terminology library for word segmentation, and dependency rules are used to dynamically modify the entity boundaries of multi-word combinations. The terminology library contains railway facility names, institutional entities, and technical standard terms. The TF-IDF algorithm is used to generate initial word weights for the segmented text. The weights are adjusted according to the preset coefficients based on the word's location in the document, and position weighting is performed. The weight of words in the title area is ×1.5, and the weight of words in the attachment area is ×0.

8. Words with weights greater than a set threshold are merged with the extracted key fields, and the final keyword set is output after verification by the terminology library. The terminology library is dynamically updated using the mutual information and left-right entropy algorithms. New terms are screened with a mutual information threshold > 5.0 and a left-right entropy threshold > 2.

5.

2. The method for extracting keywords from railway official documents according to claim 1, characterized in that: The construction of the railway-specific terminology database includes: Collect the names of institutional entities and technical terms in railway industry standard documents to form a basic terminology set; Dynamically expand new terms through the execution of mutual information and left-right entropy algorithms, including: calculating the mutual information value and left-right entropy value of candidate words, screening word combinations with mutual information > 5.0 and left-right entropy > 2.5, and adding the word combinations to the term library.

3. The method for extracting keywords from railway official documents according to claim 2, characterized in that: The execution of the mutual information and left-right entropy algorithm includes: Perform stop word filtering on the candidate words to eliminate invalid combinations of function words, including , and ; The Trie tree structure is used to store the term base to optimize the efficiency of multi-word matching.

4. The method for extracting keywords from railway official documents according to claim 1, characterized in that: The construction of the format rule library includes: Use regular expressions to match railway document numbers and set document number extraction rules; Lock the organization name after the issuing unit on the first line of text and set the issuing unit positioning rules.

5. The method for extracting keywords from railway official documents according to claim 1, characterized in that: The dependency rules dynamically modify the multi-word entity boundaries, including: When a high-frequency suffix is ​​detected, the adjacent nouns are traced back and combined into a complete entity. The high-frequency suffix includes enterprise, bureau and department; Temporarily add the combined entity to the word segmentation dictionary for Jieba word segmenter to call.

6. The method for extracting keywords from railway official documents according to claim 1, characterized in that: In the position weighting, the title area is defined as the centered text in the first line of the document, with a weighting coefficient of 1.5; the first paragraph area is defined as the first paragraph, with a weighting coefficient of 1.3; the end area is defined as the last paragraph containing the text of "this order" and "this notice", with a weighting coefficient of 1.3; and the attachment area has a weighting coefficient of 0.

7.

7. The method for extracting keywords from railway official documents according to claim 1, characterized in that: The output of the final keyword set further includes: Record new word combinations and their frequency that are not covered by the terminology database; When the frequency is greater than 10 times per 10,000 documents, the mutual information and left-right entropy calculations are triggered to generate term base update suggestions.

8. The method for extracting keywords from railway official documents according to claim 1, characterized in that: The output of the final keyword set further includes: Provides a manual intervention interface to support users to add or delete keywords or adjust weights; Based on manual operation logs, optimize the regular expressions and dependency rule thresholds of the format rule library.

9. A railway official document keyword extraction device, characterized in that: The device comprises: The format rule extraction module is used to extract key fields at fixed positions from the input text, including the issuing unit, document number, document category, and issuance time, based on the pre-built railway document format rule library through regular expression matching and position locking; The domain term segmentation module is used to load the railway-specific terminology library using the Jieba word segmenter for word segmentation and dynamically correct the boundaries of multi-word combination entities through dependency rules. The terminology library contains railway facility names, institutional entities, and technical standard terms. The weighted calculation module is used to execute the TF-IDF algorithm on the text after word segmentation to generate initial word weights. The weights are adjusted according to the preset coefficients based on the position of the words in the official document, and position weighting is performed. Among them, the weight of the words in the title area is ×1.5, and the weight of the words in the attachment area is ×0.8; The rule filtering output module is used to merge words with weights greater than a set threshold with the extracted key fields, and output the final keyword set after verification by the terminology library. The terminology library is dynamically updated through the mutual information and left-right entropy algorithms, and new terms are screened when the mutual information threshold is greater than 5.0 and the left-right entropy threshold is greater than 2.

5.

10. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the method for extracting keywords from railway official documents according to any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Enterprise culture hot word analysis method based on content tracing

    CN121435967A

  • Enterprise culture hot word analysis method based on content traceability

    CN121435967B