A method for processing document abstract content based on XML fragmentation

By converting documents to XML format and fragmenting them, the problem of traditional data media failing to meet the content organization and service needs of the multi-media and cross-media digital age is solved. This achieves high-accuracy metadata and automatic text annotation, improving the efficiency and accuracy of document abstract content processing.

CN116150346BActive Publication Date: 2026-05-01南方电网能源发展研究院有限责任公司
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
南方电网能源发展研究院有限责任公司
Filing Date
2022-11-23
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Traditional data media cannot meet the content organization and service needs of the multi-media and cross-media digital age, especially in documents such as government agencies and think tank research reports, where it is difficult to achieve remote collaborative writing, personalized customization, and automated editing.

Method used

The documents are converted into XML format and divided into fragmented data units. Data content models are formed according to the patterns of articles, chapters, sections and topics. Dynamic association and reorganization are carried out through keyword semantic relationships. A knowledge-enhanced representation model is established using XML technology, and automated annotation and dynamic association network construction are performed.

Benefits of technology

It achieved an accuracy rate of >95% for automatic metadata indexing and >90% for automatic XML annotation of the main text, improving the accuracy of target extraction and reducing the cost of text decomposition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116150346B_ABST
    Figure CN116150346B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on XML fragmentation literature abstract content processing method, including the literature being converted into XML format is divided into fragmented data unit, according to four kinds of mode of chapter, chapter, section and theme constitutes data content model, data content model is dynamically associated by key word semantic relation, extract fragmented data application unit in key word and subject content unit, content unit is dynamically reorganized according to literature abstract demand.The application has beneficial effect that metadata automatic indexing accuracy is >95%, text Xml automatic annotation accuracy is >90%, can effectively improve the prerequisite of target extraction accuracy under the condition of reducing the cost of text decomposition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of text processing technology, and in particular to a method for processing document summary content based on XML fragmentation. Background Technology

[0002] Text from various types and formats of documents issued by government agencies at all levels, including official documents, public speeches, policy documents, think tank research reports, project plans, and project feasibility studies, is broken down into text fragments with a tree-like index structure. This process provides a data foundation for subsequent semantic extraction and multi-fingerprint querying.

[0003] In the multi-media digital age, traditional data media cannot meet the needs of report writers for remote collaborative writing, personalized customization, intelligent recognition, and automated editing during content organization and service processes. Therefore, breaking the constraints of traditional processes and concepts and establishing a dynamic report generation mechanism based on content objects, collaborative work, and "one-time production, multiple distribution" has become a key technology. Summary of the Invention

[0004] The purpose of this section is to outline some aspects of embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.

[0005] In view of the problems existing in the above and / or existing XML-based document summary content processing methods, the present invention is proposed.

[0006] Therefore, the problem to be solved by this invention is how to provide a method for processing document abstract content based on XML fragmentation.

[0007] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a method for processing document abstract content based on XML fragmentation, which includes converting the document into XML format;

[0008] The documents converted into XML format are divided into fragmented data units and composed of data content models according to four modes: article, chapter, section, and topic.

[0009] By dynamically associating data content models through keyword semantic relationships, keyword and topic content units are extracted from fragmented data application units;

[0010] The content units are dynamically reorganized according to the requirements of the document abstract.

[0011] As a preferred embodiment of the document abstract content processing method based on XML fragmentation of the present invention, wherein: XML format is Extensible Markup Language, that is, a data storage language;

[0012] By using XML technology to establish a knowledge-enhanced representation model, knowledge resources can be fragmented at both the formal structure and content levels. Based on the characteristics and utilization methods of various knowledge resources, content fragmentation can be gradually realized.

[0013] As a preferred embodiment of the document abstract content processing method based on XML fragmentation of the present invention, the fragmented data unit is a data application unit form in which the document is fragmented into formulas, charts, paragraphs, chapters, content outlines, interactive operations, interactive information, note tags, external links, internal links, terminology concepts, knowledge tags, knowledge associations, programs, and experimental data according to the characteristics of the document content and the usage method.

[0014] Fragmented literature processing involves establishing a network of interconnected fragmented knowledge, which is beneficial for knowledge retrieval, reorganization, and precise services.

[0015] As a preferred embodiment of the document abstract content processing method based on XML fragmentation of the present invention, the dynamic association is to automatically annotate the fragmented content units according to the semantic engine and sort them according to the content weight of the documents to form a dynamic association network.

[0016] The weighting of literature content is based on the order of literature's technology, formula, direction, and results.

[0017] As a preferred embodiment of the document abstract content processing method based on XML fragmentation of the present invention, the data content model fragmentation is divided into two forms according to the granularity of the topic: basic metadata fragmentation and text fragmentation.

[0018] Basic metadata fragmentation includes bibliographic information and article information annotation;

[0019] Text fragmentation includes chapters, sections, paragraphs, subheadings, images, tables, footnotes, and formulas.

[0020] As a preferred embodiment of the document abstract content processing method based on XML fragmentation of the present invention, the fragmentation processing is automated.

[0021] Automated processing includes automatic XML annotation of the main text, automatic table of contents linking, garbled character detection and correction, automatic typesetting, automatic image processing, and automatic recognition.

[0022] As a preferred embodiment of the document abstract content processing method based on XML fragmentation of the present invention, the keyword and subject term content units are weighted content units that are ranked first in the document content weight sorting in the fragmentation process.

[0023] The weighting is based on the importance of specific data units in the overall selection judgment, and the calculation method is as follows:

[0024]

[0025] Where b refers to the number of times the data unit appears in the entire document, and B represents the number of data units in the document.

[0026] As a preferred embodiment of the document abstract content processing method based on XML fragmentation of the present invention, dynamic reorganization involves arranging keywords and subject terms according to the logical content library in order according to the document abstract requirements, and then achieving effective content output through the automatic typesetting, automatic recognition and correction functions of the automated processing flow.

[0027] The logical content library includes a raw material library, an article library, and an annotation library.

[0028] As a preferred embodiment of the document abstract content processing method based on XML fragmentation of the present invention, dynamic association first performs attribute tagging on all fragmented data units of the document, manages the association information of various types of content, and provides a full-text search engine to perform unified retrieval of all content;

[0029] Attribute tags are dynamically associated tags that are determined based on the characteristics of the data unit itself.

[0030] As a preferred embodiment of the document abstracting method based on XML fragmentation of the present invention, the document abstracting process is as follows:

[0031] First, perform a full-text attribute tag search to locate the relevant information;

[0032] Extract data units related to direction, technology, results, and experiments;

[0033] After extracting multiple directional data units, the weights of each directional data unit in the literature are compared, and the directional data unit with the highest weight is selected.

[0034] After extracting multiple technical data units, the frequency of use of each technical data unit in the literature is compared, and the technical data unit with the highest frequency of use is selected.

[0035] After extracting multiple result data units, perform correlation analysis on the correlation information of each result data unit in the literature, and select the result data unit with the highest correlation information.

[0036] The extracted data units are sorted in the order of directional data units, technical data units, and result data units;

[0037] New document abstracts are generated through automatic typesetting and automatic recognition.

[0038] The beneficial effects of this invention are that the accuracy of automatic metadata indexing is >95% and the accuracy of automatic text XML annotation is >90%, which can effectively improve the accuracy of target extraction while reducing the cost of text decomposition. Attached Figure Description

[0039] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:

[0040] Figure 1 This is a fragmentation style diagram of the XML-based document summary content processing method in Example 1.

[0041] Figure 2 This is a flowchart of the document abstract content processing method based on XML fragmentation in Example 1.

[0042] Figure 3 This is a fragmentation classification diagram of the XML-based document summary content processing method in Example 1.

[0043] Figure 4 This is a system architecture diagram of the XML-based document abstract content processing method in Example 2. Detailed Implementation

[0044] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0045] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0046] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0047] Example 1

[0048] Reference Figure 1 and Figure 2 This is the first embodiment of the present invention, which provides a method for processing document summary content based on XML fragmentation. The method for processing document summary content based on XML fragmentation includes...

[0049] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a method for processing document abstract content based on XML fragmentation, which includes converting the document into XML format;

[0050] The documents converted into XML format are divided into fragmented data units and composed of data content models according to four modes: article, chapter, section, and topic.

[0051] By dynamically associating data content models through keyword semantic relationships, keyword and topic content units are extracted from fragmented data application units;

[0052] The content units are dynamically reorganized according to the requirements of the document abstract.

[0053] The dynamic digital reorganization process is mainly divided into four stages: topic selection and planning, editing and processing, content management, and publishing services. While the division of stages shares some similarities with the traditional report editing and publishing process, the specific work content and characteristics within each stage differ significantly. The most important feature of the dynamic publishing process is leveraging the widespread availability and real-time nature of internet cloud services. Through data fragmentation indexing and indexing technologies, XML-based content format separation and reproduction technologies, and on-demand reorganization technologies, it achieves the structuring, fragmentation, scalability, automation, and diversity of digital publishing content. This provides readers with more convenient, faster, cheaper, and intelligent information access and knowledge services, while simultaneously realizing the commercialization of microdata from multiple data sources.

[0054] Fragmented data units are data application units that break down documents into formulas, charts, paragraphs, chapters, content outlines, interactive operations, interactive information, note tags, external links, internal links, terminology and concepts, knowledge tags, knowledge associations, programs, and experimental data based on the characteristics of the document content and the usage methods.

[0055] Fragmented literature processing involves establishing a network of interconnected fragmented knowledge, which is beneficial for knowledge retrieval, reorganization, and precise services.

[0056] Fragment indexing refers to the more detailed indexing and retrieval of knowledge from each chapter of a digital publishing resource, in addition to metadata annotation for the entire book or article. Indexed and retrievald fragments of knowledge are easier for readers to access and utilize, and their lifespan is longer and more effective than that of an entire book. The main processes of organizing fragmented digital content include:

[0057] (1) Maintain traditional publishing content, preserve author manuscripts, final manuscripts, and final layout documents, and convert the final layout documents to organize them according to the format of type, volume, document, article, chapter, and section;

[0058] (2) Classify the content of the chapters, sections and sections according to discipline, Chinese Library Classification, theme, etc., and construct a knowledge system according to a certain discipline, direction or industry.

[0059] (3) The knowledge system is further divided into knowledge units in different directions, knowledge units are further divided into knowledge points, and finally into keywords and thematic terms;

[0060] (4) Dynamically link knowledge points through semantic relationships between keywords to form a network of interconnected relationships;

[0061] (5) Reorganize the content as needed and use multi-format synchronous generation technology to achieve dynamic publishing.

[0062] Fragmentation processing uses automated methods;

[0063] Automated processing includes automatic XML annotation of the main text, automatic table of contents linking, garbled character detection and correction, automatic typesetting, automatic image processing, and automatic recognition.

[0064] In many domains, document structures are diverse, containing various textual information (reference fields), such as titles, keywords, and links. Various studies have shown that these reference fields resemble actual user queries, providing a concise summary of what the document is about and what search intents it might fulfill. These brief, highly representative fields offer evidence of which terms are of high importance within the document.

[0065] The weighting is based on the importance of specific data units in the overall selection judgment, and the calculation method is as follows:

[0066]

[0067] Where 'b' refers to the number of times a data unit appears in the overall literature, and 'B' represents the number of data units in the literature. The weight of a certain indicator refers to its relative importance in the overall evaluation. Weight interpretation in the evaluation process: Weight represents the quantitative allocation of the importance of different aspects of the evaluated object during the evaluation process, differentiating the role of each evaluation factor in the overall evaluation.

[0068] When retrieving documents, the content can sometimes be very long, ranging from thousands to tens of thousands of words. If you want to use the BERT model to predict the weight of each word, you need to overcome the 512-token length limit. A simple and straightforward approach is to cut the document into paragraphs, keeping the paragraph length within 512 tokens, and then merge the paragraphs back together to form the original document after prediction.

[0069] The process of creating a literature abstract is as follows:

[0070] First, perform a full-text attribute tag search to locate the relevant information;

[0071] Extract data units related to direction, technology, results, and experiments;

[0072] After extracting multiple directional data units, a weight comparison is performed, and the directional data unit with the higher weight is extracted.

[0073] After extracting multiple technical data units, compare the frequency of use of each technical data unit in the literature, and select the technical data units with higher frequency of use in the literature for extraction;

[0074] After extracting multiple result data units, the repetition rate of the result data unit is compared with that of the experimental data unit, and the result data unit with the higher repetition rate with the experimental data unit is selected.

[0075] The extracted data units are sorted in the order of directional data units, technical data units, and result data units;

[0076] New document abstracts are generated through automatic typesetting and automatic document recognition. The process is as follows: Figure 2 As shown.

[0077] Data content model fragmentation is divided into two forms according to the granularity of the topic: basic metadata fragmentation and text fragmentation.

[0078] Basic metadata fragmentation includes bibliographic information and article information annotation;

[0079] The main text is fragmented, including chapters, sections, paragraphs, subheadings, images, tables, footnotes, and formulas. Its structure is as follows: Figure 1 As shown.

[0080] Example 2

[0081] Reference Figure 4 This is the second embodiment of the present invention, which differs from the first embodiment in that it further includes: In the previous embodiment, the document summary content processing method based on XML fragmentation included:

[0082] China Southern Power Grid Energy Development Research Institute Co., Ltd. (hereinafter referred to as "Southern Power Grid Energy Institute") is a wholly-owned subsidiary of China Southern Power Grid Company (ranked 100th in the Fortune Global 500 in 2017). This project uses an XML-based fragmented document abstract content processing method and a security architecture that adheres to the principle of "synchronous planning, synchronous construction, and synchronous commissioning" to strengthen the network security protection of this system. Ensuring that no major or above information security incidents occur is the bottom line for network security protection in the construction and operation of this system.

[0083] In accordance with the "Basic Requirements for Cybersecurity Level Protection of Information Security Technology" (GB / T22239-2019) and the extended requirements for cloud computing, mobile internet, Internet of Things and industrial control systems, this system is classified, protected, evaluated and registered to ensure the security of critical information infrastructure.

[0084] In the literature platform of China Southern Power Grid Energy Research Institute, the data of journals and knowledge bases are preprocessed into XML data format. An XML file contains the metadata, specific content and images of the XML file.

[0085] The processing flow for XML data differs slightly from that for other data formats. Since the metadata of the content is contained in the XML file, there is no need to manually fill out the MODS form. Instead, the XML data is converted into a MODS data stream using an XSL template file during system ingestion, and then further converted to generate a DC data stream and stored.

[0086] In Embodiment 2 of the present invention, the literature "Establishment of Atmosphere-Ocean Wave Coupling Model in Finite Region and Experiment on Capital Roughness Parameterization" from the literature platform of China Southern Power Grid Energy Research Institute is selected for operation.

[0087] First, extract data units related to direction, technology, results, and experiment. After performing a full-text search of the literature, the corresponding data units related to direction, technology, results, and experiment can be extracted. After selecting multiple data units, a weight comparison is performed, and the data units with higher weights are extracted. At this point, the extracted keyword direction data is sea surface roughness.

[0088] After extracting multiple technical data units, their frequency of use in the literature is compared, and the technical data unit with the higher frequency of use in the literature is selected for extraction. The extracted technical data at this point is:

[0089] The coupling module consists of five independent C language functions. Because the Linux system supports mixed-language programming and linking, the coupling module can be easily called by the atmosphere and wave mode components. The structure and function of each function in the coupling module are shown in Table 1. Among them, `forkprog` is the most important function in the coupling module, serving as the primary process for introducing pipe communication functionality between the two mode components. The coupling mode uses the atmosphere mode component as the parent process. By calling the `forkprog` function, it first generates two pairs of simplex pipes, then creates a child process and executes the wave mode component. Finally, by controlling the pipe ports, a bidirectional communication pipe is formed, thus establishing a coupling mode system consisting of the atmosphere mode component, the wave mode component, and the bidirectional pipe connecting the two mode components.

[0090] After extracting multiple result data units, the repetition rate of these result data units is compared with that of the experimental data units. The result data unit with the higher repetition rate with the experimental data unit is selected. At this point, the extraction result is:

[0091] (1) Pipeline communication technology under the LINUX system can be conveniently and quickly used to establish air-sea coupling modes under simple physical quantity exchange. This method has low requirements for machine system and operating environment, and the established coupling mode runs stably and can be used for simulation research on air-sea interaction.

[0092] (2) Coupled models can better simulate the development and evolution of tropical cyclones. Compared with uncoupled atmospheric models, coupled models have little impact on the movement path of cyclone systems, but a significant impact on system intensity. Coupled models are more sensitive to different sea surface roughness parameterization schemes: the Smith92 scheme strengthens the cyclone system, while the coupled model weakens the cyclone system simulated under the SCOR01 and Makin05 schemes.

[0093] (3) The coupled model simulated the distribution of sea surface significant wave height well during the movement and development of tropical cyclones. Along the cyclone's movement path, there is a large area of ​​significant wave height distributed to the right rear of the cyclone center, and its intensity varies with the intensity of the cyclone system. The sea surface significant wave height simulated by the model has a good correlation with the satellite altimeter wave height, especially in deep-sea areas far from land and islands. The coupled model using the Makin05 scheme improved the simulation of sea surface significant wave height.

[0094] The extracted data units are sorted in the order of directional data units, technical data units, and result data units.

[0095] The final results of the XML-based document abstract processing method are as follows: Figure 4As shown, the accuracy rate of automatic metadata indexing is >95% (project text information: including title and responsibility statement, project name, applicant unit, project time, project leader and other more than 20 basic project information items), and the accuracy rate of automatic text XML annotation is >90%.

[0096] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for processing document abstract content based on XML fragmentation, characterized in that... include: Convert the document into XML format; The documents converted into XML format are divided into fragmented data units and composed of data content models according to four modes: article, chapter, section, and topic. The fragmented data unit refers to the process of breaking down a document into data application units such as formulas, charts, paragraphs, chapters, content outlines, interactive operations, interactive information, note tags, external links, internal links, terminology and concepts, knowledge tags, knowledge associations, programs, and experimental data, based on the characteristics of the document content and the usage methods. Fragmented document processing involves establishing a network of interconnected fragmented knowledge, which is beneficial for knowledge retrieval, reorganization, and precise services. By dynamically associating data content models through keyword semantic relationships, keyword and topic content units are extracted from fragmented data application units; The dynamic association involves automatically labeling fragmented content units using a semantic engine and sorting them according to the content weight of the documents to form a dynamic association network. The document content weight ranking is based on the order of document technology, formula, direction, and result. The dynamic association is implemented by first assigning attribute tags to all fragmented data units of the document, managing the association information of various types of content, and providing a full-text search engine to perform unified retrieval of all content; The attribute tags are dynamically associated and determined based on the characteristics of the data unit itself. The content units are dynamically reorganized according to the requirements of the document abstract; The process for forming the document abstract is as follows: First, perform a full-text attribute tag search to locate the relevant information; Extract data units related to direction, technology, results, and experiments; After extracting multiple directional data units, the weights of each directional data unit in the literature are compared, and the directional data unit with the highest weight is selected. After extracting multiple technical data units, the frequency of use of each technical data unit in the literature is compared, and the technical data unit with the highest frequency of use is selected. After extracting multiple result data units, perform correlation analysis on the correlation information of each result data unit in the literature, and select the result data unit with the highest correlation information. The extracted data units are sorted in the order of directional data units, technical data units, and result data units; New document abstracts are generated through automatic typesetting and automatic recognition.

2. The document abstract processing method based on XML fragmentation as described in claim 1, characterized in that: The XML format is Extensible Markup Language, which is a data storage language; A knowledge-enhanced representation model is established using the XML format to fragment knowledge resources at both the formal structure and content levels. Based on the characteristics and utilization methods of various knowledge resources, content fragmentation is gradually realized.

3. The document abstract processing method based on XML fragmentation as described in claim 1, characterized in that: The data content model fragmentation is divided into two forms according to the granularity of the topic: basic metadata fragmentation and text fragmentation. The basic metadata fragmentation includes bibliographic information and article information annotation; The text fragmentation includes chapters, sections, paragraphs, subheadings, images, tables, footnotes, and formulas.

4. The document abstract processing method based on XML fragmentation as described in claim 3, characterized in that: The basic metadata fragmentation and text fragmentation are performed using automated processing; The automated processing includes automatic XML annotation of the main text, automatic table of contents linking, garbled character detection and correction, automatic typesetting, automatic image processing, and automatic recognition.

5. The document summary content processing method based on XML fragmentation as described in claim 1, characterized in that: The keyword and subject term content units are weighted content units that are ranked first in the document content weighting process during the fragmentation process. The weighting is based on the importance of specific data units in the overall selection judgment, and the calculation method is as follows: Where b refers to the number of times the data unit appears in the entire document, and B represents the number of data units in the document.

6. The document abstract processing method based on XML fragmentation as described in claim 5, characterized in that: The dynamic reorganization involves arranging keywords and subject terms according to the requirements of the document abstract in a logical content library. After the arrangement, the automatic typesetting, automatic recognition and correction functions of the automated processing flow are used to achieve effective content output. The logical content library includes a raw material library, an article library, and an annotation library.

Citation Information

Patent Citations

  • Intelligent knowledge retrieval system

    CN114610847A