Document directory automatic generation method, device and computer readable storage medium

By generating adversarial network model training and generating regular expressions, the problem of difficult to identify and extract document directory structure in the prior art is solved, and the accurate and efficient generation of document directory is achieved.

CN110852079BActive Publication Date: 2025-05-02PING AN TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201910965809.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-10-11
Publication Date
2025-05-02
Estimated Expiration
2039-10-11

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify and extract the directory structure of documents, especially in the case of multi-level titles in the documents.

Method used

By extracting the initial title in the target document, determining the initial title rules, and inputting them to generate an adversarial network model training, generating regular expressions to compare and analyze the document content, extracting and arranging the titles, and generating a document directory.

Benefits of technology

It realizes accurate and efficient generation of document directories, can identify and extract the title structure in the document, and form a complete and accurate directory.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN110852079B_ABST
    Figure CN110852079B_ABST
Patent Text Reader

Abstract

The present invention relates to an artificial intelligence technology, and discloses a method for automatically generating a document directory, including: extracting the initial title in the target document, and determining the initial title rule of the target document based on the initial title; inputting the initial title rule into a pre-constructed generative adversarial network model for training to obtain the trained title rule; generating a regular expression based on the trained title rule; traversing the entire content of the target document, comparing and analyzing the content in the target document with the regular expression, extracting all the titles of the target document, arranging all the titles in the order of traversal, and generating a document directory. The present invention also proposes a device for automatically generating a document directory and a computer-readable storage medium. The present invention can realize an accurate and efficient automatic generation function of a document directory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, device and computer-readable storage medium for deep learning of document structure and generating a document directory. Background Art

[0002] The existing methods for extracting document directories are mainly to read a word document through POI (Point of Interest). The existing technology can only read by paragraphs and cannot identify the specific structure of the document. In addition, when there are multiple levels of titles in the document, the existing methods cannot completely and accurately extract the directory structure in the document. Summary of the invention

[0003] The present invention provides a method, device and computer-readable storage medium for automatically generating a document directory, the main purpose of which is to provide a method for performing deep learning on a target document to obtain a document directory.

[0004] To achieve the above object, the present invention provides a method for automatically generating a document directory, comprising:

[0005] Extracting an initial title from a target document, and determining an initial title rule of the target document based on the initial title;

[0006] Inputting the initial title rule into a pre-built generative adversarial network model for training to obtain a trained title rule;

[0007] Generate a regular expression based on the trained title rule;

[0008] The entire content of the target document is traversed, the content in the target document is compared and analyzed with the regular expression, all the titles of the target document are extracted, all the titles are arranged in the order of traversal, and a document directory is generated.

[0009] Optionally, the document directory automatic generation method further includes: constructing the generative adversarial network model, including:

[0010] A generative model and a discriminative model are established; the generative model and the discriminative model are subjected to mutual game learning to obtain an optimal solution, wherein the optimal solution includes the trained title rules.

[0011] Optionally, before generating the regular expression, the document directory automatic generation method further includes:

[0012] A state machine is generated based on the trained title rules; wherein the generated state machine includes:

[0013] Performing grammatical analysis on the trained title rules, and rewriting the trained title rules into state machine rules required for state machine construction; and constructing the state machine according to the state machine rules;

[0014] Convert the constructed state machine into the format required to generate regular expressions and store it.

[0015] Optionally, traversing all the contents of the target document, comparing and analyzing the contents of the target document with the regular expression, and extracting all the titles of the target document includes:

[0016] Traversing the entire content of the target document, and extracting one or more points of interest from the target document;

[0017] Extracting the content of the target document through the points of interest, and identifying the outline structure of the target document;

[0018] The outline structure of the target document is compared and matched with the regular expression. If the content in the target document matches the regular expression, the content in the target document is confirmed to be the title and the title is extracted. If the content in the target document does not match the regular expression, the content in the target document is confirmed to be text.

[0019] Optionally, the document directory is in Extensible Markup Language; and the file format of the target document is Microsoft Office Word.

[0020] In addition, to achieve the above-mentioned purpose, the present invention further provides a document directory automatic generation device, the device comprising a memory and a processor, the memory storing a document directory automatic generation program that can be run on the processor, and the document directory automatic generation program when executed by the processor implements the following steps:

[0021] Extracting an initial title from a target document, and determining an initial title rule of the target document based on the initial title;

[0022] Inputting the initial title rule into a pre-built generative adversarial network model for training to obtain a trained title rule;

[0023] Generate a regular expression based on the trained title rule;

[0024] The entire content of the target document is traversed, the content in the target document is compared and analyzed with the regular expression, all the titles of the target document are extracted, all the titles are arranged in the order of traversal, and a document directory is generated.

[0025] Optionally, the document directory automatic generation method also includes: constructing the generative adversarial network model, including: obtaining an optimal solution through mutual game learning between the generative model and the discriminative model, wherein the optimal solution includes the trained title rules.

[0026] Optionally, before generating the regular expression, the document directory automatic generation method further includes:

[0027] A state machine is generated based on the trained title rules, wherein the generated state machine includes:

[0028] Performing grammatical analysis on the trained title rules, and rewriting the trained title rules into state machine rules required for state machine construction;

[0029] The state machine is constructed according to the state machine rules; the constructed state machine is converted into a format required for generating a regular expression and stored.

[0030] Optionally, traversing all the contents of the target document, comparing and analyzing the contents of the target document with the regular expression, and extracting all the titles of the target document includes:

[0031] Traversing the entire content of the target document, and extracting one or more points of interest from the target document;

[0032] Extracting the content of the target document through the points of interest, and identifying the outline structure of the target document;

[0033] The outline structure of the target document is compared and matched with the regular expression. If the content in the target document matches the regular expression, the content in the target document is confirmed to be the title and the title is extracted. If the content in the target document does not match the regular expression, the content in the target document is confirmed to be text.

[0034] Optionally, the document directory is in Extensible Markup Language; and the file format of the target document is Microsoft Office Word.

[0035] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which a document directory automatic generation program is stored. The document directory automatic generation program can be executed by one or more processors to implement the steps of the document directory automatic generation method as described above.

[0036] The present invention extracts the initial title rules from the target document, and by inputting the initial title rules into a pre-built generative adversarial network model for training and obtaining the trained title rules, the computer can efficiently analyze the title without losing accuracy. Further, a regular expression is configured according to the trained title rules, and finally the content in the target document is compared and analyzed with the regular expression to extract the title. Therefore, the document directory automatic generation method, device and computer-readable storage medium proposed in the present invention can realize accurate, efficient and coherent document directory generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 A schematic diagram of a flow chart of a method for automatically generating a document directory according to an embodiment of the present invention;

[0038] Figure 2 A schematic diagram of the internal structure of a device for automatically generating a document directory according to an embodiment of the present invention;

[0039] Figure 3 A schematic diagram of modules of a document directory automatic generation program in a document directory automatic generation device provided in an embodiment of the present invention.

[0040] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0041] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0042] The present invention provides a method for automatically generating a document directory. Figure 1 FIG. 1 is a flow chart of a method for automatically generating a document directory according to an embodiment of the present invention. The method for automatically generating a document directory may be executed by a device, which may be implemented by software and / or hardware.

[0043] In this embodiment, the document directory automatic generation method includes:

[0044] S1. Extracting an initial title from a target document, and determining an initial title rule of the target document based on the initial title.

[0045] The target document is a document object for which a document directory needs to be generated in the present invention, wherein the target document is in word format. For example, the target document can be a word text of the novel "Border Town"; a word text of "How to Read a Book" and other different types of text documents. The present invention aims to identify the text content of the target document, extract the content with chapter characteristics, sort the content with chapter characteristics according to preset rules, and form a document directory of the target document.

[0046] The present invention first extracts the initial title in the target document. The initial title refers to a short sentence in the target document that indicates the content of the article, work, etc., which is generally divided into a general title, a subtitle, a subtitle, etc. The title can enable readers to understand the main content and purpose of the article. The title can be used to form a natural division of chapters, paragraphs, etc.

[0047] Furthermore, the present invention determines the initial title rule of the target document based on the initial title, including: after extracting the initial title in the target document, based on the specific form of the initial title (i.e., the grammar, semantic logic, and hierarchical relationship of each title actually contained in the initial title), abstracting the general rules in the grammar, semantic logic, and hierarchical relationship of each title of the initial title to determine the initial title rule of the target document. Among them, the initial title rule refers to the type, structure, semantic logic, and hierarchical relationship of each title of the initial title.

[0048] Specifically, the grammar refers to the word class to which the specific words in the initial title belong and the composition and word form changes (morphology) of the words in this class. For example, the text document "Animal Encyclopedia" contains the initial titles: Mammals, Birds, Reptiles, etc. The grammar of these titles is nouns; the semantic logic uses modern logic methods to reveal the relationship between language expressions and their meanings. For example, the text document "Animal Encyclopedia" contains the initial titles: Mammals, Felines, Leopards, Cats. According to the semantic logic, the titles "cats" and "leopards" both belong to the mammal class, and are in a containment relationship in semantic logic; the hierarchical relationship is from large to small: Chapter x, Section 1, Section 1.1.1, (1), etc. According to the hierarchical relationship of all the initial titles in the text; the semantic logical connection between the title content and other text content, the corresponding title rules can be determined. Taking the first chapter of the text "Research on the Non-material Incentive Strategy for Knowledge Employees of M Company" as an example, through the preset hierarchical relationship logic, such as chapter>section>subsection. The highest-level title is: Chapter 1 Introduction; the second-level titles are: Section 1 Background of the Topic Selection Section 2 Research Significance Section 3 Research Content Section 4 Research Methods; the third-level titles are: Subsection 1.2.1 Theoretical Significance Subsection 1.2.2 Practical Significance

[0049] S2. Input the initial title rules into a pre-built generative adversarial network model for training to obtain trained title rules.

[0050] Preferably, the generative adversarial network model includes a generative model and a discriminative model. The generative model and the discriminative model obtain an optimal solution through mutual game learning, wherein the optimal solution includes the trained title rule.

[0051] The present invention attempts to obtain virtual title rule data (sample G(z)) by inputting the initial title rule into a generated network model for training, and then using a discriminant model D to judge the generated virtual title rule data to train it to conform to the characteristics of the initial title rule, so as to obtain the optimal solution after training, wherein the optimal solution includes the trained title rule.

[0052] For example, the title rules trained and obtained in the constructed adversarial network model include:

[0053] First, the initial title rule is input as a variable z (hereinafter referred to as variable z) into a pre-built generative adversarial network model;

[0054] After obtaining the input of the initial title rule, the generative model G generates a sample G(z) that obeys the real data distribution;

[0055] Then, the text content of the target document and the sample G(z) are taken as input data sets, wherein the input data set may contain one or all of the text content of the target document and the sample G(z).

[0056] The input data set is input into the discriminant model D, wherein the function of the discriminant model D is to discriminate whether the input data is from the generative model G or from real data, i.e., the text content of the target document (the real data here refers to the real and specific content of the text of the target document, rather than the virtual data sample G(z) generated after learning); if the data in the current input data set comes from the sample G(z), the input data set is marked as 0 and judged as false, otherwise, if the data in the current input data set is not from G(z), the data in the current input data set comes from real data, the input data set is marked as 1 and judged as true. The goal of the generative model G here is to make the performance of the virtual data sample G(z) generated by it on the discriminant model D consistent with the performance of the real data (the text content of the target document) on D.

[0057] The mutual game learning includes: the process of mutual game learning and iterative optimization between the generation model G and the discriminant model D makes the performance of the generation model G and the discriminant model D continuously improve. As the discriminant model D's discriminative ability improves and it is unable to discriminate the source of the data input into the discriminant model D, it is considered that the generation model G has learned the true data distribution.

[0058] The initial title rule is input into the generative network model for training to obtain a sample G(z), and then the discriminant model D judges the generated sample G(z) to train it to meet the characteristics of the initial title rule. The initial title rule is input to the generative model G and the discriminant model D for mutual game learning to obtain an optimal solution, wherein the optimal solution includes the trained title rule.

[0059] S3. Generate a regular expression based on the trained title rules.

[0060] Generating a regular expression according to the method includes:

[0061] Obtaining the trained title rule;

[0062] Performing grammatical analysis on the trained title rules to extract the sentence body of the trained title rules;

[0063] Obtaining the semantic slots of the words in the sentence body;

[0064] A regular expression is generated according to the sentence body, the semantic slot and the remaining non-body part in the trained title rule.

[0065] In some embodiments of the present invention, before configuring the regular expression, the method further includes:

[0066] Generate a state machine based on the trained title rules.

[0067] Preferably, the state machine is generated to provide a stable configuration device and storage device in the process of generating the regular expression. In this embodiment, the state machine is a device that specifically configures and stores regular expressions according to corresponding title rules. The state machine matches the regular expression in the state machine according to the characters and position information in the received regular expression.

[0068] Wherein, generating the state machine comprises the following steps:

[0069] S301. Perform syntax analysis on the trained title rules to obtain a configuration file, which describes the identifier of each state of the trained title rules, response information to each event, and state transition information, and describes the hierarchical relationship between multiple states.

[0070] S302: construct the state machine according to the configuration file.

[0071] S303: Convert the constructed state machine into a format required for generating the regular expression and store it.

[0072] In some embodiments of the present invention, the state machine is composed of a state register and a combinational logic circuit, and can perform state transitions according to a preset state based on a control signal. It is a control center that coordinates related signal actions and completes specific operations.

[0073] Preferably, according to different actual application carriers, the state machine can be represented by data table entries, linked lists, instruction table entries, state diagrams, etc., which is not limited in this embodiment.

[0074] S4, traversing the entire content of the target document, comparing and analyzing the content of the target document with the regular expression, extracting all the titles of the target document, arranging all the titles in the order of traversal, and generating a document directory. According to one embodiment of the present invention, reading the target document, comparing and analyzing the content of the target document with the regular expression, and extracting the title of the target document comprises the following steps:

[0075] S401, extracting one or more points of interest (POI) from the target document.

[0076] The method comprises: obtaining data information of a POI object based on the regular expression, wherein the data information of the POI object at least includes a sentence body; comparing the obtained data information of the POI object with the content in the target document one by one, and extracting the part of the target document that has the same rules as the data information of the POI object as the POI. For example, the obtained data information of the POI object is X, Y; after comparing the content in the target document, it is found that the POI has the same rules as the content in the target document (such as logic rules, grammar rules, etc.), and then extracting the part as the POI.

[0077] Among them, the point of interest is the open source library of the Apache Software Foundation, POI provides an API for Java programs to read and write Microsoft Office format files.

[0078] In a preferred embodiment of the present invention, the POI technology is used to read the content of the Word text paragraph of the target document. Among them, a Word document contains multiple paragraphs, a paragraph contains multiple Runs, and a Run contains multiple Runs. Run is the smallest unit of the target document. For example, a paragraph contains several complete sentences, namely the Runs; and the Runs contain several phrases, namely the Runs. Specifically, the steps of reading the Word text content through POI are as follows:

[0079] (1) First, use POI to operate XWPFParagraph in XWPFDocument to obtain all paragraphs of the target document;

[0080] (2) Get all the Runs in a paragraph through the xwpfParagraph.getRuns() command:

[0081] (3) Get a Run in a Runs using the xwpfRuns.get(index) command;

[0082] (4) Traverse the entire document and obtain the title in the word document through the getPPr().getOutlineLvl() command.

[0083] Based on the above process, the word document content is traversed and all the titles in the word document are extracted.

[0084] S402, extracting the content of the target document through the point of interest (POI) and identifying the outline structure of the target document. The outline structure of the target document is to arrange all the proposed titles in a hierarchical and sequential manner according to the semantic logic and sequence to form an outline structure.

[0085] S403, comparing and matching the outline structure of the target document with the regular expression. If the content in the target document matches the regular expression, it is confirmed that the content in the target document is the title. If the content in the target document does not match the regular expression, it is confirmed that the content in the target document is text.

[0086] S5. Traverse all the contents of the target document, extract all the titles, arrange all the titles in the order of traversal, and generate a document directory.

[0087] In a preferred embodiment of the present invention, all paragraph contents of a word document are traversed, and the document content titles are identified based on regular expression comparison. The document titles are traversed and refined in the order of the document contents and integrated into a new document, which is the extracted complete word document chapter directory.

[0088] The invention also provides a device for automatically generating a document directory. Figure 2 FIG. 1 is a schematic diagram of the internal structure of a device for automatically generating a document directory according to an embodiment of the present invention.

[0089] In this embodiment, the document directory automatic generation device 1 can be a PC (Personal Computer), or a terminal device such as a smart phone, a tablet computer, a portable computer, or a server, etc. The document directory automatic generation device 1 at least includes a memory 11, a processor 12, a communication bus 13, and a network interface 14.

[0090] The memory 11 includes at least one type of readable storage medium, and the readable storage medium includes a flash memory, a hard disk, a multimedia card, a card-type memory (for example, an SD or DX memory, etc.), a magnetic memory, a disk, an optical disk, etc. In some embodiments, the memory 11 may be an internal storage unit of the document directory automatic generation device 1, such as a hard disk of the document directory automatic generation device 1. In other embodiments, the memory 11 may also be an external storage device of the document directory automatic generation device 1, such as a plug-in hard disk, a smart memory card (Smart MediaCard, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the document directory automatic generation device 1. Further, the memory 11 may also include both an internal storage unit and an external storage device of the document directory automatic generation device 1. The memory 11 may not only be used to store application software and various types of data installed in the document directory automatic generation device 1, such as the code of the document directory automatic generation program 01, but also be used to temporarily store data that has been output or is to be output.

[0091] In some embodiments, the processor 12 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor or other data processing chip, used to run program codes or process data stored in the memory 11, such as executing the document directory automatic generation program 01.

[0092] The communication bus 13 is used to realize the connection and communication between these components.

[0093] The network interface 14 may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), and is generally used to establish a communication connection between the device 1 and other electronic devices.

[0094] Optionally, the device 1 may also include a user interface, which may include a display (Display), an input unit such as a keyboard (Keyboard), and the optional user interface may also include a standard wired interface and a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, and an OLED (Organic Light-Emitting Diode) touch device, etc. The display may also be appropriately referred to as a display screen or a display unit, which is used to display information processed in the document directory automatic generation device 1 and to display a visual user interface.

[0095] Figure 2 Only the document directory automatic generation device 1 having components 11-14 and the document directory automatic generation program 01 is shown. It can be understood by those skilled in the art that Figure 1 The structure shown does not constitute a limitation on the document catalog automatic generation device 1, and may include fewer or more components than shown in the figure, or a combination of certain components, or a different arrangement of components.

[0096] exist Figure 2 In the embodiment of the device 1 shown, the memory 11 stores a document directory automatic generation program 01; when the processor 12 executes the document directory automatic generation program 01 stored in the memory 11, the following steps are implemented:

[0097] Step 1: extract the initial title in the target document, and determine the initial title rule of the target document based on the initial title.

[0098] The target document is a document object for which a document directory needs to be generated in the present invention, wherein the target document is in word format. For example, the target document can be a word text of the novel "Border Town"; a word text of "How to Read a Book" and other different types of text documents. The present invention aims to identify the text content of the target document, extract the content with chapter characteristics, sort the content with chapter characteristics according to preset rules, and form a document directory of the target document.

[0099] The present invention first extracts the initial title in the target document. The initial title refers to a short sentence in the target document that indicates the content of the article, work, etc., which is generally divided into a general title, a subtitle, a subtitle, etc. The title can enable readers to understand the main content and purpose of the article. The title can be used to form a natural division of chapters, paragraphs, etc.

[0100] Furthermore, the present invention determines the initial title rule of the target document based on the initial title, including: after extracting the initial title in the target document, based on the specific form of the initial title (i.e., the grammar, semantic logic, and hierarchical relationship of each title actually contained in the initial title), abstracting the general rules in the grammar, semantic logic, and hierarchical relationship of each title of the initial title to determine the initial title rule of the target document. Among them, the initial title rule refers to the type, structure, semantic logic, and hierarchical relationship of each title of the initial title. The grammar refers to the word class to which the specific words in the initial title belong and the composition and word form changes (morphology) of the words in this class. For example, the text document "Animal Encyclopedia" contains the initial titles: Mammals, Birds, Reptiles, etc. The grammar of these titles is nouns; the semantic logic uses modern logic methods to reveal the relationship between language expressions and their meanings. For example, the text document "Animal Encyclopedia" contains the initial titles: Mammals, Felines, Leopards, Cats. According to semantic logic, the titles "cat" and "leopard" both belong to the mammal class, and are in a containment relationship in semantic logic; the hierarchical relationship is from large to small: Chapter x, Section 1, 1.1.1, (1), etc. According to the hierarchical relationship of all the initial titles in the text and the semantic logical connection between the title content and other text content, the corresponding title rules can be determined. Taking the first chapter of the text "Research on Non-material Incentive Strategies for Knowledge Workers of M Company" as an example, the highest-level title is: Chapter 1 Introduction; the second-level titles are: Section 1 Background of the Topic Selection Section 2 Research Significance Section 3 Research Content Section 4 Research Methods; the third-level titles are: 1.2.1 Theoretical Significance 1.2.2 Practical Significance.

[0101] Step 2: Input the initial title rules into a pre-built generative adversarial network model for training to obtain trained title rules.

[0102] Preferably, the generative adversarial network model includes a generative model and a discriminative model. The generative model and the discriminative model obtain an optimal solution through mutual game learning, wherein the optimal solution includes the trained title rule.

[0103] The present invention attempts to obtain virtual title rule data (sample G(z)) by inputting the initial title rule into a generated network model for training, and then using a discriminant model D to judge the generated virtual title rule data to train it to conform to the characteristics of the initial title rule, so as to obtain the optimal solution after training, wherein the optimal solution includes the trained title rule.

[0104] For example, the title rules trained and obtained in the constructed adversarial network model include:

[0105] First, the initial title rule is input as a variable z (hereinafter referred to as variable z) into a pre-built generative adversarial network model;

[0106] After obtaining the input of the initial title rule, the generative model G generates a sample G(z) that obeys the real data distribution;

[0107] Then, the text content of the target document and the sample G(z) are taken as input data sets, wherein the input data set may contain one or all of the text content of the target document and the sample G(z).

[0108] The input data set is input into the discriminant model D, wherein the function of the discriminant model D is to discriminate whether the input data is from the generative model G or from real data, i.e., the text content of the target document (the real data here refers to the real and specific content of the text of the target document, rather than the virtual data sample G(z) generated after learning); if the data in the current input data set comes from the sample G(z), the input data set is marked as 0 and judged as false, otherwise, if the data in the current input data set is not from G(z), the data in the current input data set comes from real data, the input data set is marked as 1 and judged as true. The goal of the generative model G here is to make the performance of the virtual data sample G(z) generated by it on the discriminant model D consistent with the performance of the real data (the text content of the target document) on D.

[0109] The mutual game learning includes: the process of mutual game learning and iterative optimization between the generation model G and the discriminant model D makes the performance of the generation model G and the discriminant model D continuously improve. As the discriminant model D's discriminative ability improves and it is unable to discriminate the source of the data input into the discriminant model D, it is considered that the generation model G has learned the true data distribution.

[0110] The initial title rule is input into the generative network model for training to obtain a sample G(z), and then the discriminant model D judges the generated sample G(z) to train it to meet the characteristics of the initial title rule. The initial title rule is input to the generative model G and the discriminant model D for mutual game learning to obtain an optimal solution, wherein the optimal solution includes the trained title rule.

[0111] Step 3: Generate a regular expression based on the trained title rules.

[0112] Generating a regular expression according to the method includes:

[0113] Obtaining the trained title rule;

[0114] Performing grammatical analysis on the trained title rules to extract the sentence body of the trained title rules;

[0115] Obtaining the semantic slots of the words in the sentence body;

[0116] A regular expression is generated according to the sentence body, the semantic slot and the remaining non-body part in the trained title rule.

[0117] In some embodiments of the present invention, before configuring the regular expression, the method further includes:

[0118] Generate a state machine based on the trained title rules.

[0119] Preferably, the state machine is generated to provide a stable configuration device and storage device in the process of generating the regular expression. In this embodiment, the state machine is a device that specifically configures and stores regular expressions according to corresponding title rules. The state machine matches the regular expression in the state machine according to the characters and position information in the received regular expression.

[0120] Wherein, generating the state machine comprises the following steps:

[0121] S301. Perform syntax analysis on the trained title rules to obtain a configuration file, which describes the identifier of each state of the trained title rules, response information to each event, and state transition information, and describes the hierarchical relationship between multiple states.

[0122] S302: construct the state machine according to the configuration file.

[0123] S303: Convert the constructed state machine into a format required for generating the regular expression and store it.

[0124] In some embodiments of the present invention, the state machine is composed of a state register and a combinational logic circuit, and can perform state transitions according to a preset state based on a control signal. It is a control center that coordinates related signal actions and completes specific operations.

[0125] Preferably, according to different actual application carriers, the state machine can be represented by data table entries, linked lists, instruction table entries, state diagrams, etc., which is not limited in this embodiment.

[0126] Step 4: traverse the entire content of the target document, compare and analyze the content in the target document with the regular expression, extract all the titles of the target document, arrange all the titles in the order of traversal, and generate a document directory. According to one embodiment of the present invention, reading the target document, comparing and analyzing the content in the target document with the regular expression, and extracting the title of the target document includes the following steps:

[0127] S401, extracting one or more points of interest (POI) from the target document.

[0128] The method comprises: obtaining data information of a POI object based on the regular expression, wherein the data information of the POI object at least includes a sentence body; comparing the obtained data information of the POI object with the content in the target document one by one, and extracting the part of the target document that has the same rules as the data information of the POI object as the POI. For example, the obtained data information of the POI object is X, Y; after comparing the content in the target document, it is found that the POI has the same rules as the content in the target document (such as logic rules, grammar rules, etc.), and then extracting the part as the POI.

[0129] Among them, the point of interest is the open source library of the Apache Software Foundation, POI provides an API for Java programs to read and write Microsoft Office format files.

[0130] In a preferred embodiment of the present invention, the POI technology is used to read the content of the Word text paragraph of the target document. Among them, a Word document contains multiple paragraphs, a paragraph contains multiple Runs, and a Run contains multiple Runs. Run is the smallest unit of the target document. For example, a paragraph contains several complete sentences, namely the Runs; and the Runs contain several phrases, namely the Runs. Specifically, the steps of reading the Word text content through POI are as follows:

[0131] (1) First, use POI to operate XWPFParagraph in XWPFDocument to obtain all paragraphs of the target document;

[0132] (2) Get all the Runs in a paragraph through the xwpfParagraph.getRuns() command:

[0133] (3) Get a Run in a Runs using the xwpfRuns.get(index) command;

[0134] (4) Traverse the entire document and obtain the title in the word document through the getPPr().getOutlineLvl() command.

[0135] Based on the above process, the word document content is traversed and all the titles in the word document are extracted.

[0136] S402, extracting the content of the target document through the point of interest (POI) and identifying the outline structure of the target document. The outline structure of the target document is to arrange all the proposed titles in a hierarchical and sequential manner according to the semantic logic and sequence to form an outline structure.

[0137] S403, comparing and matching the outline structure of the target document with the regular expression. If the content in the target document matches the regular expression, it is confirmed that the content in the target document is the title. If the content in the target document does not match the regular expression, it is confirmed that the content in the target document is text.

[0138] Step 5: traverse the entire content of the target document, extract all the titles, arrange all the titles in the order of traversal, and generate a document directory.

[0139] In a preferred embodiment of the present invention, all paragraph contents of a word document are traversed, and the document content titles are identified based on regular expression comparison. The document titles are traversed and refined in the order of the document contents and integrated into a new document, which is the extracted complete word document chapter directory.

[0140] Optionally, in other embodiments, the document directory automatic generation program can also be divided into one or more modules, one or more modules are stored in the memory 11, and are executed by one or more processors (processor 12 in this embodiment) to complete the present invention. The module referred to in the present invention refers to a series of computer program instruction segments that can perform specific functions, and is used to describe the execution process of the document directory automatic generation program in the document directory automatic generation device.

[0141] For example, refer to Figure 3 As shown, it is a program module diagram of a document directory automatic generation program in an embodiment of a document directory automatic generation device of the present invention. In this embodiment, the document directory automatic generation program can be divided into a data receiving and processing module 10, a regular expression configuration module 20, a model training module 30, and a document directory output module 40. Exemplarily:

[0142] The data receiving and processing module 10 is used to receive an initial title in a target document and determine a title rule of the target document based on the initial title.

[0143] The regular expression configuration module 20 is used to configure a regular expression based on the trained title rule.

[0144] The model training module 30 is used to: input the initial title rules into a pre-built generative adversarial network model for training and obtain the trained title rules.

[0145] The document directory output module 40 is used to: receive a target document input by a user, determine the title rule of the target document, input the trained title rule and the configured regular expression into the document directory automatic generation model to generate a document directory and output it.

[0146] The functions or operation steps implemented when the above-mentioned program modules such as the data receiving and processing module 10, the regular expression configuration module 20, the model training module 30, the document directory output module 40 are executed are generally the same as those in the above-mentioned embodiments and will not be repeated here.

[0147] In addition, an embodiment of the present invention further provides a computer-readable storage medium, on which a document directory automatic generation program is stored. The document directory automatic generation program can be executed by one or more processors to implement the following operations:

[0148] An initial title in a target document is extracted, and a title rule of the target document is determined based on the initial title.

[0149] The primary article dataset and the primary summary dataset are word-vectorized and word-vector encoded to obtain a training set and a label set, respectively.

[0150] The initial title rules are input into a pre-built generative adversarial network model for training to obtain trained title rules.

[0151] Based on the trained title rule, a regular expression is configured.

[0152] The target document is read, the content in the target document is compared and analyzed with the regular expression, and the title is extracted.

[0153] Traverse the entire content of the target document, extract all the titles, arrange all the titles in the order of traversal, and generate a document directory.

[0154] It should be noted that the serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments. And the terms "including", "comprising" or any other variants thereof in this article are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "including a ..." does not exclude the presence of other identical elements in the process, device, article or method including the element.

[0155] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention.

[0156] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A method for automatically generating a document directory, characterized in that: The method comprises: Extracting an initial title from a target document, and determining an initial title rule of the target document based on the initial title; Inputting the initial title rule into a pre-built generative adversarial network model for training to obtain a trained title rule; Acquire the sentence body of the trained title rule and the semantic slot of the sentence body, and generate a regular expression according to the sentence body, the semantic slot and the remaining non-body part of the trained title rule; The entire content of the target document is traversed, the content in the target document is compared and analyzed with the regular expression, all the titles of the target document are extracted, all the titles are arranged in the order of traversal, and a document directory is generated.

2. The method for automatically generating a document directory according to claim 1, characterized in that: The document directory automatic generation method further includes: constructing the generative adversarial network model, including: Establish generative and discriminative models; The generative model and the discriminative model are subjected to mutual game learning to obtain an optimal solution, wherein the optimal solution includes the trained title rules.

3. The method for automatically generating a document directory according to claim 2, characterized in that: Before generating the regular expression, the document directory automatic generation method further includes: Generate a state machine based on the trained title rules; Wherein, the generation state machine includes: Performing grammatical analysis on the trained title rules, and rewriting the trained title rules into state machine rules required for state machine construction; Constructing a state machine according to the state machine rules; Convert the constructed state machine into the format required to generate regular expressions and store it.

4. The method for automatically generating a document directory according to claim 3, characterized in that: The traversing of the entire content of the target document, comparing and analyzing the content in the target document with the regular expression, and extracting all the titles of the target document includes: Traversing the entire content of the target document, and extracting one or more points of interest from the target document; Extracting the content of the target document through the points of interest, and identifying the outline structure of the target document; The outline structure of the target document is compared and matched with the regular expression. If the content in the target document matches the regular expression, the content in the target document is confirmed to be the title and the title is extracted. If the content in the target document does not match the regular expression, the content in the target document is confirmed to be text.

5. The method for automatically generating a document directory according to any one of claims 1 to 4, characterized in that: The document directory is in extensible markup language; The file format of the target document is Microsoft Office Word.

6. A document directory automatic generation device, characterized in that: The device includes a memory and a processor, wherein the memory stores a document directory automatic generation program that can be run on the processor, and when the document directory automatic generation program is executed by the processor, the following steps are implemented: Extracting an initial title from a target document, and determining an initial title rule of the target document based on the initial title; Inputting the initial title rule into a pre-built generative adversarial network model for training to obtain a trained title rule; Acquire the sentence body of the trained title rule and the semantic slot of the sentence body, and generate a regular expression according to the sentence body, the semantic slot and the remaining non-body part of the trained title rule; The entire content of the target document is traversed, the content in the target document is compared and analyzed with the regular expression, all the titles of the target document are extracted, all the titles are arranged in the order of traversal, and a document directory is generated.

7. The document directory automatic generation device according to claim 6, characterized in that: The document directory automatic generation method further includes: constructing the generative adversarial network model, including: Establish generative and discriminative models; The generative model and the discriminative model are subjected to mutual game learning to obtain an optimal solution, wherein the optimal solution includes the trained title rules.

8. The document directory automatic generation device according to claim 7, characterized in that: Before configuring the regular expression, the document directory automatic generation method further includes: Generate a state machine based on the trained title rules; Wherein, the generation state machine includes: Performing grammatical analysis on the trained title rules, and rewriting the trained title rules into state machine rules required for state machine construction; Constructing a state machine according to the state machine rules; Convert the constructed state machine into the format required to generate regular expressions and store it.

9. The document directory automatic generation device according to claim 8, characterized in that: Said Traversing the entire content of the target document, comparing and analyzing the content in the target document with the regular expression, and extracting all the titles of the target document, including: Traversing the entire content of the target document, and extracting one or more points of interest from the target document; Extracting the content of the target document through the points of interest, and identifying the outline structure of the target document; The outline structure of the target document is compared and matched with the regular expression. If the content in the target document matches the regular expression, the content in the target document is confirmed to be the title and the title is extracted. If the content in the target document does not match the regular expression, the content in the target document is confirmed to be text.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a document directory automatic generation program, which can be executed by one or more processors to implement the steps of the document directory automatic generation method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Automatic cataloguing method and system and computer readable storage medium

    CN109766433A