Document information extraction method and device, and storage medium
By parsing the document twice, and using machine learning models and extraction rules to extract software quality indicator information, the problems of low accuracy and poor efficiency of information extraction in the existing technology are solved, and more efficient and accurate extraction of software quality indicators are achieved.
Patent Information
- Application Number
- PCT/CN2024/090583
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-20
- Filing Date
- 2024-04-29
- Publication Date
- 2025-06-26
AI Technical Summary
When extracting software quality evaluation information from documents, the prior art has problems with low accuracy of information extraction and poor extraction efficiency.
A document information extraction method is proposed. By analyzing the to be processed documents twice, a document tree is constructed, and a software quality index category to which the node belongs is determined based on the trained machine learning model, and a software quality index information is extracted from the document tree in combination with the extraction rules.
It improves the efficiency and accuracy of software quality indicator extraction, and enhances the automation level and evaluation accuracy of software quality evaluation.
Smart Images

Figure CN2024090583_26062025_PF_FP_ABST
Abstract
Description
Document information extraction method, device and storage medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This disclosure is based on and claims priority to an application with CN application number 202311760500.0 and filing date December 20, 2023. The disclosure content of this CN application is hereby incorporated into this disclosure as a whole. Technical Field
[0003] The present disclosure relates to the field of computer technology, and in particular to a document information extraction method, device, and storage medium. Background Art
[0004] Throughout the software lifecycle, effectively measuring software quality helps manage and improve software, which is of great significance to both developers and users. In software development, the information required to evaluate software quality is usually scattered across multiple documents, such as requirements documents, design documents, test documents, maintenance documents, and user manuals.
[0005] In related technologies, information in documents is extracted based on extraction rules or based on machine learning models.
[0006] Summary of the Invention
[0007] The present disclosure provides a document information extraction method, device and storage medium.
[0008] According to a first aspect of the present disclosure, a document information extraction method is proposed, comprising: performing a first parsing of a document to be processed to obtain an order of nodes in the document to be processed; performing a second parsing of the document to be processed to obtain information contained in the nodes in the document to be processed; constructing a document tree according to the order of nodes in the document to be processed and the information contained in the nodes in the document to be processed; determining, based on a trained machine learning model, the software quality indicator category to which the nodes in the document tree belong; and extracting software quality indicator information from the nodes in the document tree according to extraction rules corresponding to the software quality indicator category.
[0009] In some embodiments, determining the software quality indicator category to which the nodes in the document tree belong based on the trained machine learning model includes: constructing an input vector according to the order of the nodes in the document to be processed and the information contained in the nodes in the document to be processed; and processing the input vector using the trained machine learning model to obtain the software quality indicator category to which the nodes in the document tree belong.
[0010] In some embodiments, the trained machine learning model is a transformer-based bidirectional encoder representation model.
[0011] In some embodiments, the document to be processed is a word document, and the second parsing of the document to be processed to obtain the information contained in the nodes in the document to be processed includes: creating a document object instance based on the document to be processed; and extracting the information contained in the node from the document object instance according to the type of the node in the document to be processed.
[0012] In some embodiments, the type of the node includes at least one of a paragraph, a table, and a directory, and the type of the paragraph includes at least one of a picture and text.
[0013] In some embodiments, the information of the node includes hierarchical representation information of the node, and constructing the document tree according to the order of the nodes in the document to be processed and the information contained in the nodes in the document to be processed includes: traversing the nodes in the document to be processed according to the order of the nodes in the document to be processed; determining the hierarchical structure to which the nodes in the document to be processed belong according to the hierarchical representation information of the nodes in the document to be processed; and constructing the document tree according to the hierarchical structure to which the nodes in the document to be processed belong and the information contained in the nodes in the document to be processed.
[0014] In some embodiments, the hierarchical representation information of the node includes at least one of the prefix information of the node and the text style of the node, and whether the node is a title text.
[0015] In some embodiments, the determining of the hierarchical structure to which the nodes in the document to be processed belong based on the hierarchical representation information of the nodes in the document to be processed includes: when the i-th node in the document to be processed is a title text, determining whether there is a traversed node that matches at least one of the prefix information and text style of the i-th node, where i is an integer greater than 1 and less than N; if so, the i-th node and the matching traversed node are at the same level; if not, the i-th node is the next level below the lowest level of the traversed nodes. Wherein, N is the total number of nodes,
[0016] In some embodiments, determining the hierarchical structure to which the nodes in the document to be processed belong based on the hierarchical representation information of the nodes in the document to be processed also includes: when the i-th node in the document to be processed is non-title text and the previous node of the i-th node is title text, the i-th node is the next layer of the previous node.
[0017] In some embodiments, determining the hierarchical structure to which the nodes in the document to be processed belong based on the hierarchical representation information of the nodes in the document to be processed also includes: when the i-th node in the document to be processed is non-title text and the previous node of the i-th node is non-title text, judging whether at least one of the prefix information and text style of the i-th node and the previous node of the i-th node matches; if they match, the i-th node and the previous node of the i-th node are at the same level; if they do not match, the i-th node is the next layer below the previous node of the i-th node.
[0018] In some embodiments, the document to be processed is a word document, and the first parsing of the document to be processed to obtain the node order in the document to be processed includes: modifying the suffix of the document to be processed to a compressed file format to obtain a compressed package; decompressing the compressed package to obtain a target folder, and the target folder includes a document file in XML format; parsing the document file in XML format to obtain the node order in the document to be processed.
[0019] In some embodiments, the extraction rules include regular matching rules.
[0020] According to a second aspect of the present disclosure, a document information extraction device is proposed, comprising: a first parsing module, configured to perform a first parsing on a document to be processed to obtain an order of nodes in the document to be processed; a second parsing module, configured to perform a second parsing on the document to be processed to obtain information contained in the nodes in the document to be processed; a construction module, configured to construct a document tree according to the order of nodes in the document to be processed and the information contained in the nodes in the document to be processed; a classification module, configured to determine, based on a trained machine learning model, a software quality indicator category to which the nodes in the document tree belong; and an extraction module, configured to extract software quality indicator information from the nodes in the document tree according to extraction rules corresponding to the software quality indicator category.
[0021] According to a third aspect of the present disclosure, a document information extraction device is proposed, comprising: a memory; and a processor coupled to the memory, wherein the processor is configured to execute the document information extraction method as described above based on instructions stored in the memory.
[0022] According to a fourth aspect of the present disclosure, a computer storable medium is proposed, on which computer program instructions are stored, and when the instructions are executed by a processor, the document information extraction method as described above is implemented.
[0023] Other features and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0025] The present disclosure can be more clearly understood from the following detailed description with reference to the accompanying drawings.
[0026] FIG1 is a flow chart of a document information extraction method according to some embodiments of the present disclosure.
[0027] FIG2 is a schematic diagram of a process of constructing a document tree according to some embodiments of the present disclosure.
[0028] FIG3 is a schematic diagram of a document tree constructed according to some embodiments of the present disclosure.
[0029] FIG4 is a schematic diagram of a process of determining a hierarchical structure of a node according to some embodiments of the present disclosure.
[0030] FIG5 is a schematic structural block diagram of a document information extraction device according to some embodiments of the present disclosure.
[0031] FIG6 is a schematic structural diagram of a document information extraction device according to other embodiments of the present disclosure.
[0032] FIG7 is a schematic diagram of the structure of a computer system according to some embodiments of the present disclosure. DETAILED DESCRIPTION
[0033] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that unless otherwise specifically stated, the relative arrangement of components and steps, numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure.
[0034] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0035] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.
[0036] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.
[0037] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not limiting. Therefore, other examples of the exemplary embodiments may have different values.
[0038] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.
[0039] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.
[0040] The inventors of the present disclosure have discovered that the document information extraction methods in related technologies have problems such as low information extraction accuracy and poor extraction efficiency. How to efficiently and accurately extract the required software quality assessment information from documents remains a challenging problem.
[0041] In view of this, the present disclosure proposes a document information extraction method, device, electronic device and storage medium to improve the efficiency and accuracy of software quality indicator extraction.
[0042] Figure 1 is a flow chart of a document information extraction method according to some embodiments of the present disclosure. As shown in Figure 1 , the document information extraction method includes steps S101 to S105.
[0043] In step S101 , a first parsing is performed on the document to be processed to obtain the node sequence in the document to be processed.
[0044] In some embodiments, the document information extraction method is performed by a document information extraction device.
[0045] In some embodiments, the document to be processed is a Word document. For example, the document to be processed is a Word document with a suffix of docx.
[0046] Nodes in the document to be processed are also called elements. Node types include at least one of paragraphs, tables, and directories. Paragraphs can be divided into images and text.
[0047] In some embodiments, the document to be processed is a Word document, and the first parsing step of the document to be processed includes: modifying the suffix of the document to be processed to a compressed file format to obtain a compressed package; decompressing the compressed package to obtain a target folder, and the target folder includes a document file in XML format; parsing the document file in XML format to obtain the node order in the document to be processed.
[0048] In some examples, the suffix of the word document to be processed is changed from "docx" to "zip" to obtain a compressed package, and then the compressed package is decompressed to obtain a target folder. The target folder contains multiple document files in Extensible Markup Language (XML) format, such as a document.xml file. Next, the document.xml file is parsed based on the etree.parse method in the Lxml library to obtain an element tree object (ElementTree object), and then all nodes in the element tree object are traversed to obtain the order of the nodes, and the order of the nodes is stored. The Lxml library is a parser that can parse XML files. Its main function is to parse and extract data from XML files.
[0049] In some examples, the order of the nodes is stored in a list. The order of the nodes includes an overall number and a local number. The overall number is the numbering of all types of nodes in the order in which they are in the document, and the local numbering is the numbering of the nodes of the same type in the order in which they are in the document. For example, the overall numbers are 001, 002, 003, 004, 005, 006, where 001 to 003 are the overall numbers of the nodes of the paragraph type, and the local numbers of the nodes of these three paragraph types are (p, 0), (p, 1), and (p, 2); 004 is the overall number of the node of the table type, and the local number of the node of the table type is (t, 0); 005 to 006 are the overall numbers of the nodes of the paragraph type, and the local numbers of the nodes of these two paragraph types are (p, 3) and (p, 4).
[0050] In step S102, a second analysis is performed on the document to be processed to obtain information contained in the nodes in the document to be processed.
[0051] In some embodiments, the document to be processed is a word document, and the second analysis of the document to be processed includes: creating a document object instance based on the document to be processed; and extracting information contained in the node from the document object instance according to the type of the node in the document to be processed.
[0052] In some embodiments, the type of the node includes at least one of a paragraph, a table, and a directory, and the type of the paragraph includes at least one of a picture and a text. In some examples, the picture includes a Visio diagram.
[0053] In some examples, the Python docx library is used to create document object instances (i.e., document object instances) for the documents to be processed. The Python docx library is a third-party library for processing Word documents. It directly operates on Word documents and supports reading, querying, and modifying files in formats such as docx. This allows for efficient access to node information such as paragraphs and tables.
[0054] In some embodiments, according to the node order determined in step S101 and in accordance with the extraction method corresponding to the node type, node information is extracted from the document to be processed.
[0055] In some examples, when the node type is a paragraph, the information contained in the paragraph is extracted as follows: determine whether the paragraph is a picture or text; if the paragraph is a picture (such as a Visio diagram), query the information contained in the picture based on the picture storage path; if the paragraph is text, query the information contained in the text based on the text storage path.
[0056] In some examples, if the paragraph is a picture, such as a Visio diagram, a Visio object instance is first created, and then based on the Visio object instance, information of the Visio diagram is obtained from a storage path of the Visio diagram and stored.
[0057] For example, based on the XML query language Xpath, it is necessary to query whether the paragraph object contains<w:object> Tags; if included<w:object> Tag, use the getElementsByTagName method to find out whether the tag contains <o:oleobject>Tags; if included <o:oleobject>Tag, use the getAttribute method to get the RID attribute, find the image storage path based on the RID attribute; then extract the image information based on the image storage path.
[0058] In some examples, if the paragraph is text, the storage path of the text is found, and the paragraph style, font size, paragraph type (such as whether it is a title), text information in the paragraph, and paragraph prefix information are extracted from the storage path of the text.
[0059] In some examples, when extracting the prefix of a paragraph, if there is no automatic numbering prefix, determine whether there is a regular type prefix; if so, truncate the prefix; if not, mark the prefix as empty; if there is an automatic numbering prefix, extract the prefix.
[0060] In some examples, when the node type is a table, the information in the table is queried according to the storage path of the table. For example, from the document object instance<w:tbl> Under the label, extract the information of each cell in the table by row and column, and store the extracted information.
[0061] In some instances, when the node type is a directory, the information in the directory is queried based on the directory's storage path. For example, from the document object instance<w:sdt> Under the tag, extract the text information of each line in the directory and store the extracted information.
[0062] In some embodiments, the first parsing step and the second parsing step are performed based on different parsing tools, for example, the first parsing step is performed based on the Lxml library, and the second parsing step is performed based on the Python docx library. By combining the above two parsing tools, the efficiency of document information extraction is improved.
[0063] In step S103, a document tree is constructed according to the order of the nodes in the document to be processed and the information contained in the nodes in the document to be processed.
[0064] In some embodiments, the document tree is constructed according to the process shown in FIG. 2 .
[0065] In step S104, based on the trained machine learning model, the software quality indicator category to which the node in the document tree belongs is determined.
[0066] In some embodiments, in step S104, an input vector is constructed based on the information contained in the nodes in the document to be processed; the input vector is processed using a trained machine learning model to obtain the software quality indicator category to which the node in the document tree belongs.
[0067] In some embodiments, in step S104, an input vector is constructed based on the order of nodes in the document to be processed and the information contained in the nodes in the document to be processed; the input vector is processed using a trained machine learning model to obtain the software quality indicator category to which the node in the document tree belongs.
[0068] In some examples, the software quality indicator category to which a node belongs includes one or more of the following primary indicators: functional adaptability, performance efficiency, compatibility, usability, reliability, information security, maintainability, and portability. A trained machine learning model can be used to determine the software quality indicator category to which each node belongs, facilitating the subsequent rapid and accurate extraction of more granular indicator information.
[0069] In some examples, the trained machine learning model is a transformer-based bidirectional encoder representation model (BERT). The BERT model uses a bidirectional transformer model with a self-attention mechanism to learn and obtain the semantic relationship between sentences. This makes it easier to parse and obtain the semantic expression of text language.
[0070] In some examples, the information contained in the node is preprocessed before using the trained BERT model. This preprocessing includes tokenizing the text contained in the node. For example, the BertTokenizer method in the transformers package is used to tokenize the text content in the node.
[0071] In some examples, before using the trained BERT model, the loss function of the BERT model is optimized based on the stochastic gradient descent method.
[0072] In the embodiment of the present disclosure, by comprehensively considering the node order and the information contained in the node when determining the category of the node based on the machine learning model, it is helpful to improve the accuracy of node classification.
[0073] In step S105 , software quality indicator information is extracted from the nodes in the document tree according to extraction rules corresponding to the software quality indicator categories.
[0074] In some embodiments, the extraction rules include regular matching rules, which are matching rules corresponding to specific indicators included in the software quality indicator category.
[0075] For example, when the software quality indicator category to which the node belongs is functional adaptability, the specific indicators included in this category are one or more of functional coverage, functional correctness, functional suitability of usage targets, and functional suitability of the system. The regular matching rules corresponding to these specific indicators include: one or more of the regular matching rules corresponding to functional coverage, the regular matching rules corresponding to functional correctness, the regular matching rules corresponding to functional suitability of usage targets, and the regular matching rules corresponding to the functional suitability of the system.
[0076] For example, when the software quality indicator category to which the node belongs is performance efficiency, the specific indicators included in this category are one or more of average response time, sufficiency of response time, average turnaround time, sufficiency of turnaround time, average throughput, average processor occupancy, average memory occupancy, average input and output device occupancy, bandwidth occupancy, transaction processing capacity, user access volume, and sufficiency of user access growth. The regular matching rules corresponding to these specific indicators include: regular matching rules corresponding to average response time, regular matching rules corresponding to sufficiency of response time, regular matching rules corresponding to average turnaround time, regular matching rules corresponding to sufficiency of turnaround time, regular matching rules corresponding to average throughput, regular matching rules corresponding to average processor occupancy, regular matching rules corresponding to average memory occupancy, regular matching rules corresponding to average input and output device occupancy, regular matching rules corresponding to bandwidth occupancy, regular matching rules corresponding to transaction processing capacity, regular matching rules corresponding to user access volume, and one or more of regular matching rules corresponding to sufficiency of user access growth.
[0077] In some examples, the regular matching rule is represented by a regular expression. For example, the regular matching rule for the average response time designed for a test document can be represented as ".*?([1-9]\d*\.\d+|0\.\d+|0|[1-9]\d*)(ms|s).*".
[0078] In some examples, different regular expression matching rules are designed for different document types. For example, the document types include one or more of requirements documents, design documents, test documents, maintenance documents, and user manuals. By designing different regular expression matching rules for different document types, not only can the information extraction requirements of various documents be met, improving the applicability of the document information extraction method, but also helping to improve the accuracy of information extraction.
[0079] For example, the regular matching rule designed for the "adequacy of response time" indicator in the requirement document is .*? ([1-9]\d*\.\d+|0\.\d+|0|[1-9]\d*)(ms|s).*".
[0080] For example, the regular matching rule designed for the "average turnaround time" indicator in the test document is ".*?([1-9]\d*\.\d+|0\.\d+|0|[1-9]\d*)(ms|s).*".
[0081] In an embodiment of the present disclosure, after the nodes in the document are classified by a machine learning model, software quality indicator information is extracted in a targeted manner based on the extraction rules corresponding to the classification, which can improve the accuracy and efficiency of information extraction.
[0082] In the embodiments disclosed herein, on the one hand, by extracting not only the information contained in each node but also the order of each node, and constructing a document tree based on the above two aspects of information, it helps to improve the efficiency and accuracy of subsequent document information extraction; on the other hand, by combining the machine learning model with the extraction rules to extract software quality indicator information, compared with the related technology of extracting information based only on the machine learning model or only based on the extraction rules, it can improve the efficiency and accuracy of software quality indicator extraction, thereby helping to improve the automation level and evaluation accuracy of software quality assessment.
[0083] Figure 2 is a schematic diagram of a process for constructing a document tree according to some embodiments of the present disclosure. As shown in Figure 2, the process for constructing a document tree includes steps S201 to S203.
[0084] In step S201 , the nodes in the document to be processed are traversed in the order of the nodes in the document to be processed.
[0085] For example, the order of the nodes in the document to be processed is 001, 002, 003, 004, 005, 006, where 001 to 003 are the overall numbers of the nodes of the paragraph type, 004 is the overall number of the node of the table type, and 005 to 006 are the overall numbers of the nodes of the paragraph type. Then, each node is traversed in the above order.
[0086] In step S202 , the hierarchical structure to which the nodes in the document to be processed belong is determined based on the hierarchical representation information of the nodes in the document to be processed.
[0087] In some embodiments, the hierarchical representation information of the node in the document to be processed includes at least one of the prefix information of the node and the text style of the node, and whether the node is a title text.
[0088] In some examples, the title text includes title types such as Heading, Title, and subtitle; the non-title text includes non-title types such as List Paragraph, Table, and template file (such as Normal).
[0089] In some embodiments, the hierarchical structure to which a node belongs is determined as follows: when the i-th node in the document to be processed is title text, it is determined whether there is a traversed node that matches at least one of the prefix information and text style of the i-th node, where i is an integer greater than 1 and less than N; if so, the i-th node and the matching traversed node are at the same level; if not, the i-th node is at the next level below the lowest level of the traversed nodes, where N is the total number of nodes.
[0090] In some embodiments, the hierarchical structure to which a node belongs is determined as follows: when the i-th node in the document to be processed is non-title text and the previous node of the i-th node is title text, the i-th node is the next layer of the previous node.
[0091] In some embodiments, the hierarchical structure to which a node belongs is determined as follows: when the i-th node in the document to be processed is non-title text and the previous node of the i-th node is non-title text, determine whether at least one of the prefix information and text style of the i-th node and the previous node of the i-th node matches; if they match, the i-th node and the previous node of the i-th node are at the same level; if they do not match, the i-th node is at the next level below the previous node of the i-th node.
[0092] In some embodiments, the hierarchical structure to which the node belongs is determined according to the embodiment shown in FIG. 4 .
[0093] In step S203, a document tree is constructed according to the hierarchical structure to which the nodes in the document to be processed belong and the information contained in the nodes in the document to be processed.
[0094] In the embodiment of the present disclosure, the hierarchical structure in the document to be processed is determined based on the above multiple hierarchical representation information, and a document tree is constructed based on the hierarchical structure and node information, so as to facilitate subsequent rapid and accurate information extraction.
[0095] FIG3 is a schematic diagram of a document tree constructed according to some embodiments of the present disclosure. As shown in FIG3 , the document tree includes multiple nodes. In FIG3 , circles represent nodes, the numbers within the circles represent the order of the nodes, and the relationships between the nodes represent the hierarchical structure of the nodes.
[0096] In some examples, when extracting the software quality indicator "Number of Considered Features" from a document to be processed, first find node 59 (which corresponds to "Chapter 3" in the document to be processed), then search for the child nodes of node 59 and find node 65 (which corresponds to "Section 3.2" in the document to be processed). Then, search for the grandchild nodes of node 65 (which correspond to the fourth-level heading "3.2.XX" in the document to be processed), and the number of grandchild nodes is the value of "Number of Considered Features".
[0097] Figure 4 is a schematic diagram of a process for determining a hierarchical structure of a node according to some embodiments of the present disclosure. As shown in Figure 4 , the process for determining a hierarchical structure of a node includes steps S401 to S411.
[0098] In step S401, the i-th node in the document to be processed is obtained.
[0099] In the embodiment of the present disclosure, i is an integer greater than 1 and less than N. In addition, the node with i equal to 1 is the root node, that is, the node at the highest level, where N is the total number of nodes.
[0100] In some embodiments, before executing step S401, the process further includes: normalizing the formats of the nodes in the document to be processed. For example, the normalization process includes normalizing the text style, text prefix, etc. in the document.
[0101] In step S402, it is determined whether the i-th node is a title text.
[0102] When the i-th node is a title text, step S403 is executed; when the i-th node is not a title text, step S407 is executed.
[0103] In step S403 , it is determined whether there is a node in the traversed nodes that matches the text prefix of the i-th node.
[0104] If the judgment result of step S403 is yes, step S405 is executed; if the judgment result of step S403 is no, step S404 is executed.
[0105] For example, the text prefix of the i-th node is "3.2", which matches a node with a text prefix of "3.1" existing in the traversed nodes. In this case, step S405 is executed.
[0106] For example, the text prefix of the i-th node is "3.1", and there is no node matching its text prefix in the traversed nodes. In this case, step S404 is executed.
[0107] In step S404 , it is determined whether there is a node in the traversed nodes that matches the text style of the i-th node.
[0108] If the judgment result of step S404 is yes, execute step S405; otherwise, execute step S406.
[0109] For example, the text style of the i-th node is font size 3, bold, and indented by 2 characters. If there is a node in the traversed nodes that matches the text style of the i-th node, step S405 is executed; otherwise, step S406 is executed.
[0110] In step S405 , it is determined whether the level of the i-th node is consistent with the level of the matched traversed nodes.
[0111] In step S406 , it is determined that the i-th node is the next layer of the lowest layer of the traversed nodes.
[0112] For example, when the lowest level of traversed nodes is the third level, the i-th node is regarded as the fourth level.
[0113] In step S407 , it is determined whether the previous node of the i-th node is a title text.
[0114] If the judgment result of step S407 is yes, execute step S408; otherwise, execute step S409.
[0115] In step S408 , it is determined that the i-th node is a next level below the previous node.
[0116] In step S409 , it is determined whether the text prefix of the i-th node matches the text prefix of the previous node.
[0117] If the judgment result of step S409 is yes, execute step S410; otherwise, execute step S411.
[0118] In step S410 , it is determined that the i-th node and the previous node are at the same level.
[0119] In step S411 , it is determined whether the text style of the i-th node matches that of the previous node.
[0120] If the judgment result of step S411 is yes, step S410 is executed, that is, determining that the i-th node and the previous node are at the same level; otherwise, step S408 is executed, that is, determining that the i-th node is at the next level of the previous node.
[0121] In the embodiment of the present disclosure, the above process can quickly construct a document tree corresponding to the document to be processed, thereby facilitating subsequent information extraction based on the document tree.
[0122] FIG5 is a schematic structural block diagram of a document information extraction apparatus according to some embodiments of the present disclosure. As shown in FIG5 , the document information extraction apparatus 500 includes a first parsing module 501 , a second parsing module 502 , a construction module 503 , a classification module 504 , and an extraction module 505 .
[0123] The first parsing module 501 is configured to perform a first parsing on the document to be processed to obtain a node sequence in the document to be processed.
[0124] In some embodiments, the documents to be processed include one or more of requirement documents, design documents, test documents, maintenance documents, and user manuals.
[0125] In some embodiments, the first parsing module 501 obtains the node order in the document to be processed in the following manner: modifying the suffix of the document to be processed to a compressed file format to obtain a compressed package; decompressing the compressed package to obtain a target folder, which includes a document file in XML format; parsing the document file in XML format to obtain the node order in the document to be processed.
[0126] The second parsing module 502 is configured to perform a second parsing on the document to be processed to obtain information contained in the nodes in the document to be processed.
[0127] In some embodiments, the second parsing module 502 obtains information contained in nodes in the document to be processed in the following manner: creating a document object instance based on the document to be processed; and extracting information contained in the nodes from the document object instance according to the type of the nodes in the document to be processed.
[0128] The construction module 503 is configured to construct a document tree according to the order of the nodes in the document to be processed and the information contained in the nodes in the document to be processed.
[0129] In some embodiments, the construction module 503 constructs a document tree in the following manner: traverse the nodes in the document to be processed in the order of the nodes in the document to be processed; determine the hierarchical structure to which the nodes in the document to be processed belong based on the hierarchical representation information of the nodes in the document to be processed; and construct a document tree based on the hierarchical structure to which the nodes in the document to be processed belong and the information contained in the nodes in the document to be processed.
[0130] The classification module 504 is configured to determine the software quality indicator classification to which the node in the document tree belongs based on the trained machine learning model.
[0131] The extraction module 505 is configured to extract software quality indicator information from the nodes in the document tree according to extraction rules corresponding to the software quality indicator classification.
[0132] In the embodiments of the present disclosure, the above apparatus can improve the efficiency and accuracy of extracting software quality indicators, thereby helping to improve the automation level and evaluation accuracy of software quality assessment.
[0133] FIG6 is a schematic structural diagram of a document information extraction device according to other embodiments of the present disclosure.
[0134] As shown in FIG6 , a document information extraction apparatus 600 includes a memory 601 and a processor 602 coupled to the memory 601. The memory 601 is configured to store instructions for executing the document information extraction method according to the embodiments. The processor 602 is configured to execute the document information extraction method according to any of the embodiments of the present disclosure based on the instructions stored in the memory 601.
[0135] FIG7 is a schematic diagram of the structure of a computer system according to some embodiments of the present disclosure.
[0136] As shown in Figure 7, a computer system 700 may be a general-purpose computing device and includes a memory 701, a processor 702, and a bus 703 connecting various system components.
[0137] Memory 701 may include, for example, system memory, non-volatile storage media, and the like. System memory may store, for example, an operating system, application programs, a boot loader, and other programs. System memory may include volatile storage media, such as random access memory (RAM) and / or cache memory. Non-volatile storage media may store, for example, instructions for executing at least one embodiment of the document information extraction method. Non-volatile storage media include, but are not limited to, disk storage, optical storage, flash memory, and the like.
[0138] The processor 702 can be implemented as a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, or as discrete hardware components such as discrete gates or transistors. Accordingly, each module, such as the first parsing module, the second parsing module, the building module, the classification module, and the extraction module, can be implemented by a central processing unit (CPU) executing instructions in a memory for executing the corresponding steps, or by dedicated circuits for executing the corresponding steps.
[0139] The bus 703 may use any of a variety of bus architectures, including, but not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, and a Peripheral Component Interconnect (PCI) bus.
[0140] The computer system 700 interfaces 704, 705, and 706, as well as the memory 701 and the processor 702, can be connected via a bus 703. The input / output interface 704 provides a connection interface for input / output devices such as a display, mouse, and keyboard. The network interface 705 provides a connection interface for various networked devices. The storage interface 706 provides a connection interface for external storage devices such as floppy disks, USB flash drives, and SD cards.
[0141] Here, various aspects of the present disclosure are described with reference to flowcharts and / or block diagrams of methods, devices, and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks, can be implemented by computer-readable program instructions.
[0142] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable device to produce a machine, so that the processor executes the instructions to produce means for implementing the functions specified in one or more blocks in the flowcharts and / or block diagrams.
[0143] These computer-readable program instructions may also be stored in a computer-readable memory, which cause the computer to operate in a specific manner to produce an article of manufacture, including instructions for implementing the functions specified in one or more blocks in the flowcharts and / or block diagrams.
[0144] The present disclosure can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects.
[0145] The document information extraction method, device, and storage medium in the above-mentioned embodiments can improve the accuracy and efficiency of extracting software quality indicator information.
[0146] The document information extraction method, apparatus, and storage medium disclosed herein have been described in detail. To avoid obscuring the scope of the present disclosure, some details known in the art have been omitted. Based on the above description, those skilled in the art will readily understand how to implement the technical solutions disclosed herein.< / o:oleobject> < / o:oleobject>
Claims
1. A document information extraction method, comprising: Performing a first analysis on the document to be processed to obtain a node sequence in the document to be processed; Performing a second analysis on the document to be processed to obtain information contained in the nodes in the document to be processed; Constructing a document tree according to the order of nodes in the document to be processed and the information contained in the nodes in the document to be processed; Based on the trained machine learning model, determining the software quality indicator category to which the node in the document tree belongs; Software quality indicator information is extracted from the nodes in the document tree according to an extraction rule corresponding to the software quality indicator category.
2. The document information extraction method according to claim 1, wherein: Determining the software quality indicator category to which the node in the document tree belongs based on the trained machine learning model includes: Constructing an input vector according to the order of nodes in the document to be processed and the information contained in the nodes in the document to be processed; The input vector is processed using the trained machine learning model to obtain the software quality indicator category to which the node in the document tree belongs.
3. The document information extraction method according to claim 2, wherein: The trained machine learning model is a transformer-based bidirectional encoder representation model.
4. The document information extraction method according to claim 1, wherein: The document to be processed is a word document, and the second parsing of the document to be processed to obtain information contained in the nodes in the document to be processed includes: Based on the document to be processed, create a document object instance; According to the type of the node in the document to be processed, the information contained in the node is extracted from the document object instance.
5. The document information extraction method according to claim 4, wherein: The types of the nodes include at least one of a paragraph, a table, and a directory, and the types of the paragraphs include at least one of a picture and a text.
6. The document information extraction method according to any one of claims 1 to 5, wherein: The node information includes hierarchical representation information of the node, and constructing the document tree according to the order of the nodes in the document to be processed and the information contained in the nodes in the document to be processed includes: Traversing the nodes in the document to be processed according to the order of the nodes in the document to be processed; Determining the hierarchical structure to which the nodes in the document to be processed belong according to the hierarchical representation information of the nodes in the document to be processed; A document tree is constructed according to the hierarchical structure to which the nodes in the document to be processed belong and the information contained in the nodes in the document to be processed.
7. The document information extraction method according to claim 6, wherein: The hierarchical representation information of the node includes at least one of the prefix information of the node and the text style of the node, and whether the node is a title text.
8. The document information extraction method according to claim 7, wherein: The step of determining the hierarchical structure to which the node in the to-be-processed document belongs according to the hierarchical representation information of the node in the to-be-processed document comprises: When the i-th node in the document to be processed is a title text, determining whether there is a traversed node that matches at least one of the prefix information and the text style of the i-th node, where i is an integer greater than 1 and less than N; If so, the i-th node and the matched traversed node are at the same level; If not present, the i-th node is the next layer of the lowest level among the traversed nodes.
9. The document information extraction method according to claim 8, wherein: The step of determining the hierarchical structure to which the node in the to-be-processed document belongs according to the hierarchical representation information of the node in the to-be-processed document further comprises: In the case that the i-th node in the document to be processed is non-title text and the previous node of the i-th node is title text, the i-th node is the next layer of the previous node.
10. The document information extraction method according to claim 8, wherein: The step of determining the hierarchical structure to which the node in the to-be-processed document belongs according to the hierarchical representation information of the node in the to-be-processed document further comprises: The i-th node in the document to be processed is non-title text, and the previous node of the i-th node In the case of non-title text, determining whether at least one of the prefix information and the text style of the i-th node matches that of the previous node of the i-th node; If they match, the i-th node and the previous node of the i-th node are at the same level; If there is no match, the i-th node is the next layer of the previous node of the i-th node.
11. The document information extraction method according to any one of claims 1 to 5, wherein: The document to be processed is a word document, and performing a first analysis on the document to be processed to obtain a node sequence in the document to be processed includes: Modifying the suffix of the document to be processed into a compressed file format to obtain a compressed package; Decompressing the compressed package to obtain a target folder, wherein the target folder includes a document file in XML format; The document file in the XML format is parsed to obtain the node sequence in the document to be processed.
12. The document information extraction method according to any one of claims 1 to 5, wherein: The extraction rules include regular matching rules.
13. A document information extraction device, comprising: A first parsing module is configured to perform a first parsing on the document to be processed to obtain a node sequence in the document to be processed; A second parsing module is configured to perform a second parsing on the document to be processed to obtain information contained in the nodes in the document to be processed; A construction module, configured to construct a document tree according to the order of nodes in the document to be processed and the information contained in the nodes in the document to be processed; A classification module, configured to determine the software quality indicator classification to which the node in the document tree belongs based on the trained machine learning model; The extraction module is configured to extract software quality indicator information from the nodes in the document tree according to an extraction rule corresponding to the software quality indicator classification.
14. A document information extraction device, comprising: Memory; as well as A processor coupled to the memory, wherein the processor is configured to execute the document information extraction method according to any one of claims 1 to 12 based on instructions stored in the memory.
15. A computer storable medium having computer program instructions stored thereon, wherein the instructions, when executed by a processor, implement the document information extraction method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Document title tree construction method and device, electronic equipment and storage medium
CN111460083A
Method and device for checking software system design document
CN116680199A
Document information extraction method and device and storage medium
CN117669501A
Methods and systems for automated software natural language documentation
US20220261241A1
Methods and systems for scoring quality of open source software documentation
US20220276862A1