A Fine-Grained Automatic Extraction Method and System for Third-Party Component Documents Based on a Question-Answering Model

Through a question-and-answer model-based method, the attention model and natural language processing model are used to extract fine-grained usage rules in third-party component documents, which solves the problem of inaccurate extraction rules in the prior art and improves the accuracy of misuse detection.

CN114841124BActive Publication Date: 2025-05-30ZHEJIANG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210331439.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-30
Publication Date
2025-05-30
Estimated Expiration
2042-03-30

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively extract the fine-grained usage rules in third-party component documents, which affects the accuracy of third-party component misuse detection.

Method used

Using a method based on the question-and-answer model, document preprocessing is carried out through the attention model, document warehouse is built and question-and-answer model is designed, and documents are extracted in fine-grained manner using natural language processing models.

Benefits of technology

It realizes accurate extraction of third-party component usage rules, can effectively adapt to third-party component documents in different formats, and improves the accuracy of misuse detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114841124B_ABST
    Figure CN114841124B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for fine-grained automatic extraction of third-party component documents based on a question-and-answer model, belonging to the technical field of third-party component testing. The system includes: a third-party component document preprocessing module that preliminarily filters the third-party component documents to obtain coarse-grained third-party component usage rules; a document question-and-answer tree construction module that deeply analyzes the misuse types of third-party components, designs query questions for each type of misuse, and manually marks the document to be tested according to the questions; a third-party component usage rule extraction module based on question-and-answer that uses a natural language processing model based on the RoBERTa model to perform question-and-answer information extraction on the document to obtain fine-grained usage rules related to the third-party components. The system of the present invention solves the problem of coarse-grained refinement of third-party component documents without a unified format and can perform fine-grained automatic extraction of the usage rules in the third-party component documents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of third-party component testing, and in particular to a method and system for fine-grained automatic extraction of third-party component documents based on a question-and-answer model. Background Art

[0002] With the continuous advancement of the open-source community, various third-party components have seen booming development and are currently widely used in the development of software in all walks of life. However, recent research and actual incidents have shown that while using third-party components brings convenience to software developers, the security issues arising during the use of third-party components are a cause for concern. Due to the lack of effective methods to strictly regulate developers during the use of third-party components, software developed based on various third-party components may pose serious security threats. For example, the functions of some third-party components require to be released after being called, and developers may overlook or omit similar usage rules, thus causing serious security threats. Such threats can range from affecting user privacy to endangering the security of national critical equipment.

[0003] In order to detect whether developers strictly follow the usage rules when using third-party components, researchers have proposed various detection systems for mining such misuse cases of third-party components. These detection systems have one thing in common, that is, they all need to accurately obtain the usage rules of third-party components. Currently, researchers mainly obtain the corresponding usage rules from third-party component documents through methods such as manual acquisition, regular expression matching, and syntactic dependency trees. However, the method of obtaining usage rules by manually reading third-party components is time-consuming and laborious. A third-party component often contains hundreds of functions, and each function has multiple usage rules. Therefore, a third-party component can contain up to thousands of usage rules. Secondly, the usage rules obtained by the regular expression matching method often result in a large number of false negatives and cannot comprehensively obtain the usage rules of third-party components, thus affecting the accuracy of subsequent detection of misuse cases of third-party components. In addition, mining rules through the syntactic dependency tree method requires a large amount of document preprocessing work, and it has poor effects on the third-party component documents with loose structures and is difficult to apply to large-scale detection.

[0004] There are the following challenges in designing an effective fine-grained automated extraction method for third-party component documents: (1) Adapting to third-party component documents in different formats. Currently, there is no unified format for writing third-party component documents, and there may be significant differences between any two third-party component documents. The method of only using regular expression matching cannot be applied to different categories of third-party component documents. (2) Comprehensively obtaining the usage rules of third-party components. Since there are a large number of interfering statements in third-party component documents, such as the functional descriptions of functions, it greatly affects the difficulty of comprehensively obtaining the usage rules of third-party components, resulting in a large number of false alarms. Secondly, the usage rules of some third-party components are described ambiguously, and sometimes it is impossible to determine whether they are real usage rules even through manual judgment.

[0005] Since the documents of third-party components do not have a unified format and their usage rules vary greatly, there is currently no effective method for automatically and fine-grainedly obtaining the usage rules of third-party components. Designing a method that can automatically and fine-grainedly extract the corresponding usage rules from third-party component documents is important and necessary for subsequent detection of vulnerabilities caused by misuse of third-party components. Summary of the Invention

[0006] Aiming at the deficiencies in the fine-grained automated extraction of third-party component usage rules, the present invention provides a fine-grained automated extraction method and system for third-party component documents based on a question-answering model, which can accurately extract the usage rules of each function in third-party components.

[0007] The specific technical solution of the present invention is as follows:

[0008] The first object of the present invention is to provide a fine-grained automated extraction method for third-party component documents based on a question-answering model, including the following steps:

[0009] Step 1: Collect documents of multiple different third-party components, preprocess the documents, and build a document repository; use an attention model to refine the sentences of the documents to be tested in the document repository to obtain the coarse-grained usage rules of third-party components;

[0010] Step 2: Design corresponding questions for the question-answering model according to the misuse types of third-party components; select some documents from the documents to be tested in the document repository and mark the answers according to the designed questions;

[0011] Step 3: Divide the marked documents to be tested into a training set and a validation set, use the training set to train a natural language processing model until the test accuracy of the validation set meets the preset requirements; use the trained natural language processing model to perform fine-grained mining on the coarse-grained usage rules of the remaining unmarked answer documents in the document repository.

[0012] The second objective of the present invention is to provide a third - party component document fine - grained automatic extraction system based on a question - answering model for implementing the above - mentioned method. The extraction system includes:

[0013] A third - party component document pre - processing module, which is used to collect documents of multiple different third - party components, pre - process the documents, and build a document repository; use an attention model to refine sentences of the documents to be tested in the document repository, and obtain the coarse - grained usage rules of the third - party components;

[0014] A document question - answering tree construction module, which is used to design corresponding questions of the question - answering model according to the misuse types of the third - party components; select some documents from the documents to be tested in the document repository, and mark the answers according to the designed questions;

[0015] A third - party component usage rule extraction module based on question - answering, which is used to divide the marked documents to be tested into a training set and a validation set, train a natural language processing model using the training set until the test accuracy rate of the validation set meets the preset requirements; use the trained natural language processing model to conduct fine - grained mining on the coarse - grained usage rules of the remaining unmarked answer documents in the document repository.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0017] (1) The present invention provides a third - party component document fine - grained automatic extraction system, and proposes a third - party component document fine - grained automatic extraction method based on a question - answering model, which solves the problem of fine - grained automatic extraction of third - party component usage rules, can effectively extract the usage rules of different types of third - party components, and has practicality;

[0018] (2) The present invention provides a document pre - processing method based on an attention model, which solves the problem of coarse - grained refinement of third - party component documents without a unified format; the present invention provides a document content extraction method based on a question - answering model, which provides a basis for constructing a document question - answering tree and fine - grained extraction of third - party component usage rules without a unified format. Brief Description of the Drawings

[0019] Figure 1 It is a schematic diagram of the overall module structure of the third - party component document fine - grained automatic extraction system based on a question - answering model;

[0020] Figure 2 It is a schematic diagram of the process of the third - party component document fine - grained automatic extraction method based on a question - answering model;

[0021] Figure 3 It is a schematic diagram of the third - party component document pre - processing method;

[0022] Figure 4Schematic diagram of the method for constructing the Q&A tree of the third-party component documentation

[0023] Figure 5 Schematic diagram of the method for extracting the usage rules of third-party components based on Q&A Specific implementation manner

[0024] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be noted that the following embodiments are intended to facilitate the understanding of the present invention and do not impose any limitations on it.

[0025] As Figure 1 shown, the fine-grained automated extraction system for third-party component documentation based on the Q&A model of the present invention includes a third-party component documentation preprocessing module, a documentation Q&A tree construction module, and a third-party component usage rule extraction module based on Q&A.

[0026] The working process of the entire fine-grained automated extraction system for third-party component documentation is as Figure 2 shown and includes the following steps:

[0027] (1) Collect the documentation of multiple different third-party components, such as the documentation of OpenSSL, SQLite, and uClibc, etc., preprocess the documentation to form a large-scale documentation repository. Use the attention model to refine the sentences of the to-be-tested documents in the documentation repository to obtain the coarse-grained usage rules of the third-party components;

[0028] (2) Design corresponding questions for the subsequent Q&A model according to the types of third-party component misuse; Select at least 5 documents from the to-be-tested documents in the documentation repository, and mark the artificial answers for the documents according to the designed questions;

[0029] (3) Divide the marked to-be-tested documents into a training set and a validation set, input the training set into the natural language processing model for training, and test the accuracy rate on the validation set. Use the trained model for the unmarked to-be-tested documents in the documentation repository to perform fine-grained mining on the usage rules of each involved third-party component.

[0030] In the present invention, the core of step (1) is to obtain the features of the third-party component usage rules, and use the attention model to roughly filter out the content irrelevant to the usage rules. Based on experience, even for documents with different writing styles, the content related to the usage rules contains an emphasized tone and will use modal mood words such as can (could), may (might), must, need, ought to, dare (dared), shall (should), will (would), etc. Therefore, in order to roughly filter out the irrelevant content in the third-party component documentation, the present invention proposes a filtering method based on the attention-based natural language processing model, which mainly includes:

[0031] (1-1) Collect the relevant documents of the third-party components recorded in the Linux help document. Filter out the third-party components with unclear or severely missing documents.

[0032] (1-2) According to the above characteristics of the third-party component usage rules, use the attention model to conduct coarse-grained filtering on the content of the third-party component documents, retain the sentences with an emphasized tone in the documents, and obtain the usage rules after coarse-grained filtering.

[0033] In the present invention, the core of step (2) is to deeply and manually analyze the misuse types of third-party components, design corresponding questions according to the misuse categories, and thus construct a question-and-answer tree, which mainly includes:

[0034] (2-1) Based on the publicly available third-party component misuse dataset, deeply and manually analyze the misuse types of third-party components. The obtained misuse types include four categories: misuse of obsolete functions, misuse of return values, misuse of call order, and misuse of parameters.

[0035] (2-2) Design corresponding query questions for each misuse type for the construction of the question-and-answer model. The query questions include seven categories: whether the function is obsolete, whether the function has a return value, in which cases the function has a return value, what the return values of the function are in the above cases, whether there are other functions that need to be called in advance, whether there are other functions that need to be called later, and what the parameter type is.

[0036] (2-3) According to the designed questions, select the to-be-tested documents in the document repository for manual marking, and construct a question-and-answer tree structure for each third-party component document.

[0037] In the present invention, the core of step (3) is to use the question-and-answer model to conduct fine-grained extraction of the usage rules. Based on experience, methods such as regular expressions cannot handle a large number of third-party component documents with loose structures and different writing styles. Therefore, in order to automatically process a large number of third-party component documents with different styles, the present invention proposes a fine-grained extraction method for third-party component documents based on the question-and-answer model, which mainly includes:

[0038] (3-1) Divide the to-be-tested documents marked in step (2-3) into a training set and a validation set according to a ratio of 8:2. Use the natural language processing model developed based on RoBERTa to perform iterative training on the training set and perform validation on the validation set until the loss function converges.

[0039] (3-2) Use the trained model to perform fine-grained extraction on the remaining third-party component documents to be tested in the document repository. The model takes the answer with the highest confidence probability for each question as the correct answer, and at the same time generates a question-and-answer tree for each document to be tested. According to the processing results of the question-and-answer model, fine-grained usage rules are extracted for each third-party component.

[0040] The following will explain each module separately:

[0041] 1. Third-party component document preprocessing module, which uses an attention model to perform coarse-grained filtering on the third-party component usage rules. As Figure 3 shown, the process is as follows:

[0042] First, according to the Linux help documents, collect the relevant documents of the third-party components recorded therein. Filter out the third-party components with unclear documentation or serious document deficiencies;

[0043] Then, according to the characteristics of the third-party component usage rules, use the attention model to perform coarse-grained filtering on the content of the third-party component documents, retain the sentences with an emphasized tone in the documents, and obtain the usage rules after coarse-grained filtering.

[0044] 2. Document question-and-answer tree construction module, which is used to obtain the misuse types of third-party components, design query questions for each type, and manually mark the relevant question answers for the documents to be tested, so as to construct a document question-and-answer tree. As Figure 4 shown, the process is as follows:

[0045] First, based on the publicly available third-party component misuse dataset, deeply and manually analyze the misuse types of third-party components. The obtained misuse types include: misuse of obsolete functions, misuse of return values, misuse of call order, and misuse of parameters;

[0046] Then, design corresponding query questions for each misuse type for the construction of the question-and-answer model. The query questions include: whether the function is obsolete, whether the function has a return value, in which cases the function has a return value, what the return values of the function are in the above cases, whether there are other functions that need to be called in advance, whether there are other functions that need to be called later, and what the parameter types are;

[0047] Finally, according to the designed questions, select the documents to be tested in the document repository for manual marking, and construct a question-and-answer tree structure for each third-party component document.

[0048] 3. Third-party component usage rule extraction module based on question-and-answer, which, based on the corresponding questions designed according to the third-party component misuse types, combines with a natural language processing model based on the RoBERTa model to obtain fine-grained usage rules related to the third-party components. As Figure 5 shown, the process is as follows:

[0049] First, divide the to-be-tested documents marked by the document Q&A tree construction module into a training set and a validation set according to the ratio of 8:2. Use the RoBERTa model to perform iterative training on the training set and validate on the validation set until the loss function converges;

[0050] Next, use the trained model to perform fine-grained extraction on the remaining to-be-tested third-party component documents in the document repository. The model takes the answer with the highest confidence probability for each question as the correct answer, and at the same time generates a Q&A tree for each to-be-tested document. According to the processing results of the Q&A model, generate the fine-grained usage rules for each third-party component.

[0051] The above-described embodiments have detailed the technical solutions and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, supplements, equivalent replacements, etc. made within the scope of the principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A fine-grained automated extraction method for third-party component documents based on a question-and-answer model, characterized in that, it includes the following steps: Step 1: Collect documents of multiple different third-party components, preprocess the documents, and build a document repository; use an attention model to refine the sentences of the documents to be tested in the document repository to obtain the coarse-grained usage rules of the third-party components; Step 2: Design corresponding questions for the question-and-answer model according to the misuse types of the third-party components; select some documents from the documents to be tested in the document repository and mark the answers according to the designed questions; The misuse types of the third-party components include obsolete function misuse, return value misuse, call order misuse, and parameter misuse; The corresponding questions of the question-and-answer model include: a. Whether the function is obsolete; b. Whether the function has a return value; c. In which cases does the function have a return value; d. What are the return values of the function in the above cases respectively; e. Whether there are other functions that need to be called in advance; f. Whether there are other functions that need to be called later; g. What is the parameter type; each question has optional answers; Step 3: Divide the marked documents to be tested into a training set and a validation set, use the training set to train a natural language processing model until the test accuracy of the validation set meets the preset requirements; use the trained natural language processing model to perform fine-grained mining on the coarse-grained usage rules of the remaining unmarked answer documents in the document repository.

2. The fine-grained automated extraction method for third-party component documents based on a question-and-answer model according to claim 1, characterized in that, when preprocessing the documents of the third-party components, filter out the third-party components with unclear descriptions or serious document deficiencies.

3. The fine-grained automated extraction method for third-party component documents based on a question-and-answer model according to claim 1, characterized in that, the working method of the attention model is: retain the sentences with emphasis words in the document and filter out the content irrelevant to the usage rules.

4. The fine-grained automated extraction method for third-party component documents based on a question-and-answer model according to claim 3, characterized in that, the emphasis words are at least one of can, could, may, might, must, need, ought to, dare, dared, shall, should, will, would.

5. The fine-grained automated extraction method for third-party component documents based on a question-and-answer model according to claim 1, characterized in that, the natural language processing model described in Step 3 adopts the RoBERTa model; during the training process, use the coarse-grained usage rules of the document as the input of the model, output the confidence of each answer corresponding to each question, and use the answer corresponding to the highest confidence as the result to generate a question-and-answer tree for each document.

6. A fine-grained automated extraction system for third-party component documents based on a question-and-answer model, used to implement the method described in claim 1, characterized in that, the extraction system includes: A third-party component document preprocessing module, which is used to collect documents of multiple different third-party components, preprocess the documents, and build a document repository; use an attention model to refine sentences of the documents to be tested in the document repository to obtain coarse-grained usage rules of the third-party components; A document Q&A tree construction module, which is used to design corresponding questions of a Q&A model according to the misuse types of third-party components; select some documents from the documents to be tested in the document repository and mark the answers according to the designed questions; A third-party component usage rule extraction module based on Q&A, which is used to divide the marked documents to be tested into a training set and a validation set, use the training set to train a natural language processing model until the test accuracy rate of the validation set meets the preset requirements; use the trained natural language processing model to perform fine-grained mining on the coarse-grained usage rules of the remaining unmarked answer documents in the document repository.

Citation Information

Patent Citations

  • Transaction type function point structured extraction method and system of software requirement document

    CN112817561A