A method for fusing test tool results based on text embedding
By building a standard rule set and using text embedding technology for rule mapping, the automatic analysis and integration of heterogeneous test data is solved, and efficient and accurate fusion and processing of test results are achieved.
Patent Information
- Application Number
- CN202411611444.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-12
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-11-12
AI Technical Summary
The heterogeneous test data generated by existing testing tools is difficult to automatically parse and integrate, resulting in complex and inefficient analysis and integration of test results.
Using a text embedding method, a standard rule set is constructed, and the results of different test tools are mapped to the standard rule set through multi-source polymorphic data analysis and rule mapping to achieve automatic resolution and integration of results.
It improves the standardization and consistency of test results, significantly improves the efficiency and accuracy of test results processing, reduces manual intervention, and enhances the intelligence and automation of test results.
Smart Images

Figure CN119829416B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of software testing, and particularly relates to a method for fusing test tool results based on text embedding. Background Art
[0002] With the rapid development of software development, the quality and security of software code have become the focus of attention for enterprises and developers. To ensure the reliability of software products, developers usually use a variety of automated test tools, such as static code analysis tools (such as klockwork, Testbed, Checkmarx, Fortify, etc.) and dynamic test tools (such as Appscan). However, due to the lack of a unified standard, the test result formats generated by these tools are diverse, and at the same time, there are gaps in the capabilities of each tool when detecting different rules and different codes. It is necessary to confirm all the different test results generated by different tools, resulting in complex and inefficient analysis and integration of multi-source test data. The existing test process usually relies on manual analysis of the test results of different tools, which is not only time-consuming and laborious, but also prone to errors such as omissions in the manual confirmation of results, resulting in the failure to timely discover and repair some potential software defects or security vulnerabilities. To solve these problems, there is an urgent need for an intelligent method that can automatically parse and integrate heterogeneous test results to improve the efficiency and accuracy of test result processing and reduce the workload of developers.
[0003] Existing test result fusion methods often rely on manual configuration of rules and cannot flexibly handle changes in new tools and new rules. In addition, due to the heterogeneity of test tools, there may be significant semantic differences in the results of each tool, and traditional rule matching methods are difficult to accurately capture these differences. Therefore, there is an urgent need for an intelligent fusion method based on deep learning that can automatically parse, map, and fuse the results from different test tools, reduce manual intervention, and improve the consistency and automation level of test result processing. Summary of the Invention
[0004] (1) Technical Problems to be Solved
[0005] The technical problem to be solved by the present invention is how to provide a method for fusing test tool results based on text embedding to solve the problem of automatic parsing and integration of heterogeneous test data generated by multiple automated test tools.
[0006] (2) Technical Solutions
[0007] To solve the above technical problems, the present invention proposes a method for fusing test tool results based on text embedding, and the method includes the following steps:
[0008] S1. Construct a standard rule set: Construct a standard rule set R0 based on national and military standards. This rule set contains several core rules for subsequent mapping and fusion processes.
[0009] S2. Parse multi-source and polymorphic data: Different testing tools have their own rule sets, which may adopt different formats and standards. First, obtain the rule sets of different testing tools through manual extraction. Through multi-source and polymorphic data parsing, convert the multi-source and polymorphic rule sets into intermediate rule sets with the same rules as those in the standard rule set R0.
[0010] S3. Rule mapping: Map the intermediate rule sets of different testing tools to the standard rule set R0 through a text embedding model.
[0011] S4. Result fusion: Through predefined rule mapping, dock the output results of each testing tool with the standard rule set to ensure the accurate aggregation of the output results. In this process, all results mapped to the same standard rule will be processed uniformly to eliminate differences between different tools and achieve the consistency and comparability of test results. Finally, generate a complete and standardized test report.
[0012] (3) Beneficial effects
[0013] The present invention proposes a method for fusing test tool results based on text embedding. The method for fusing test tool results based on text embedding of the present invention has the following characteristics:
[0014] First, by constructing the standard rule set R0, the unity and standardization of the rules are ensured, enabling the effective fusion of test tool results from different sources through a standardized framework.
[0015] Second, the introduction of multi-source and polymorphic data parsing can handle various data formats, achieving flexibility and adaptability and reducing compatibility problems caused by data format differences.
[0016] In addition, rule mapping is proposed, which utilizes text embedding technology and cosine similarity to improve the accuracy and efficiency of rule mapping and ensure the intelligence and automation of the fusion process. These series of advantages make this method have significant advantages in realizing the efficient integration of test results and improving test quality.
[0017] Through the above characteristics, the method for fusing test tool results based on text embedding of the present invention not only improves the standardization and consistency of software test results, but also significantly improves the efficiency and accuracy of test result processing through an intelligent processing mechanism, providing an innovative solution for the software testing field. Description of the drawings
[0018] Figure 1It is the architecture diagram of the method described in the present invention;
[0019] Figure 2 It is the flow chart of the method described in the present invention;
[0020] Figure 3 It is the rule mapping diagram in the method described in the present invention;
[0021] Figure 4 It is the text embedding technology based on the CLIP pre-trained model in the method described in the present invention. Detailed implementation manners
[0022] To make the objectives, contents and advantages of the present invention clearer, the following further describes in detail the specific implementation manners of the present invention with reference to the drawings and embodiments.
[0023] The present invention relates to the field of software testing, and particularly to an intelligent fusion method for test tool results based on text embedding, aiming to improve the efficiency of software testing and the accuracy of result fusion.
[0024] The present invention provides a method for fusing test tool results based on text embedding, aiming to solve the problem of automatic parsing and integration of heterogeneous test data generated by multiple automated test tools.
[0025] This method constructs a benchmark rule set R0 based on national standards and military standards, including several core rules R01, R02, etc. It uses the CLIP pre-trained model to generate text vectors for each rule description in the rule set and stores them in the vector database. Through multi-source and polymorphic data parsing, the rule sets of different test tools are converted into a unified format identical to the standard rule set to obtain the rules and rule descriptions in the results. This method also proposes rule mapping, mapping the rules of the test results to the rules of the standard rule set R0. Specifically, it converts the rule description of the test results into vectors, matches them with the rules of the existing standard rule set R0 through vector cosine similarity. For the matched rules, the rules with a significantly different number of detected problems from other tools are manually verified, and the rest are fused into the standard set. For the unmatched rules, they are extended to the standard set through manual verification.
[0026] The present invention significantly improves the efficiency and accuracy of test result processing, helps developers quickly locate and solve security vulnerabilities and quality problems in software, and provides an intelligent test result fusion solution.
[0027] Embodiment 1:
[0028] Regarding the method for fusing test tool results based on text embedding, its architecture diagram is as Figure 1 shown, and the flow chart is as Figure 2As shown, the implementation steps include the following key parts:
[0029] S1. Construct a standard rule set:
[0030] Constructing a standard rule set is a key step in achieving the integration of the results of different testing tools. To ensure the accuracy and consistency of the integration, a benchmark rule set R0 was first constructed based on relevant national standards and industry specifications. This rule set contains a number of core rules that cover various common vulnerability detection and security audit items. The definition of each rule has been strictly verified to ensure that its content conforms to the provisions of national and military standards and has sufficient applicability, in the form of:
[0031]
[0032] During the construction process, not only will the specific content of each rule be described in detail, but also its applicable scope, application conditions, and test environment requirements will be clearly defined. At the same time, when constructing the standard rule set, future expansion requirements will also be considered, leaving room for rule expansion so that the rule set can support emerging testing tools and new security standards.
[0033] S2. Parse multi-source and polymorphic data:
[0034] The rule set formats of different testing tools may be different, and these formats may include XML, CSV, JSON, or log file formats, etc., in the form of:
[0035]
[0036] To achieve seamless docking between heterogeneous data sources generated by different testing tools, an intermediate rule format was constructed, and its core function is to map source data in various formats - including XML, CSV, JSON, or log files, etc. - to a unified intermediate rule format.
[0037] First, identify the format of the input data, which is the starting point of the mapping process. Whether it is XML, CSV, or JSON, it can be accurately identified and the corresponding processing flow can be prepared. With a plug-in architecture, it is possible to flexibly access and parse various different formats of test data. For common data formats such as XML and CSV, standard parsing plug-ins are built-in. For special or customized data formats, developers can develop new plug-ins to extend the support for new formats. After identifying and parsing the source data, these data are converted into a standardized intermediate rule format, which ensures that data from different tools can be understood and processed within a unified framework. This format facilitates subsequent processing and fusion because it provides a common and standardized representation for all source data, while retaining the rule ID and rule description fields of the source data for more accurate mapping in subsequent steps. For example:
[0038]
[0039] This format facilitates subsequent processing and fusion because it provides a common and standardized representation for all source data. By mapping the source data to the intermediate rule format, the consistency and compatibility of data from different test tools during further processing are ensured. This not only improves the efficiency of data processing but also enhances the quality of data fusion. It enables the efficient integration of data from different test tools, laying a solid foundation for subsequent analysis and processing.
[0040] S3. Rule mapping: Rule mapping is the core link of the method of the present invention. As Figure 3 shown, its purpose is to map the rules generated by different test tools into the rules of a unified standard rule set R0. This process ensures that test results from different sources can be compared and analyzed under the same standard, thereby improving the consistency and comparability of test results.
[0041] After the multi-source and multi-polymorphic data parsing is completed and converted into a standardized format, an intelligent mapping algorithm is used to map the rule descriptions generated by different tools to the rule descriptions in the standard rule set R0 by vectorizing the rule descriptions in the rules generated by different tools. The core of the intelligent mapping algorithm lies in the application of text embedding technology. By introducing text embedding technology based on the CLIP pre-trained model, as Figure 4 shown, the CLIP pre-trained model has been widely trained and has powerful text understanding and embedding capabilities. The algorithm can convert each rule description into a high-dimensional vector representation Text embedding technology enables the capture of semantic features of rule descriptions and vectorization processing without relying on a specific format.
[0042] During the mapping process, the cosine similarity of the vectors generated by each rule is first calculated and compared with the rule vectors in the standard rule set. In this way, mapping rules can be automatically generated based on the semantic similarity scores between rule descriptions. To ensure the accuracy of mapping, the rule with the highest similarity score is preferentially selected for mapping, and a similarity threshold w is set s , to ensure the quality of the mapping results. For rules below the threshold that still need to be mapped, a manual verification function is provided to ensure that the mapping of these rules can accurately reflect the differences and commonalities between different tools.
[0043] Vector retrieval mechanism: To further improve the processing efficiency and expansion ability, an efficient vector retrieval mechanism is constructed based on text embedding. Each time a new rule set is parsed and vectors are generated, these vectors are stored in the Milvus vector database. The vector database has the advantages of fast retrieval and high-dimensional space similarity calculation, and can support the efficient storage and real-time retrieval of large-scale rule sets. When mapping a new rule set, the most similar standard rule set vectors are first found by retrieving the stored Milvus vector database. This vector retrieval-based mechanism not only improves the speed of rule mapping but also ensures the efficient integration of results from different tools. By combining text vectorization and vector retrieval, a large amount of rule data can be processed in a short time, greatly improving the processing efficiency and adapting to application scenarios with large-scale data sets and complex rule sets.
[0044] Similarity calculation and mapping rule generation: In the implementation of the intelligent mapping algorithm, the calculation of similarity plays a crucial role. First, we use the text embedding technology pre-trained by the CLIP model to convert each rule description into a vector form in a high-dimensional space. Subsequently, the system calculates the cosine similarity between these newly generated rule vectors and the vectors stored in the standard rule set. This process enables us to quantitatively evaluate the similarity between different rules. The higher the similarity between rules, the closer their matching degree. The system preferentially selects the rule with the highest similarity for mapping, and this step is to automatically create mapping rules.
[0045] For rules whose similarity exceeds the preset threshold, the system will automatically create mapping rules and store these rules in the mapping rule table for subsequent calls and applications. For rules whose similarity does not reach the threshold but still need to be mapped, the system will mark them as needing manual review to ensure the accuracy and consistency of the mapping results. This method combining similarity calculation and mapping rule generation not only improves the efficiency of rule mapping but also ensures the quality of the integration of results from different test tools. In this way, we can ensure the accuracy and efficiency of rule mapping while maintaining the consistency and reliability of the rule set.
[0046] Mapping rule verification: To ensure the accuracy of mapping rules, a mechanism combining automatic and manual verification is designed. After the automatic generation of mapping rules, automatic verification of logical consistency and data consistency is first carried out to ensure that each mapping rule meets the expectations. At the same time, for some mapping rules with low similarity or large data differences, they will be automatically marked and notified to manual verification personnel for inspection. For example, if a certain rule R21 is mapped to the standard rule R01, but the number of vulnerabilities or other attributes are significantly different from those of similar rules from other testing tools, the manual verification personnel will be prompted to review and adjust this mapping. In the manual verification stage, experts can adjust the mapping rules according to the actual situation to ensure that the generated mapping results can accurately reflect the actual situations of different tools. Through this combined automatic and manual verification mechanism, the accuracy of rule mapping is ensured, and data inconsistency problems caused by rule differences are effectively avoided.
[0047] S4. Result fusion:
[0048] Result fusion docks the output results of each testing tool with the standard rule set through predefined rule mapping to ensure the accurate summary of the output results. In this process, all results mapped to the same standard rule will be processed uniformly, eliminating the differences between different tools and achieving the consistency and comparability of test results. The system stores the summarized results in a unified data structure for subsequent data analysis and visual display, ensuring the transparency and usability of test results, thereby improving the overall efficiency and effectiveness of software testing.
[0049] Embodiment 2:
[0050] A method for fusing test tool results based on text embedding, including:
[0051] S1. Construct a standard rule set: Construct a standard rule set R0 based on national and military standards. This rule set contains several core rules R01, R02, etc., for subsequent mapping and fusion processes;
[0052] S2. Parse multi-source and polymorphic data: Different testing tools have their own rule sets ( etc.), and these rule sets may adopt different formats and standards. First, obtain the rule sets of different testing tools through manual extraction ( etc.), and through parsing multi-source and polymorphic data, convert the multi-source and polymorphic rule sets into intermediate rule sets (R1, R2, etc.) with the same format as the rules in the standard rule set R0;
[0053] S3. Rule mapping: Map the intermediate rule sets (R1, R2, etc.) of different testing tools to the standard rule set R0 through a text embedding model.
[0054] S4, Result Fusion: Through predefined rule mapping, the output results of each test tool are docked with the standard rule set to ensure the accurate aggregation of the output results. During this process, all results mapped to the same standard rule are processed uniformly, eliminating differences between different tools and achieving the consistency and comparability of test results. Finally, a complete and standardized test report is generated.
[0055] Furthermore, the construction of the standard rule set includes:
[0056] Rule Definition: Clearly define the specific content of each rule such as R01, R02, etc., in the form of:
[0057]
[0058]
[0059] Rule Verification: Manually verify each rule to ensure it complies with national and military standards.
[0060] Furthermore, the construction of the standard rule set also includes:
[0061] Vector Generation: Use text embedding technology to generate text vectors corresponding to the descriptions of rules R01, R02, …, R0 in the rule set; n
[0062] Establish Retrieval: Store the rule IDs and the corresponding generated text vectors in the Milvus vector database for efficient retrieval of vector data;
[0063] Provide Extension: The rules of different test tools may contain rules outside the standard rule set. Therefore, the constructed rule set R0 also includes extension rules R0 e0 , R0 e1 , …, R0 em , facilitating rule extension.
[0064] Furthermore, the multi-source and polymorphic data parsing includes:
[0065] Heterogeneous Data Conversion: The rule sets of different test tools extracted manually ( etc.) belong to multi-source, polymorphic, and heterogeneous data, including different formats such as CSV, XML, JSON, etc. Step S2 is responsible for converting the rule sets of different test tools ( etc.) into the same format as the standard rule set as the intermediate rule set (R1, R2, etc.) for subsequent processing; Assume that the code scanning results generated by the test tool are in the following format:
[0066]
[0067]
[0068] For this data format, neither the internal key-value nor the rule description is determined. The intermediate format is to uniformly convert the results generated by different testing tools into the same format. The intermediate rule set after conversion is in the following format:
[0069]
[0070] Furthermore, the text embedding model in step S3 is a CLIP pre-trained model, and the CLIP pre-trained model is used to generate text vectors of test data.
[0071] Furthermore, it includes the following steps:
[0072] Similarity calculation: For the rule descriptions in the intermediate rule sets (R1, R2, etc.) of different testing tools, use CLIP to calculate the text vectors of the rule descriptions, and calculate the similarity between the rule description vectors of different rules in the R0 rule set.
[0073] Rule mapping implementation: According to the similarity calculation results, select the rule with the highest similarity and map it to the standard rule set R0;
[0074] Rule mapping verification: There will be different situations when mapping the rules to the standard rule set R0. For some situations, manual verification is performed.
[0075] Furthermore, the similarity calculation preferably uses cosine similarity to match the test results with the rules to ensure the accuracy of rule mapping.
[0076] Furthermore, step S3 also includes:
[0077] Fusion into the standard set: If the rule description of a certain rule, such as R21, can be mapped to R01 through an algorithm, it will be mapped to R01 and fused into the standard set;
[0078] Expansion into the standard set: The rule description of a certain rule, such as R31, has an algorithm similarity lower than the set similarity threshold w s , and after manual verification, it will be expanded into the standard set.
[0079] Furthermore, step S3 also includes:
[0080] If the rule description of a certain rule, such as R22, can be mapped to R02 through an algorithm, but the specific number of problems in the rule is too different from the rules that can be mapped to R02 by other testing tools, manual verification will be performed and mapping will be carried out.
[0081] Further, in step S4, through predefined rule mapping, the output results of each test tool are docked with the standard rule set to ensure accurate result aggregation. All results mapped to the same standard rule are uniformly processed to eliminate differences between tools, achieve consistency and comparability of test results, and finally generate a standardized test report. The integrated data is stored in a unified structure for subsequent analysis and display, improving the visualization effect and usability of the results.
[0082] The present invention proposes an intelligent fusion method for test tool detection results based on text embedding, belonging to the field of software testing. The method includes four parts: constructing a standard rule set, parsing multi-source and polymorphic data, rule mapping, and result fusion. First, based on standards such as national standards and military standards or independently establish a standard rule set R0. Through the text embedding algorithm, the rule descriptions of the rules (R01, R02, etc.) in the rule set are vectorized and stored in a vector database. Manually extract the multi-source and polymorphic rule sets of different test tools ( etc.). The multi-source and polymorphic data parsing is responsible for converting these rule sets into intermediate rule sets (R1, R2, etc.) with the same format as the standard rule set R0. The rule mapping maps the intermediate rule sets R1, R2 of the test tool to the standard rule set R0. Finally, the result fusion ensures that the output results of the test tool can be accurately docked to the standard rule set according to the established rule mapping mechanism, and effectively aggregates the detection results of multiple tools, thereby achieving consistency and comparability of the detection results. This method can significantly improve the efficiency and accuracy of software testing, reduce the differences in test results caused by inconsistent standards, and provide a new solution for the field of software testing.
[0083] The result fusion method of the test tool based on text embedding of the present invention has the following characteristics:
[0084] First, by constructing the standard rule set R0, the unity and standardization of the rules are ensured, so that the test tool results from different sources can be effectively fused through a standardized framework.
[0085] Second, the introduction of multi-source and polymorphic data parsing can process various data formats, achieve flexibility and adaptability, and reduce compatibility problems caused by data format differences.
[0086] In addition, the rule mapping is proposed. Its implementation utilizes text embedding technology and cosine similarity, improving the accuracy and efficiency of rule mapping, and ensuring the intelligence and automation of the fusion process. These series of advantages make this method have significant advantages in realizing the efficient integration of test results and improving the test quality.
[0087] Through the above characteristics, the method for fusing the results of the test tool based on text embedding of the present invention not only improves the standardization and consistency of software test results, but also significantly enhances the efficiency and accuracy of test result processing through an intelligent processing mechanism, providing an innovative solution for the field of software testing.
[0088] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A test tool result fusion method based on text embedding, characterized in that: The method comprises the following steps: S1. Build a standard rule set: Build a standard rule set based on national standards and national military standards R 0, this rule set contains several core rules, which are used in the subsequent mapping and fusion process; S2. Multi-source and multi-state data analysis: Different testing tools have their own rule sets, which may use different formats and standards. First, the rule sets of different testing tools are obtained through manual extraction, and then the multi-source and multi-state rule sets are converted into format and standard rule sets through multi-source and multi-state data analysis. R The intermediate rule set is the same as the rules in 0; S3, rule mapping: Mapping the intermediate rule sets of different testing tools to the standard rule set through the text embedding model R 0 in; S4, Result Fusion: Through predefined rule mapping, the output results of each test tool are connected with the standard rule set to ensure accurate summary of the output results; In this process, all results mapped to the same standard rules will be processed uniformly to eliminate differences between different tools, achieve consistency and comparability of test results, and ultimately generate a complete and standardized test report; in, Said S1 also includes: vector generation, search establishment and extension provision; Vector generation: Use text embedding technology to generate text vectors corresponding to the descriptions of the rules in the rule set; Establish search: store the rule ID and the corresponding generated text vector into the Milvus vector database to facilitate efficient retrieval of vector data; Provide extensions: The rules of different testing tools may contain rules outside the standard rule set, so the constructed rule set R 0 also contains extension rules to facilitate rule expansion; The S3 includes: Similarity calculation: The rule descriptions in the intermediate rule sets of different testing tools are calculated using CLIP to calculate the text vectors of the rule descriptions and the similarity with the rule set. R Similarity of rule description vectors of different rules in 0; Rule mapping implementation: Based on the similarity calculation results, select the rule with the highest similarity and map it to the standard rule set R 0 in; Rule mapping verification: rules are mapped to standard rule sets R There will be different situations in 0, and for some situations, manual verification is performed.
2. The test tool result fusion method based on text embedding according to claim 1, characterized in that: The S1 includes: rule definition and rule verification. The rule definition is used to clearly define the specific content of each rule, and the rule verification is used to manually verify each rule to ensure that it complies with national standards and national military standards.
3. The test tool result fusion method based on text embedding according to any one of claims 1-2, characterized in that: The S2 includes: the rule sets of different testing tools extracted manually belong to multi-source polymorphic heterogeneous data, and contain different formats, including: CSV, XML and JSON. Step S2 is responsible for converting the rule sets of different testing tools into the same format as the standard rule set as an intermediate rule set for subsequent processing.
4. The test tool result fusion method based on text embedding according to claim 3, characterized in that: The text embedding model of S3 is a CLIP pre-trained model, and the CLIP pre-trained model is used to generate text vectors for test data.
5. The test tool result fusion method based on text embedding according to claim 4, characterized in that: The similarity calculation uses cosine similarity to match the test result with the rule to ensure the accuracy of rule mapping.
6. The test tool result fusion method based on text embedding according to claim 4, characterized in that: The step S3 further comprises: Fusion into the standard set: If the rule description of a rule can be mapped to the standard rule set through the algorithm, it will be mapped and merged into the standard set; Expand to the standard set: The rule description of a rule is lower than the set similarity threshold after the algorithm. w s , and after manual verification will be expanded to the standard set.
7. The test tool result fusion method based on text embedding according to claim 6, characterized in that: The S3 also includes: if the rule description of a rule can be mapped to a standard rule through an algorithm, but the number of specific questions in the rule is too different from the number of questions that can be mapped to the standard rule by other testing tools, manual verification and mapping will be performed.
8. The test tool result fusion method based on text embedding according to claim 4, characterized in that: The S4 integrated data is stored in a unified structure, which is convenient for subsequent analysis and display, and improves the visualization effect and use value of the results.
Citation Information
Patent Citations
Static report consolidation analysis techniques
CN111367789A
Merging analysis method and system based on static report and medium
CN113742214A