Method for establishing open source component license index based on assistance of large model RAG

The establishment of open source component license indexes through large-scale RAG assists in establishing open source component license indexes, solving the problems of network dependence and incomplete information in the existing technology, and achieving efficient and accurate construction of license index libraries in any environment, reducing legal risks.

CN120010856APending Publication Date: 2025-05-16BEIJING SIMPLE POINT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510111417.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When analyzing open source component licenses, the existing technology has problems such as strong network dependency, poor flexibility and incomplete information, which leads to the inability to fully and accurately establish a license index library, increasing legal risks.

Method used

The large-scale model RAG is used to assist in establishing open source component license indexes, and vectorized storage is carried out by combining the open source license knowledge base, and using the large-scale model for knowledge enhancement, processing non-standard license descriptions, and combining component warehouse interface query and manual verification to build a standard license index library.

Benefits of technology

It realizes efficient and accurate construction of open source component license index library in any network environment, improves the coverage and accuracy of license risk analysis, and reduces legal risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010856A_ABST
    Figure CN120010856A_ABST
Patent Text Reader

Abstract

The invention discloses a method for assisting in establishing an open source component license index based on a large model RAG. The method comprises the following steps: step 1, carrying out large model RAG in combination with an open source license knowledge base; step 2, constructing an open source component license index database; according to the method, knowledge enhancement is performed on the large model based on the large model RAG technology in combination with the official license knowledge base, and non-standard license description information of the open source component is further processed in combination with the large model. The obtained license name meeting the standard is used for establishing an open source component license index database, and compared with a license alias limited enumeration or keyword matching technology, the license information is corrected; according to the method, the coverage range of the license of the index database component is further expanded, a more comprehensive license index database of the open source component is established, and the component coverage rate and accuracy of risk analysis of the license of the open source component are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of LLM large language model and SCA software supply chain security, and in particular to a method for establishing an open source component license index based on a large model RAG. Background Art

[0002] Most software will reference various open source components during the development process, which can reduce a lot of basic development work. Developers can focus on developing business-related code, which can improve software development efficiency. However, there is another problem: these introduced open source libraries will have different open source protocols, and these open source protocols generally have some claims. If the project does not meet the claims of the open source component license, it may face some legal risks. For example, if a project references an open source component of the PGL protocol, then the project itself must also follow the same protocol for open source, and our project characteristics may not meet such requirements, which will cause risks. In order to avoid this risk, it is necessary to understand the licenses and risks used by each open source component referenced in the project;

[0003] The methods for analyzing the open source component licenses of existing projects are as follows:

[0004] Based on the project's software bill of materials (SBOM), analyze the license information for each component:

[0005] Method 1: Query the component license information through the open interfaces of various open source component repositories;

[0006] Method 2: Combine the project compilation environment, download the dependent components during the compilation process, and obtain the component license information by parsing the license feature file;

[0007] Method 3: Establish a license information library for open source components, and directly query the established license information library to obtain the license information of the components when performing license risk analysis. The specific method of establishing the information library is basically based on 1 or 2. The original license description information of the component is obtained by interface query or feature file parsing, and then the non-standard license information is processed through limited enumeration mapping, keyword matching, manual correction, etc. to obtain the standard license name to establish the component license information library.

[0008] The disadvantages of the above methods 1-3 are as follows:

[0009] Method 1: The analysis process depends on the network environment. The official component repository is requested to query the component license information in real time. Normal analysis cannot be performed in an environment where the external network cannot be accessed.

[0010] Method 2: During the analysis process, you need to combine the project compilation environment to obtain the components that the project depends on and the license information of the components. This method requires relying on the compilation environment and downloading components, which is not lightweight and flexible enough.

[0011] Method 3: Although an open source component license information library can be established in advance based on Method 1 and Method 2, the component license information library is not comprehensive and the license information is not standardized. Because the license information queried through the official interface of the component repository is likely not a standard license name, it may be just a text description of the license, or a non-standard license name. Since it is filled in manually and there are no clear specifications and constraints, there are many possibilities. Or the license field returned by the interface does not contain license information, but the license information appears in other fields. At this time, the obtained license information needs to be processed by a program and converted into a standard license. For example, the non-standard description Apache License version 2, Apache License 2.0 or a section of the license agreement content obtained from the interface is mapped to the standard license Apache-2.0, such as Figure 3 As shown, Figure 3 This is the information of an open source component queried by the python pypi interface, for reference: the license field returned by the interface does not contain license information, but the classifiers field contains license information. Although it contains license information, it is not the standard license name GPL-2.0, and cannot be directly used to build a license index library, nor can it be used for license risk analysis. Summary of the invention

[0012] The purpose of the present invention is to provide a method for establishing an open source component license index based on a large model RAG to solve the problems raised in the above background technology.

[0013] To achieve the above object, the present invention provides the following technical solutions:

[0014] The method for establishing an open source component license index based on the large model RAG is as follows:

[0015] Step 1: Combine the open source license knowledge base to conduct large model RAG;

[0016] Step 2: Build an open source component license index library;

[0017] Step 3: Store the open source component license index library.

[0018] As a further solution of the present invention: Step 1 is specifically to organize the open source license information based on the official SPDX and OSI to obtain an open source license knowledge base, cut each open source license content and perform vectorized storage, and enhance the open source license knowledge of the general large model.

[0019] As a further solution of the present invention: the steps of the method for constructing the open source component license index library in step 2 are as follows:

[0020] S1: Based on the official interface of the component repository, query the list of released components and the list of recently updated components, and query and obtain detailed information of the components in turn;

[0021] S2: Based on the component information obtained in S1, parse the field content that may contain license description information. If the license description of the component is already a standard license name, it can be directly used for the subsequent construction of the license index library;

[0022] S3: If it is a non-standard license name, it is further converted into a standard license name based on the enumeration of license aliases or keyword matching, which is used for the subsequent construction of the license index library;

[0023] S4: If S3 still fails to resolve the standard license name, then based on the contents that may contain the license description obtained in S2, assemble them into prompt words to further request the knowledge-enhanced large model, and let the large model return a standard license name that best matches the license description. If there is no best match, return empty without further elaboration. The obtained standard license name is used for the subsequent construction of the license index library;

[0024] S5: If S1-S4 fail to resolve the standard license name of the component, the component is included in the list that requires manual verification, and the license information is supplemented after subsequent manual confirmation.

[0025] As a further solution of the present invention: the open source component license index library storage method in step three is specifically to establish a map format index file based on the component information obtained in step two and the standard license information of the component. When performing component license risk analysis, all components and license information that meet the prefix of the component name are retrieved from the index file, and the license list of the component is further retrieved for subsequent license risk analysis.

[0026] As a further solution of the present invention: the component repository in S1 includes but is not limited to maven, npm, pypi, and go components.

[0027] As a further solution of the present invention: the standard license names in S2 include Apache-1.0, Apache-2.0, MIT, and GPL-2.0.

[0028] As a further solution of the present invention: the data structure in step three is in the form of: {"a":{"a1":["Apache-2.0"], "a2":["Apache 2.0", "MIT"]}, ...}, and is stored in a binary file.

[0029] Compared with the prior art, the present invention has the following beneficial effects:

[0030] 1. The present invention enhances the knowledge of the big model based on the big model RAG technology combined with the official license knowledge base, and further processes the non-standard license description information of the open source component in combination with the big model to obtain the license name that meets the standard for establishing the open source component license index library, compared with only correcting the license information through limited enumeration of license aliases or keyword matching technology;

[0031] 2. The present invention further improves the coverage of index library component licenses, establishes a more comprehensive license index library for open source components, and improves component coverage and accuracy of open source component license risk analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 Module relationship diagram of the method for establishing an open source component license index based on the large model RAG.

[0033] Figure 2 The effect diagram of building the open source component license index library in the method of establishing the open source component license index based on the large model RAG is presented.

[0034] Figure 3 This is an information diagram of an open source component queried from the Python PyPI interface. DETAILED DESCRIPTION

[0035] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0036] See also Figures 1 to 3 In an embodiment of the present invention, a method for establishing an open source component license index based on a large model RAG is as follows:

[0037] Step 1: Combine the open source license knowledge base to conduct large model RAG;

[0038] Step 2: Build an open source component license index library;

[0039] Step 3: Open source component license index library storage;

[0040] The step 1 specifically includes: based on the open source license information maintained by the official organizations such as SPDX (https: / / spdx.org / licenses / ) and OSI (https: / / opensource.org / license), an open source license knowledge base is obtained, and each open source license content is cut and vectorized for storage, and the open source license knowledge is enhanced for the general large model;

[0041] The steps of the method for constructing the open source component license index library in step 2 are as follows:

[0042] S1: Based on component repositories, including but not limited to the official interfaces of common maven, npm, pypi, go and other components, query the list of released components and regularly query the list of recently updated components, and query and obtain component detailed information in sequence, usually in json or xml format;

[0043] S2: Based on the component information obtained in S1, parse the field content that may contain license description information. If the license description of the component is already a standard license name, such as Apache-1.0, Apache-2.0, MIT, GPL-2.0, etc., no further processing is required and it can be directly used for the subsequent construction of the license index library;

[0044] S3: If it is a non-standard license name, it is further converted into a standard license name based on the enumeration of license aliases or keyword matching, which is used for the subsequent construction of the license index library;

[0045] S4: If S3 still fails to resolve the standard license name, then based on the content that may contain the license description obtained in S2, assemble them into prompt words to further request the knowledge-enhanced large model, so that the large model returns a standard license name that best matches the license description. If there is no best match, it returns nothing without further explanation. The effect is as follows: Figure 2 As shown, the obtained standard license name is used for the subsequent construction of the license index library;

[0046] S5: If S1-S4 fail to resolve the standard license name of the component, the component is included in the list that requires manual verification, and the license information is supplemented after subsequent manual confirmation.

[0047] The open source component license index library storage method of step three is specifically to establish a map format index file based on the component information obtained in step two and the standard license information of the component, and the data structure is as follows: {"a":{"a1":["Apache-2.0"], "a2":["Apache 2.0", "MIT"]},......}, and store it in a binary file. When performing component license risk analysis, all components and license information that match the prefix of the component name are retrieved from the index file, and the license list of the component is further retrieved for subsequent license risk analysis.

[0048] Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments, or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for establishing an open source component license index based on a large model RAG, characterized in that: The method steps are as follows: Step 1: Combine the open source license knowledge base to conduct large model RAG; Step 2: Build an open source component license index library; Step 3: Store the open source component license index library.

2. The method for establishing an open source component license index based on a large model RAG according to claim 1, characterized in that: The step one specifically includes organizing the open source license information based on the official SPDX and OSI maintained open source license information to obtain an open source license knowledge base, cutting and vectorizing the content of each open source license, and enhancing the open source license knowledge of the general large model.

3. The method for establishing an open source component license index based on a large model RAG according to claim 1, characterized in that: The steps of the method for constructing the open source component license index library in step 2 are as follows: S1: Based on the official interface of the component repository, query the list of released components and the list of recently updated components, and query and obtain detailed information of the components in turn; S2: Based on the component information obtained in S1, parse the field content that may contain license description information. If the license description of the component is already a standard license name, it can be directly used for the subsequent construction of the license index library; S3: If it is a non-standard license name, it is further converted into a standard license name based on the enumeration of license aliases or keyword matching, which is used for the subsequent construction of the license index library; S4: If S3 still fails to resolve the standard license name, then based on the contents that may contain the license description obtained in S2, assemble them into prompt words to further request the knowledge-enhanced large model, and let the large model return a standard license name that best matches the license description. If there is no best match, return empty without further elaboration. The obtained standard license name is used for the subsequent construction of the license index library; S5: If S1-S4 fail to resolve the standard license name of the component, the component is included in the list that requires manual verification, and the license information is supplemented after subsequent manual confirmation.

4. The method for establishing an open source component license index based on a large model RAG according to claim 1, characterized in that: The open source component license index library storage method of step three is specifically to establish a map format index file based on the component information obtained in step two and the standard license information of the component. When performing component license risk analysis, all components and license information that match the prefix of the component name are retrieved from the index file, and the license list of the component is further retrieved for subsequent license risk analysis.

5. The method for establishing an open source component license index based on a large model RAG as claimed in claim 3, characterized in that: The component repositories in S1 include but are not limited to maven, npm, pypi, and go components.

6. The method for establishing an open source component license index based on a large model RAG as claimed in claim 3, characterized in that: The standard license names in S2 include Apache-1.0, Apache-2.0, MIT, GPL-2.

0.

7. The method for establishing an open source component license index based on a large model RAG as claimed in claim 4, characterized in that: The data structure in step three is in the form of: {"a":{"a1":["Apache-2.0"],"a2":["Apache2.0","MIT"]},......}, and is stored in a binary file.