Open source malicious code security intelligence collection method and system based on large model and medium
Through the open source malicious code security intelligence collection method based on large-scale models, the problem of insufficient attention to malicious code intelligence in the open source package manager is solved, and timely identification and response to package manager security threats is realized, improving the overall security of the open source ecosystem.
Patent Information
- Application Number
- CN202510209939.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-17
AI Technical Summary
The prior art has relatively insufficient attention to malicious code-specific intelligence in open source package managers such as NPM and PyPI, resulting in malicious code that may lurk in various mirroring and downstream projects for a long time, posing a persistent security threat.
The open source malicious code security intelligence collection method based on large models is adopted, and the recursive expansion process driven by trusted sources is used to guide the search direction using high-quality initial keywords, expand the information coverage through link recursive access, and finally establish a comprehensive package manager security intelligence library. The specific steps include identifying the very nouns in the intelligence source, extracting SSC entity information through the LLM model, performing consistency analysis and data aggregation.
By promptly identifying and responding to security threats in the package manager, mitigating risks, improving network security situation, improving the speed and efficiency of response to existing threats, and achieving a more comprehensive protection of the security of the open source ecosystem.
Smart Images

Figure CN120162397A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information processing, and particularly relates to an open-source malicious code security intelligence collection method, system and medium based on a large model. Background Art
[0002] With the wide application of open-source software globally, its security threats have become increasingly prominent, especially in third-party library registries such as NPM and PyPI. These platforms are often subject to the upload and potential penetration of malicious code. As shown in the event-stream incident in 2017, a widely used open-source package was implanted with malicious code, affecting millions of users. In response to this problem, researchers in academia and industry have carried out extensive research and practical work, including developing machine learning-based malicious code pattern recognition technologies, which can automatically identify latent malicious behaviors by analyzing the structure and behavior patterns of the code. At the same time, some advanced dynamic behavior analysis tools have been designed, such as real-time monitoring within Docker containers to detect and handle abnormal behaviors in a timely manner.
[0003] In addition, software composition analysis (SCA) tools such as Snyk, BlackDuck, OWASPDependencyCheck, and Dependabot have also been widely used. These tools combine database records of third-party libraries and corresponding hidden security threats, and update their security databases by relying on existing public security platforms such as GitHubAdvisory, NVD, and OSV, greatly enhancing the ability to identify and protect vulnerable dependencies. At the same time, the sharing of security information by organizations such as the OpenSourceSecurityFoundation (OpenSSF) has greatly promoted information exchange and cooperation among members of the global open-source community. This not only helps community members identify and prevent potential threats in advance but also strengthens the security of the entire open-source ecosystem.
[0004] Open-source intelligence (OSINT) collects information through public channels and is widely used in the field of cybersecurity. Researchers have developed a series of automated systems that can collect key threat intelligence from diverse sources such as the Internet, social media, developer communities, security forums, and code repositories. This intelligence is effectively utilized in attack detection and threat analysis, significantly enhancing the capabilities of cybersecurity defenses. For example, data extracted from developer communities and code repositories can reveal potential code vulnerabilities and malicious code patterns, while information provided by social media and forums reflects the latest security threats and attack trends.
[0005] During the analysis process of OSINT, natural language processing (NLP) techniques and knowledge graphs are widely used to automatically extract and organize cybersecurity knowledge from a large amount of unstructured text. Named entity recognition techniques are particularly used to accurately identify key security-related entities from network texts and social media content, such as specific malware, attack techniques, or security vulnerabilities. Relationship extraction techniques further analyze the semantic connections between these entities, revealing the interaction and influence paths between them.
[0006] Deep learning techniques have also been integrated into intelligence analysis to improve the classification and prediction accuracy of intelligence. For example, through deep learning models, the collected data can be analyzed at a deeper level to identify complex attack patterns and potential threat dynamics. The application of large language models (LLMs) has brought a revolutionary improvement in text analysis, making open source intelligence analysis not only more efficient but also able to conduct more in-depth analysis and understanding.
[0007] However, there are still some deficiencies in the current intelligence analysis for supply chain security, mainly focusing on the analysis of vulnerabilities and malware, while relatively less attention is paid to the specific intelligence of malicious code in package managers such as malicious NPM
[0008] (Node Package Manager) components and malicious PyPI (package index) components. This oversight may cause malicious code to lurk in various images and downstream projects for a long time, forming a continuous security threat. Therefore, there is an urgent need to extend existing advanced analysis techniques to the monitoring of package managers, by timely identifying and addressing security threats in package managers to mitigate risks and enhance the overall cybersecurity posture. This extension can not only improve the response speed and efficiency to existing threats but also more comprehensively protect the security of the open source ecosystem. Summary of the Invention
[0009] The purpose of the present invention is to overcome the above-mentioned defects in the prior art and provide an open source malicious code security intelligence collection method based on a large model, a recursive expansion process driven by a trusted source. The search direction is guided by high-quality initial keywords, and then the information coverage is expanded through recursive access to links, and finally a comprehensive package manager security intelligence library is established.
[0010] To achieve the above purpose, the present invention provides an open source malicious code security intelligence collection method based on a large model, including the following steps:
[0011] S1. Identify uncommon nouns in the intelligence source through a trusted source and mark them as potential malicious component names;
[0012] S2. Extract the SSC entity information of potential malicious component names through the LLM model, and verify whether the extracted SSC entity information is malicious code intelligence;
[0013] S3. Conduct a consistency analysis on the malicious component information related to the same components obtained from different intelligence, aggregate the information of related malicious code, and integrate it into the database.
[0014] Furthermore, in the step S1, there is also a preset trusted intelligence source library. Generate an SSC report according to the trusted intelligence source. This SSC report includes specific keywords related to malicious packages and general keywords related to the software supply chain. Retrieve using the combination of the above specific keywords and general keywords to obtain the initial intelligence source.
[0015] Furthermore, screen the initial intelligence source, select the intelligence source related to package manager security, extract the corresponding links, and analyze the content within the links using regular matching to extract non-common nouns related to software supply chain security, and mark them as potential malicious component names.
[0016] Furthermore, the SSC entity information in the step S2 includes: package name, package manager, package version, discovery date, repository address, attack method, discoverer, affected system, attack vector, and threat indicator.
[0017] Furthermore, after extracting the entity information in the step S2, entity relationship analysis is also required. The entity relationship analysis is based on the package name, and one by one, the package manager, package version, discovery date, repository address, attack method, discoverer, affected system, attack vector, and threat indicator are matched and assigned values.
[0018] Furthermore, the extracted entity information and entity relationship are verified against the original text.
[0019] Furthermore, the consistency analysis in the step S3 is as follows: Match the entity information of different malicious code intelligence, use the package name as the key, collect all the information of the same malicious code from different intelligence sources, and unify its format to achieve data standardization.
[0020] Furthermore, the aggregation is carried out by using a differential voting method. The voting method is as follows: For the six fields of version, repository address, attack method, discoverer, affected system, and attack vector, a voting mechanism is adopted. When the number of votes is the same, select the record with the latest timestamp; for the discovery date, special processing is carried out, and directly select the earliest date as the final value; for the threat indicator, a merging strategy is adopted to merge all non-repeated values.
[0021] An open-source malicious code security intelligence collection system based on a large model, comprising:
[0022] An intelligence identification module, which is used to generate an SSC report according to a trusted source, and extract corresponding keywords from the SSC report to screen initial intelligence sources that meet the requirements from the network;
[0023] A data collection module, which is used to screen and process the web content in the initial intelligence sources, identify non-common nouns therein, and mark them as potential malicious component names;
[0024] An intelligence processing module, which uses an LLM big data model to extract SSC entity information of potential malicious component names, and verifies whether the extracted SSC entity information is malicious code intelligence;
[0025] A data aggregation module, which is used to perform consistency analysis on the above-mentioned malicious code intelligence, and aggregate and integrate information related to the malicious code into a database.
[0026] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps in the above-mentioned open-source malicious code security intelligence collection method based on a large model are implemented.
[0027] A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps in the above-mentioned open-source malicious code security intelligence collection method based on a large model are implemented.
[0028] The present invention aims to start from existing known malware package reports, explore and collect intelligence sources by summarizing domain-specific keywords and using a search engine for snowball search to identify as many intelligence sources as possible.
[0029] Secondly, the present invention introduces large language models (LLMs) to accurately extract key information of malware packages. Specifically, in order to handle potential inaccuracies, such as hallucination phenomena, etc., the present invention uses the domain knowledge of existing malware packages and combines engineering techniques based on chain of thought to accurately guide the LLMs model.
[0030] In addition, the present invention also introduces cross-verification of intelligence from different sources for the same malware package. When information conflicts are identified, a voting mechanism based on intelligence popularity and recency is further adopted to determine the correct information. Description of the Drawings
[0031] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments and descriptions thereof of the invention are used to explain the invention and do not unduly limit the invention.
[0032] Figure 1 It is a schematic diagram of the architecture of the open-source malicious code security intelligence collection method based on a large model in an embodiment of the invention;
[0033] Figure 2 It is a schematic diagram of the automated analysis process of malicious code intelligence;
[0034] Figure 3 It is a schematic diagram of the intelligence release timeline of the malicious component colorwed;
[0035] Figure 4 It is a schematic diagram of the distribution of malicious code component intelligence sources;
[0036] Figure 5 : Schematic diagram of the analysis of the intelligence coverage rate of malicious components. Detailed implementation manners
[0037] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are one embodiment of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0038] It should be noted that all directional indications (such as up, down, left, right, front, back...) in the embodiments of the present invention are only used to explain the relative position relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.
[0039] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features.
[0040] As Figures 1 to 5 shown, the embodiments of the present invention provide an open-source malicious code security intelligence collection method based on a large model, including the following steps:
[0041] S1. Identify the uncommon nouns in the intelligence source and mark them as potential malicious component names;
[0042] This step S1 also includes a preset trusted intelligence source library. The core of this trusted intelligence source library is based on two authoritative security databases, OSV (Open Source Vulnerability Database) and Snyk. The reason they are used as trusted sources is that each malicious package record in these databases has been deeply analyzed and verified by professional security researchers. Each malicious package record contains a complete list of references, which point to the original information sources when the malicious package was first discovered and disclosed, such as blog posts, security bulletins, technical analysis reports, etc. Choosing these databases as trusted sources is precisely because they provide verified and traceable malicious package information.
[0043] Generate an SSC report (Software Supply Chain Security Report) based on the trusted intelligence source. This SSC report is a specialized report form based on authoritative security databases. Its core content comes from the web content pointed to by the malicious package reference information that has been strictly verified by security experts in the OSV (Open Source Vulnerability Database) and Snyk databases. These reference information point to the original information sources when the malicious package was first discovered and disclosed, including professional technical blog posts, official security bulletins, in-depth technical analysis reports, and authoritative vulnerability disclosure documents, which constitute the basic information sources for software supply chain security analysis. The content of the SSC report varies according to the detail level of the original disclosure information. Its basic information usually includes key identification features of the malicious package, such as the package name and the package manager it belongs to (PyPI / NPM), etc. According to the detail level of the disclosure, the report may also include additional technical details, such as malicious behavior analysis, attack payload implementation, scope of impact, etc. This flexible content structure ensures that the report can accurately reflect the integrity of the original disclosure information, regardless of the depth of its technical details.
[0044] In addition, the following rules need to be followed when generating the SSC report: The generation of the SSC report follows a strict technical processing process. Starting from obtaining the reference URL of the malicious package from the trusted source database, web content collection is carried out through a specially designed customized crawler, achieving precise extraction of the core technical content. The crawler system can identify and filter non-core content in the web page (such as navigation bars, advertisements, etc.), and only retain the content of the main body part of the web page.
[0045] The SSC reports generated from the preset trusted sources thus contain specific keywords related to malicious packages (such as specific identifiers like package names and version numbers) and general keywords related to the software supply chain (such as "malicious", "supplychainattack", etc.). Therefore, when performing operations, we first need to extract them from the SSC reports. The importance of this design lies in addressing the issue of the incompleteness of trusted source data. Because in the OSV and Snyk databases, some malicious packages may lack complete reference information or the reference is not comprehensive enough. Therefore, we combine the extracted specific keywords and general keywords and conduct searches through web search engines such as Google to discover more relevant security reports and technical analysis articles, thereby expanding the scope of the intelligence sources to ensure that even when the original reference is incomplete, we can still obtain more comprehensive initial intelligence sources related to malicious packages. This trusted-source-driven and keyword-guided expansion mechanism enables us to establish a more complete and comprehensive malicious package intelligence collection system, effectively making up for the information missing problem that may be brought about by simply relying on preset trusted sources.
[0046] Next, screen the initial intelligence sources, select the intelligence sources related to package manager security, and extract the corresponding referenced links. Then, visit and analyze these new links one by one, and use regular matching to determine whether the web content is related to software supply chain security, and extract the uncommon nouns related to software supply chain security and mark them as potential malicious component names (SSC content).
[0047] Such as Figure 2 (1) shows the web content screened from the initial intelligence sources. It can be seen from the figure that the web content involves nouns related to software supply chain security, and the obtained Figure 2 (2) The uncommon nouns ('PyProto2', 'Pyg-utils','sonatype', 'typosquatting','malware’, 'AWs', 'pygrata.com', 'pymocks.com’, 'PyP!', 'pymocks', 'password-stealing”DNS', 'CheckPoint'), these are the potential malicious component names that need to be marked.
[0048] S2. Extract the SSC entity information of potential malicious component names through the LLM model. For the extraction of entity information, in this embodiment, an entity rule for extracting ten standardized information, namely package name, package manager, package version, discovery date, repository address, attack method, discoverer, affected system, attack vector, and threat indicator, is used as the entity information for extraction.
[0049] In large language model prompts such as Figure 2 (3a), classify, define, and establish entity rules for SSC entity information in the software supply chain model, and have the large language model sequentially extract software supply chain entities existing in potential malicious component names.
[0050] After extracting entity information, entity relationship analysis is also required Figure 2 (3b). The core of entity relationship analysis is to use the package name as a benchmark to verify whether the other nine entity attributes actually belong to the malicious package. That is, based on the package name, pair and assign values to the package manager, package version, discovery date, repository address, attack method, discoverer, affected system, attack vector, and threat indicator one by one. Specifically, after determining a malicious package name, the large language model will carefully verify the attribution relationship of each attribute against the original text content: which package manager platform (such as PyPI or NPM) this package belongs to, the specific version number, the specific date of the first discovery or report, the repository address storing this malicious package, the specific attack method used by this package, who first discovered this malicious package, which systems are affected by this package, what attack vector this package uses, and the specific threat indicator related to this package. Thus, it can be verified whether the extracted SSC entity information is malicious code intelligence;
[0051] Finally, it is also necessary to check and verify the extracted entity information and entity relationships against the original text, which is entity verification. The core of entity verification is a two-step verification mechanism: first, conduct entity accuracy verification, that is, ensure that the entity information extracted from the original text is accurate (such as verifying that the extracted package name is indeed a malicious package in a package manager, the extracted version number is the correct version of the package, and whether the date format is standardized, etc.); second, conduct entity relationship verification, that is, based on the package name, verify against the original text whether the other nine entities (package manager, version number, discovery date, repository address, attack method, discoverer, affected system, attack vector, threat indicator) actually belong to this malicious package, ensuring that the extracted information is not only accurate but also has the correct interrelationship. This dual verification mechanism guarantees the reliability and integrity of the finally extracted malicious package intelligence and provides a high-quality data basis for subsequent security analysis.
[0052] S3. Conduct consistency analysis on the malicious component information related to the same components obtained from different intelligence. The consistency analysis is as follows: match the entity information of different malicious code intelligence, use the package name as the key, collect all information of the same malicious code from different intelligence sources, and standardize its format to achieve data standardization; then adopt a differential voting method to aggregate the information of relevant malicious code and integrate it into the database.
[0053] The intelligence aggregation in this step S3 adopts a voting mechanism based on differential processing. First is the data standardization stage, which unifies the formats of malicious package intelligence from all sources to achieve data standardization. That is, each malicious package P contains ten standardized entity fields (package name N, version V, discovery date F, repository URL R, attack method M, discoverer D, affected system I, attack vector A, threat indicator C, timestamp T).
[0054] Then it enters the data aggregation stage, as Figure 2 (4) The system uses the malicious package name N as the key to collect all information of the same malicious package from different intelligence sources. In the voting processing stage, the system adopts different processing strategies for different fields: for the six fields of version V, repository URL R, attack method M, discoverer D, affected system I, and attack vector A, a voting mechanism is adopted. When the number of votes is the same, the record with the latest timestamp is selected; for the discovery date F, special processing is adopted, and the earliest date is directly selected as the final value; for the threat indicator C, a merging strategy is adopted to merge all non-repeated values. This differential processing mechanism ensures that different types of intelligence can be aggregated and processed most appropriately, thus providing more accurate and complete malicious package information.
[0055] Next, taking the hypothetical malicious package "malicious-pkg" as an example, we have collected the following information from three different intelligence sources (Intelligence Source A, Intelligence Source B, and Intelligence Source C):
[0056] Intelligence Source A (timestamp: 2023-10-01):
[0057]
[0058]
[0059] Intelligence Source B (timestamp: 2023-09-15):
[0060] 1 Package Name: malicious-pkg 2 Version: 1.0.0 3 Discovery Date: 2023-01-10 4 Repository URL: github.com / repo2 5 Attack Method: Data Theft 6 Discoverer: Researcher Y 7 Threat Indicator: IP-2, URL-1
[0061] Intelligence Source C (timestamp: 2023-08-20):
[0062]
[0063]
[0064] According to the voting mechanism, the final aggregation result is:
[0065]
[0066] The entire intelligence source construction process starts from trusted sources (OSV and Snyk databases). The core role of trusted sources is to provide high-quality initial keywords for subsequent web searches. Specifically, first, obtain the reference links of malicious packages verified by experts from trusted sources, extract the text content corresponding to these links (i.e., SSC reports) through customized crawlers, and extract specific keywords (such as package names, version numbers, etc.) and general keywords (such as supply chain security-related terms) from them. These keywords extracted from trusted sources are used as the input for web searches to help discover a wider range of open-source intelligence content. For each web page returned by the web search, first analyze whether its content is related to package manager security. For the confirmed relevant web pages, then extract the links cited in its text, access and analyze these new links, and determine whether their web page content is related to software supply chain security. This recursive expansion process driven by trusted sources guides the search direction through high-quality initial keywords, and then expands the information coverage through recursive access to links, such as Figure 4 and Figure 5 the integrity of the collected open-source malicious component intelligence has been significantly improved compared to mainstream open-source security data platforms; at the same time, it can effectively solve problems such as Figure 3 shown in the lag of malicious component intelligence information, effectively improving the timely and accurate utilization of malicious component intelligence information for downstream applications, such as software composition analysis security detection, open-source ecosystem governance, etc. By continuously repeating the above steps, a comprehensive package manager security intelligence library is finally established.
[0067] In some preferred embodiments, an open-source malicious code security intelligence collection system based on a large model is also provided, including:
[0068] An intelligence identification module, which is used to generate an SSC report according to trusted sources and extract corresponding keywords from the SSC report to screen initial intelligence sources that meet the requirements from the network;
[0069] A data collection module, which is used to screen and process the web page content in the initial intelligence source, identify the uncommon nouns therein, and mark them as potential malicious component names;
[0070] An intelligence processing module, which uses the LLM big data model to extract SSC entity information of potential malicious component names and verify whether the extracted SSC entity information is malicious code intelligence;
[0071] A data aggregation module, which is used to perform consistency analysis on the above malicious code intelligence and aggregate and integrate the information of relevant malicious code into the database.
[0072] The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the first embodiment above. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.
[0073] In the above embodiments, the descriptions of each embodiment have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0074] The proposed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the above modules is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.
[0075] In some preferred embodiments, a computer-readable storage medium is also provided, on which a computer program is stored. When the program is executed by a processor, the steps in the above-mentioned method for collecting open-source malicious code security intelligence based on a large model are implemented.
[0076] In some preferred embodiments, a computer device is also provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps in the above-mentioned method for collecting open-source malicious code security intelligence based on a large model are implemented.
[0077] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.
[0078] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more flows or multiple flows and / or blocks Figure 1 one or more blocks or multiple blocks.
[0079] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one process or more processes and / or one block or more blocks in the process Figure 1 one process or more processes and / or Figure 1 one block or more blocks.
[0080] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process or more processes and / or one block or more blocks in the process Figure 1 one process or more processes and / or Figure 1 one block or more blocks.
[0081] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0082] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A method for collecting security intelligence of open source malicious code based on a large model, characterized in that: The steps include: S1. Identify uncommon nouns in intelligence sources through trusted sources and mark them as potential malicious component names; S2. Extract the SSC entity information of the potential malicious component name through the LLM model, and verify whether the extracted SSC entity information is malicious code intelligence; S3. Perform consistency analysis on the malicious component information related to the same component obtained from different intelligence, aggregate the information of related malicious codes and integrate it into the database.
2. According to the large model-based open source malicious code security intelligence collection method of claim 1, it is characterized in that: In the step S1, an SSC report is generated based on a trusted source, and the SSC report includes specific keywords related to malicious packages and general keywords related to the software supply chain. The above specific keywords and general keywords are used in combination to perform a search to obtain an initial intelligence source.
3. The method for collecting security intelligence of open source malicious code based on a large model according to claim 2, characterized in that: The initial intelligence sources are screened, and those related to package manager security are selected. The corresponding links are extracted, and the content in the links is analyzed using regular matching and uncommon nouns related to software supply chain security are extracted and marked as potential malicious component names.
4. The method for collecting security intelligence of open source malicious code based on a large model according to claim 1, characterized in that: The SSC entity information in step S2 includes: package name, package manager, package version, discovery date, warehouse address, attack method, discoverer, affected system, attack vector, and threat indicator.
5. The method for collecting security intelligence of open source malicious code based on a large model according to claim 4 is characterized in that: After extracting the entity information in step S2, an entity relationship analysis is required. The entity relationship analysis is based on the package name and matches and assigns values to the package manager, package version, discovery date, warehouse address, attack method, discoverer, affected system, attack vector, and threat indicator one by one.
6. The method for collecting security intelligence of open source malicious code based on a large model according to claim 5, characterized in that: The extracted entity information and entity relationships are checked and verified with the original text.
7. The method for collecting security intelligence of open source malicious code based on a large model according to claim 4, characterized in that: The consistency analysis in step S3 is as follows: matching entity information of different malicious code intelligence, using the malicious package name as a key, collecting all information of the same malicious code from different intelligence sources, and unifying its format to achieve data standardization.
8. The method for collecting security intelligence of open source malicious code based on a large model according to claim 7, characterized in that: The aggregation is performed using a differentiated voting method, which is: a voting mechanism is used for the six fields of version, warehouse address, attack method, discoverer, affected system, and attack vector. When the number of votes is the same, the record with the latest timestamp is selected; Special treatment is applied to the discovery date, and the earliest date is directly selected as the final value; a merging strategy is adopted for threat indicators, and all non-duplicate values are merged.
9. An open source malicious code security intelligence collection system based on a large model, characterized in that: include: An intelligence identification module, the intelligence identification module is used to generate an SSC report based on a trusted source, and extract corresponding keywords from the SSC report to screen initial intelligence sources that meet the requirements from the network; A data collection module, which is used to screen and process the webpage content in the initial intelligence source, identify uncommon nouns therein, and mark them as names of potential malicious components; Data module module, this intelligence processing module uses the LLM big data model to extract the SSC entity information of the potential malicious component name, and verifies whether the extracted SSC entity information is malicious code intelligence; The data aggregation module is used to perform consistency analysis on the above malicious code intelligence, and aggregate the information of related malicious codes and integrate it into the database.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in a method for collecting open source malicious code security intelligence based on a large model as described in any one of claims 1-8 are implemented.