Original vulnerability data translation and compression method based on large language model and related equipment
The large language model extracts structured facts from vulnerability scan data and establishes MSF module mapping, which solves the problem of vulnerability reporting redundancy and realizes efficient vulnerability data conversion and automated processing.
Patent Information
- Application Number
- CN202510651008.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The original report generated by the existing vulnerability scanning tool contains a large number of natural language descriptions and information redundant, making it difficult to directly provide the mapping of vulnerabilities with the Metasploit Framework (MSF) module, affecting work efficiency and automated processing.
Using a method based on a large language model, structured facts are extracted from the original vulnerability data through preset Prompt templates, including network service facts, vulnerability existence facts and vulnerability attribute facts, and a direct mapping with the MSF utilization module is established to generate a compressed vulnerability data set.
It realizes the conversion from original vulnerable data to structured and information refinement, reduces the amount of data, improves work efficiency, facilitates machine automated processing, and ensures data accuracy and consistency.
Smart Images

Figure CN120567458A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of network security technology, and in particular relates to a large language model-based original vulnerability data translation and compression method and related equipment. Background Art
[0002] Vulnerability scanning is a crucial tool for identifying potential security risks during network security assessments and penetration testing. Scanning tools (such as Nmap with scripts, Nessus, OpenVAS, or the built-in scanning module of MSF) can identify services, ports, and known vulnerabilities on target systems. However, the raw reports generated by these tools often contain extensive natural language descriptions, diverse data formats, and high levels of information redundancy.
[0003] Security analysts need to extract key information from these complex reports, such as the affected IP address, port, service type, and CVE number of the vulnerability, and determine the nature of the vulnerability (e.g., remote code execution, SQL injection, etc.). Furthermore, they need to find available exploit code or methods for these vulnerabilities. Especially in frameworks like Metasploit Framework (MSF), which integrate a large number of pre-built vulnerability exploit modules, how to quickly and accurately match the vulnerabilities discovered by scanning with specific exploit modules in MSF is a key step affecting work efficiency.
[0004] Currently, this process often relies on manual experience or requires secondary processing with external scripts or tools. Raw scan data itself doesn't directly provide this precise mapping to MSF modules, and its redundant text descriptions hinder automated processing and efficient storage. For example, to convey the fact that "the httpd service at IP address 10.20.0.68 has the CVE-2007-1603 vulnerability, which is a remotely exploitable privilege escalation vulnerability and has a corresponding exploit module in MSF," traditional reports might require a significant amount of text, making such a description difficult to use directly in automated processes. Summary of the Invention
[0005] In response to the problems existing in the prior art, the present invention provides a method and related equipment for translating and compressing raw vulnerability data based on a large language model, which can convert raw and lengthy vulnerability scanning data into a highly structured and information-refined representation.
[0006] In order to solve the above technical problems, the present invention is implemented through the following technical solutions: According to a first aspect of the present invention, a method for translating and compressing original vulnerability data based on a large language model is provided, comprising: Based on a preset prompt template, a large language model is used to extract structured facts from the raw vulnerability data. The structured facts include network service facts and vulnerability existence facts. The network service facts are used to describe the target IP address, the name of the running service, the protocol type, the port number, and the service running permissions. The vulnerability existence facts are used to describe the vulnerabilities existing in the service running on the target IP address, and the existing vulnerabilities are identified by standard vulnerability numbers. Associating standardized vulnerability attribute labels with vulnerabilities with standard vulnerability numbers to generate vulnerability attribute facts; Establish a direct mapping relationship between vulnerabilities associated with standardized vulnerability attribute labels and MSF exploit module paths to generate MSF exploit module mapping facts; The network service facts, vulnerability existence facts, vulnerability attribute facts and MSF utilization module mapping facts are combined, translated and compressed to form a compressed vulnerability data set, and the compressed vulnerability data set is represented by standardized predicate logic or key-value pair set.
[0007] In a possible implementation of the first aspect, the preset Prompt template is specifically: Extract the target IP address, running service name, protocol type, port number, service running permissions, standard vulnerability number, vulnerability attribute category and vulnerability attribute value from the original vulnerability data and output: networkSI (target IP address, running service name, protocol type, port number and service running permissions); vulE (target IP address, standard vulnerability number, running service name); vulP (standard vulnerability number, vulnerability attribute category, vulnerability attribute value).
[0008] In a possible implementation of the first aspect, the network service fact is represented as follows: networkSI (target IP address, running service name, protocol type, port number and service running permissions).
[0009] In a possible implementation of the first aspect, the fact that the vulnerability exists is expressed as follows: vulE (target IP address, standard vulnerability number, running service name).
[0010] In a possible implementation of the first aspect, the vulnerability attribute fact is represented as follows: vulP (standard vulnerability number, vulnerability attribute category, vulnerability attribute value).
[0011] In a possible implementation of the first aspect, the MSF utilizes a module mapping fact in the form of: mapExploit (standard vulnerability number, exploit module path).
[0012] In a possible implementation manner of the first aspect, the compressed vulnerability dataset is stored in text, JSON, or XML format. When stored in text format, each row represents a fact.
[0013] According to a second aspect of the present invention, there is provided a device for translating and compressing original vulnerability data based on a large language model, comprising: An extraction module is used to extract structured facts from raw vulnerability data using a large language model based on a preset prompt template. The structured facts include network service facts and vulnerability existence facts. The network service facts describe the target IP address, the name of the running service, the protocol type, the port number, and the service running permissions. The vulnerability existence facts describe the vulnerabilities existing in the service running on the target IP address and identify the existing vulnerabilities using standard vulnerability numbers. An association module is used to associate standardized vulnerability attribute labels with vulnerabilities with standard vulnerability numbers to generate vulnerability attribute facts; A mapping module is used to establish a direct mapping relationship between vulnerabilities associated with standardized vulnerability attribute labels and MSF exploit module paths, generating MSF exploit module mapping facts; The compression module is used to combine the network service facts, vulnerability existence facts, vulnerability attribute facts and MSF utilization module mapping facts, translate and compress them to form a compressed vulnerability data set, and the compressed vulnerability data set is represented by standardized predicate logic or key-value pair set.
[0014] According to a third aspect of the present invention, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for translating and compressing original vulnerability data based on a large language model is implemented.
[0015] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for translating and compressing original vulnerability data based on a large language model is implemented.
[0016] Compared with the prior art, the present invention has at least the following beneficial effects: The present invention provides a method for translating and compressing raw vulnerability data based on a large language model. By presetting a prompt template and a large language model, structured facts are extracted from the raw vulnerability data, including network service facts, vulnerability existence facts, vulnerability attribute facts, and MSF utilization module mapping facts, and the data is translated and compressed to form a compressed vulnerability data set represented by a standardized predicate logic or key-value pair set, effectively removing lengthy natural language descriptions, retaining only core operational information, reducing the amount of data, and improving the refinement of the information. The present invention utilizes a large language model and a preset prompt template to quickly and accurately extract key information from the raw vulnerability data, generate structured facts, and establish a direct mapping relationship between the vulnerability and the MSF utilization module. Compared with the current method of relying on manual experience or using external scripts and tools for secondary processing, the present invention greatly shortens the time required for this process, improves work efficiency, and facilitates security analysts to carry out subsequent penetration testing and vulnerability repair work more efficiently. This invention converts raw vulnerability data into standardized predicate logic or key-value pairs, replacing lengthy text descriptions. This allows for the compression and structuring of raw data, making it easier for machines to read and process. Automated systems can quickly analyze and make decisions based on compressed vulnerability datasets. This invention uses a unified, pre-set prompt template and a large language model to process raw vulnerability data, ensuring the accuracy and consistency of the extracted data in both format and content. Regardless of the scanning tool used for the raw data, this invention processes it to produce structured data in a uniform and reliable format, reducing errors caused by tool differences and manual processing.
[0017] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the specific embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the specific embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 This is a flow chart of a method for translating and compressing original vulnerability data based on a large language model according to the present invention; Figure 2 This is a schematic diagram of the structure of the compressed vulnerability dataset of the present invention. DETAILED DESCRIPTION
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0021] Combine Figure 1 and Figure 2 As shown, an embodiment of the present invention provides a method for translating and compressing original vulnerability data based on a large language model, which specifically includes the following steps: Step 1: Based on the preset prompt template, a large language model is used to extract structured facts from the original vulnerability data. The structured facts include network service facts and vulnerability existence facts.
[0022] Specifically, based on the preset prompt template, the large language model is guided to extract key information from raw vulnerability data reports in different formats. The preset prompt template is as follows: Extract the target IP address, running service name, protocol type, port number, service running permissions, standard vulnerability number, vulnerability attribute category and vulnerability attribute value from the original vulnerability data and output: networkSI (target IP address, running service name, protocol type, port number and service running permissions); vulE (target IP address, standard vulnerability number, running service name); vulP (standard vulnerability number, vulnerability attribute category, vulnerability attribute value).
[0023] Preferably, the preset prompt template also has the following features: through the few-sample thinking chain technology, combined with a small number of annotated examples, it guides the large language model to accurately extract key information from vulnerability scanning reports in different formats. The network service fact is used to describe the target IP address, the name of the running service, the protocol type, the port number and the service running permissions. In this embodiment, the network service fact is represented as: networkSI (target IP address, the name of the running service, the protocol type, the port number and the service running permissions).
[0024] That is, for each identified open service, a network service fact (networkSI) is generated (target IP address, running service name, protocol type, port number, and service running permissions). For example, if the scan finds that the httpd service is running on port 80 of 10.30.0.29 using the TCP protocol, and it is inferred that the service is running with root permissions, the generated network service fact is: networkSI('10.30.0.29', httpd, tcp, 80, root).
[0025] The vulnerability existence fact is used to describe a vulnerability existing in the service running at the target IP address, and the vulnerability is identified by a standard vulnerability number (such as a CVE ID). In this embodiment, the vulnerability existence fact is represented as: vulE(target IP address, standard vulnerability number, running service name).
[0026] That is, for each identified vulnerability associated with a service running on the target IP address, a vulnerability existence fact is generated: vulE(target IP address, standard vulnerability number, running service name). For example, if a scan discovers that the httpd service running on 10.20.0.68 has the CVE-2007-1603 vulnerability, the vulnerability existence fact generated is: vulE('10.20.0.68', 'CVE-2007-1603', 'httpd').
[0027] It should be understood that the raw vulnerability data is obtained from one or more vulnerability scanning tools (such as Nmap's NSE script output, Nessus report XML file, vulns table in MSF database, etc.).
[0028] Step 2: Associate standardized vulnerability attribute labels with vulnerabilities with standard vulnerability numbers to generate vulnerability attribute facts. Vulnerability attribute facts are used to describe the technical characteristics or attack types corresponding to specific vulnerability numbers, such as remote code execution (RCE), SQL injection (SQLi), and privilege escalation (privEscalation).
[0029] It should be noted that when associating standardized vulnerability attribute tags with vulnerabilities with standard vulnerability numbers, a vulnerability knowledge base is utilized, which contains attribute information of known standard vulnerability numbers CVE.
[0030] In this embodiment, the vulnerability attribute fact is represented as: vulP (standard vulnerability number, vulnerability attribute category, vulnerability attribute value). For example, for CVE-2007-1603, the vulnerability attribute fact may be generated as follows: vulP('CVE-2007-1603', 'vulType', 'privEscalation'). vulP('CVE-2007-1603', 'accessVector', 'remote'). Common vulnerability attribute categories include but are not limited to deserialization, cross-site scripting (XSS), information leakage and buffer overflow. Vulnerability attribute values include RCE, SQLi, XSS, Deserialization, AuthBypass, InfoLeak, BufferOverflow, etc.
[0031] Step 3: Establish a direct mapping relationship between the vulnerability associated with the standardized vulnerability attribute label and the MSF (Metasploit Framework) exploit module path, and generate an MSF exploit module mapping fact. The MSF exploit module mapping fact clearly indicates the full path of the MSF exploit module corresponding to the specific vulnerability number.
[0032] It should be understood that Metasploit Framework (MSF) is an open-source penetration testing and security development platform widely used by security researchers, penetration testers, and ethical hackers for vulnerability exploitation, vulnerability research, and security assessments.
[0033] It should be noted that when establishing a direct mapping relationship between a vulnerability associated with a standardized vulnerability attribute label and an MSF exploit module path, a mapping table of known standard vulnerability numbers CVE to MSF exploit module paths is used.
[0034] In this embodiment, the MSF exploit module mapping fact is represented as: mapExploit (standard vulnerability number, exploit module path). For example: mapExploit('CVE-2017-12629', 'exploit / multi / http / solr_velocity_rce'). mapExploit('CVE-2018-20062', 'exploit / unix / webapp / thinkphp_rce'). mapExploit('CVE-2019-15107', 'exploit / linux / http / webmin_backdoor'). If a CVE does not have a direct corresponding MSF exploit module path, this fact can be empty or point to a generic auxiliary module or manual exploitation guide.
[0035] Step 4: Combine the network service facts, vulnerability existence facts, vulnerability attribute facts, and MSF utilization module mapping facts, translate and compress them to form a compressed vulnerability dataset. The compressed vulnerability dataset is represented by standardized predicate logic or key-value pair sets, replacing lengthy natural language descriptions, thereby compressing and structuring the original vulnerability data.
[0036] In this embodiment, the compressed vulnerability dataset is stored in text, JSON or XML format. When stored in text format, each row represents a fact.
[0037] In other words, the facts are collected to form a logical fact set. This set can be stored simply as a text file, with one fact per line; or in a more structured form such as a JSON object array, an XML document, or a database table. For example, a portion of the compressed vulnerability dataset is as follows: networkSI('10.30.0.29', httpd, tcp, 80, root). networkSI('10.20.0.68', httpd, tcp, 80, user). vulE('10.20.0.68', 'CVE-2007-1603', 'httpd'). vulP('CVE-2007-1603', 'vulType', 'privEscalation'). vulP('CVE-2007-1603', 'accessVector', 'remote'). mapExploit('CVE-2007-1603', 'exploit / windows / some_module_for_cve_2007_1603'). vulE('10.30.0.29', 'CVE-2019-15107', 'httpd'). mapExploit('CVE-2019-15107', 'exploit / linux / http / webmin_backdoor'). vulP('CVE-2019-15107', 'vulType', 'RCE'). This fact set is a compressed representation of the original vulnerability data, removing descriptive text and retaining only the core, actionable, machine-readable information.
[0038] When a user or automated script needs to take action against a target IP and vulnerability (e.g., vulE('10.20.0.68', 'CVE-2007-1603', 'httpd')), the MSF exploit module mapping fact mapExploit associated with CVE-2007-1603 is used to obtain the MSF exploit module path (e.g., exploit / windows / some_module_for_cve_2007_1603). The target IP (10.20.0.68) and port (80) are extracted from the related network service fact networkSI (e.g., networkSI('10.20.0.68',httpd, tcp, 80, user).).
[0039] As a more preferred embodiment, a large language model-based method for translating and compressing raw vulnerability data also includes a dynamic optimization process: adjusting the preset prompt template and the large language model's translation strategy based on manual verification or automated feedback to improve translation accuracy and adaptability to new report formats. Specifically, the prompt template is adjusted based on historical translation results and user feedback.
[0040] It also features a multimodal input processing mechanism that supports automatic recognition and processing of various report formats (such as text, XML, and JSON), and optimizes the mapping of vulnerability attributes to Metasploit Framework (MSF) exploit modules through a dynamic knowledge graph. It supports multimodal input processing, automatically identifying and processing vulnerability scan reports in various formats, such as text, XML, and JSON, and dynamically selecting the appropriate prompt template based on the report format. It maintains the mapping relationship between vulnerability attributes and MSF exploit modules through a dynamic knowledge graph, updating graph nodes and edges in real time to accommodate new vulnerability types and incremental changes to MSF modules.
[0041] This implementation method uses a dynamic optimization process and a multimodal input processing mechanism, combined with manual verification or automated feedback to adjust the translation strategy, supports automatic recognition and processing of multiple report formats, and uses dynamic knowledge graphs to optimize mapping and adapt to changes in the network security environment.
[0042] The present invention is described below with reference to specific cases.
[0043] Original data: interfaces: ############################################################ # DMZ Layer (10 hosts: 10.1.0.11 ~ 10.1.0.20) ############################################################ - hostname: dmz-web-11 os_type: debian iface: eth0 mac_address: "00:50:56:a1:11:01" address: "10.1.0.11 / 24" secondary_ips: ["10.1.0.111 / 24"] gateway: "10.1.0.1" dns: ["8.8.8.8", "8.8.4.4"] mtu: 1500 vlan: 100 state: up description: "DMZ Apache Web Server" services: - name: dmz_web_11 protocol: tcp port: 80 permissions: root vulnerabilities: - cve: CVE-2024-1234 description: "Apache httpd " attributes: vulType: RCE accessVector: remote - hostname: dmz-web-12 os_type: windows iface: Ethernet0 mac_address: "00:50:56:a1:11:02" address: "10.1.0.12 / 24" gateway: "10.1.0.1" dns: ["8.8.8.8", "8.8.4.4"] mtu: 1500 state: up description: "DMZ IIS Server" services: - name: dmz_web_12 protocol: tcp port: 80 permissions: user - hostname: dmz-web-13 os_type: rhel iface: ens192 mac_address: "00:50:56:a1:11:03" address: "10.1.0.13 / 24" gateway: "10.1.0.1" dns: ["8.8.8.8", "8.8.4.4"] state: up description: "DMZ Nginx Reverse Proxy" services: - name: dmz_web_13 protocol: tcp port: 80 permissions: user ############################################################ # Application Layer (10.2.0.11 ~ 10.2.0.30) ############################################################ - hostname: app-11 os_type: debian iface: eth0 mac_address: "00:50:56:a1:21:01" address: "10.2.0.11 / 24" gateway: "10.2.0.1" dns: ["10.3.0.11", "10.3.0.12"] state: up bonding: false description: "Tomcat App Server #1" services: - name: app_11 protocol: tcp port: 8080 permissions: user - hostname: app-12 os_type: windows iface: Ethernet0 mac_address: "00:50:56:a1:21:02" address: "10.2.0.12 / 24" gateway: "10.2.0.1" dns: ["10.3.0.11", "10.3.0.12"] state: up description: ".NET App Server #2" services: - name: app_12 protocol: tcp port: 8080 permissions: user vulnerabilities: - cve: CVE-2024-5678 description: ".NET" attributes: vulType: infoLeak accessVector: remote ############################################################ # Database Layer (10.3.0.11 ~ 10.3.0.15) ############################################################ - hostname: db-11 os_type: debian iface: eth0 mac_address: "00:50:56:a1:31:01" address: "10.3.0.11 / 24" gateway: "10.3.0.1" dns: ["8.8.8.8"] mtu: 9000 state: up description: "MySQL Primary Database" services: - name: db_11 protocol: tcp port: 3306 permissions: root - hostname: db-12 os_type: windows iface: Ethernet0 mac_address: "00:50:56:a1:31:02" address: "10.3.0.12 / 24" gateway: "10.3.0.1" dns: ["8.8.8.8"] state: up description: "MSSQL Database Server" services: - name: db_12 protocol: tcp port: 1433 permissions: user ############################################################ # File Server Layer (10.3.0.16 ~ 10.3.0.20) ############################################################ - hostname: fs-16 os_type: debian iface: eth0 mac_address: "00:50:56:a1:32:01" address: "10.3.0.16 / 24" gateway: "10.3.0.1" dns: ["8.8.8.8"] state: up description: "FTP Server (Internal)" services: - name: fs_16 protocol: tcp port: 21 permissions: user ############################################################ # Workstation Layer (10.3.0.31 ~ 10.3.0.50) ############################################################ - hostname: ws-31 os_type: windows iface: Ethernet0 mac_address: "00:50:56:a1:33:01" address: "10.3.0.31 / 24" gateway: "10.3.0.1" dns: ["8.8.8.8", "8.8.4.4"] state: up vlan: 300 description: "Employee Windows Workstation" services: - name: ws_31 protocol: tcp port: 3389 permissions: user - hostname: ws-32 os_type: debian iface: eth0 mac_address: "00:50:56:a1:33:02" address: "10.3.0.32 / 24" gateway: "10.3.0.1" dns: ["8.8.8.8", "8.8.4.4"] state: up description: "Employee Linux Workstation" services: - name: ws_32 protocol: tcp port: 22 permissions: user The data compressed by the method of the present invention is as follows: networkSI('10.1.0.11', dmz_web_11, tcp, 80, root). networkSI('10.2.0.12', app_12, tcp, 8080, user). vulE('10.1.0.11', 'CVE-2024-1234', dmz_web_11). vulE('10.2.0.12', 'CVE-2024-5678', app_12). vulP('CVE-2024-1234', 'vulType', 'RCE'). vulP('CVE-2024-1234', 'accessVector', 'remote'). vulP('CVE-2024-5678', 'vulType', 'infoLeak'). vulP('CVE-2024-5678', 'accessVector', 'remote'). In the original data, each host's information includes numerous fields, such as host name, operating system type, interface name, MAC address, gateway, DNS, and so on. This is both voluminous and redundant. Compressed data, on the other hand, retains only the network service facts, vulnerability existence facts, and vulnerability attribute facts that are closely related to network security analysis. Non-core information such as host name and operating system type is removed, significantly reducing the data volume and making it more concise and clear, making it easier to quickly obtain key information.
[0044] Despite data compression, all key information is preserved. Network service facts clearly provide key elements such as the target IP address, running service name, protocol type, port number, and service permissions, clearly demonstrating the operational status of network services. Vulnerability facts accurately identify the vulnerable IP address, standard vulnerability number, and running service name, enabling analysts to quickly locate the vulnerability. Vulnerability attribute facts further provide important attributes such as the vulnerability type and access vector.
[0045] The compressed data uses a standardized predicate logic representation with a uniform format and clear structure. This representation facilitates automated machine processing and parsing. For example, security analysis tools or automated scripts can quickly read this data and perform further analysis, statistics, or take appropriate security measures based on the factual information. This eliminates the need for complex processing and parsing of the original complex data structure, improving the efficiency and accuracy of automated processing.
[0046] By combining the existence of a vulnerability with its attributes, we can quickly understand its specific circumstances. For example, the compressed data clearly shows that the dmz_web_11 service running on the host with the IP address 10.1.0.11 is vulnerable to CVE-2024-1234, a remotely exploitable remote code execution vulnerability. This provides security analysts with clear vulnerability information, enabling them to quickly assess the risk and take appropriate measures, such as updating software and deploying protection policies, thereby effectively improving work efficiency and the timeliness of vulnerability resolution.
[0047] The embodiment of the present invention provides a large language model-based original vulnerability data translation and compression device, which is used to implement the aforementioned large language model-based original vulnerability data translation and compression method, and specifically includes the following modules: The extraction module is used to extract structured facts from the original vulnerability data based on a preset prompt template and a large language model. The structured facts include network service facts and vulnerability existence facts. The network service facts are used to describe the target IP address, the name of the running service, the protocol type, the port number and the service running permissions; the vulnerability existence facts are used to describe the vulnerabilities existing in the service running on the target IP address, and the existing vulnerabilities are identified by standard vulnerability numbers.
[0048] The association module is used to associate standardized vulnerability attribute labels with vulnerabilities with standard vulnerability numbers to generate vulnerability attribute facts.
[0049] The mapping module is used to establish a direct mapping relationship between vulnerabilities associated with standardized vulnerability attribute labels and MSF exploit module paths, and generate MSF exploit module mapping facts.
[0050] The compression module is used to combine the network service facts, vulnerability existence facts, vulnerability attribute facts and MSF utilization module mapping facts, translate and compress them to form a compressed vulnerability data set, and the compressed vulnerability data set is represented by standardized predicate logic or key-value pair set.
[0051] All relevant contents of each step involved in the embodiment of the aforementioned method for translating and compressing raw vulnerability data based on a large language model can be referred to the functional description of the functional module corresponding to the device for translating and compressing raw vulnerability data based on a large language model in the embodiment of the present invention, and will not be repeated here. The division of modules in the embodiment of the present invention is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, the functional modules in various embodiments of the present invention can be integrated into one processor, or they can exist physically separately, or two or more modules can be integrated into one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0052] In another embodiment of the present invention, a computer device is provided, which includes a processor and a memory, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions in a computer storage medium to implement a corresponding method flow or corresponding function; the processor described in the embodiment of the present invention can be used for the operation of a method for translating and compressing original vulnerability data based on a large language model.
[0053] In another embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device, used to store programs and data. It is understood that the computer-readable storage medium herein may include both built-in storage media in the computer device and, of course, extended storage media supported by the computer device. The computer-readable storage medium provides storage space, which stores the terminal's operating system. Furthermore, the storage space also stores one or more instructions suitable for being loaded and executed by a processor. These instructions may be one or more computer programs (including program code). It should be noted that the computer-readable storage medium herein may be high-speed RAM memory or non-volatile memory, such as at least one disk storage device. The processor may load and execute the one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the method for translating and compressing raw vulnerability data based on a large language model in the above-mentioned embodiment.
[0054] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0055] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0056] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0057] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0058] The present invention also provides a computer program product for executing any of the aforementioned methods for translating and compressing raw vulnerability data based on a large language model. Because the computer program product provided by the present invention and the aforementioned method for translating and compressing raw vulnerability data based on a large language model are based on the same inventive concept, the computer program product provided by the present invention possesses all the advantages of the aforementioned method for translating and compressing raw vulnerability data based on a large language model. Therefore, the beneficial effects of the computer program product provided by the present invention will not be detailed here.
[0059] In the present invention, the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples without contradiction.
[0060] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A method for translating and compressing original vulnerability data based on a large language model, characterized in that: include: Based on a preset prompt template, a large language model is used to extract structured facts from the raw vulnerability data. The structured facts include network service facts and vulnerability existence facts. The network service facts are used to describe the target IP address, the name of the running service, the protocol type, the port number, and the service running permissions. The vulnerability existence facts are used to describe the vulnerabilities existing in the service running on the target IP address, and the existing vulnerabilities are identified by standard vulnerability numbers. Associating standardized vulnerability attribute labels with vulnerabilities with standard vulnerability numbers to generate vulnerability attribute facts; Establish a direct mapping relationship between vulnerabilities associated with standardized vulnerability attribute labels and MSF exploit module paths to generate MSF exploit module mapping facts; The network service facts, vulnerability existence facts, vulnerability attribute facts and MSF utilization module mapping facts are combined, translated and compressed to form a compressed vulnerability data set, and the compressed vulnerability data set is represented by standardized predicate logic or key-value pair set.
2. The method for translating and compressing original vulnerability data based on a large language model according to claim 1, characterized in that: The preset prompt template is specifically: Extract the target IP address, running service name, protocol type, port number, service running permissions, standard vulnerability number, vulnerability attribute category and vulnerability attribute value from the original vulnerability data and output: networkSI (target IP address, running service name, protocol type, port number and service running permissions); vulE (target IP address, standard vulnerability number, running service name); vulP (standard vulnerability number, vulnerability attribute category, vulnerability attribute value).
3. The method for translating and compressing original vulnerability data based on a large language model according to claim 2, characterized in that: The representation of the network service fact is: networkSI (target IP address, running service name, protocol type, port number and service running permissions).
4. The method for translating and compressing original vulnerability data based on a large language model according to claim 2, characterized in that: The fact that the vulnerability exists is expressed as follows: vulE (target IP address, standard vulnerability number, running service name).
5. The method for translating and compressing original vulnerability data based on a large language model according to claim 2, characterized in that: The vulnerability attribute fact is represented as follows: vulP (standard vulnerability number, vulnerability attribute category, vulnerability attribute value).
6. The large language model-based original vulnerability data translation and compression method according to claim 1, characterized in that: The MSF utilizes the module mapping fact to express the form: mapExploit (standard vulnerability number, exploit module path).
7. The method for translating and compressing original vulnerability data based on a large language model according to claim 1, characterized in that: The compressed vulnerability dataset is stored in text, JSON, or XML format. When stored in text format, each row represents a fact.
8. A device for translating and compressing original vulnerability data based on a large language model, characterized in that: include: An extraction module is used to extract structured facts from raw vulnerability data using a large language model based on a preset prompt template. The structured facts include network service facts and vulnerability existence facts. The network service facts describe the target IP address, the name of the running service, the protocol type, the port number, and the service running permissions. The vulnerability existence facts describe the vulnerabilities existing in the service running on the target IP address and identify the existing vulnerabilities using standard vulnerability numbers. An association module is used to associate standardized vulnerability attribute labels with vulnerabilities with standard vulnerability numbers to generate vulnerability attribute facts; A mapping module is used to establish a direct mapping relationship between vulnerabilities associated with standardized vulnerability attribute labels and MSF exploit module paths, generating MSF exploit module mapping facts; The compression module is used to combine the network service facts, vulnerability existence facts, vulnerability attribute facts and MSF utilization module mapping facts, translate and compress them to form a compressed vulnerability data set, and the compressed vulnerability data set is represented by standardized predicate logic or key-value pair set.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the processor implements the original vulnerability data translation and compression method based on a large language model as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for translating and compressing original vulnerability data based on a large language model according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Security vulnerability scanning and publishing system and method based on Nmap and Metasploit
CN113221124A
Method and device for automatically mapping vulnerabilities to attack techniques and tactics based on large language model
CN118368103A
Vulnerability identification and load matching method based on large pre-training language model
CN118606950A
Method for Threat Control in a Computer Network Security System
US20200007560A1
Cited By
Multi-source vulnerability scanning result detection method and device based on AI
CN121479780A