GPT-based high-interaction honey point design method and system

By deploying GPT interactive programs in honey dots, simulating the network interface and generating simulation responses, the problem that existing honey dots cannot effectively respond to attack payloads is solved, and a low-cost and high-security high-interaction honey dot design is achieved.

CN119996025AActive Publication Date: 2025-05-13GUANGZHOU UNIVERSITY +2
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510231072.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-13
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

The existing high-interactive honey point design method cannot effectively respond to the attack payload issued by attackers, and there is a risk of honeypot escape and high deployment cost.

Method used

Using a high-interactive honey dot design method based on GPT, the network front-end scene and service information is crawled through network crawling technology, simulate the scene network interface and deploy vulnerable services, point the service to the GPT interactive program, eliminate the system call program, and encapsulate it into a Docker container to deploy it to the honey dot host. GPT is used to perform natural language processing and response generation, identify illegal instructions and malicious payloads, generate error signals and record attack instructions, and analyze attack intentions in combination with historical databases to generate simulation responses.

Benefits of technology

It significantly reduces the deployment cost and security risks of honey dots, improves the deception effect on attackers, reduces resource consumption, and reduces deployment and maintenance difficulty.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996025A_ABST
    Figure CN119996025A_ABST
Patent Text Reader

Abstract

The invention provides a GPT-based high-interaction honey spot design method and system, and relates to the technical field of large model security analysis. The GPT-based high-interaction honey point design method comprises the following steps of simulating a scene network interface based on a network front-end scene and service information, deploying services with vulnerabilities, modifying a pointed shell program to a GPT interaction program, packaging the interface and the services into a Docker container, and deploying the Docker container to a honey point host; receiving attack data of an attacker based on the honey point host, recording an attack instruction, and analyzing the intention of the attacker in combination with a historical attack database to obtain an analysis result; embedding a pre-designed cue word template based on an analysis result, inputting the pre-designed cue word template into a GPT response terminal, searching an instruction database based on a maximum approximate search RAG algorithm, generating response content, and making a response to an attack instruction based on the response content. According to the design method provided by the invention, the response is realized based on the deployed GPT server, the perceiving of an attacker is effectively avoided, and the deployment cost and the maintenance cost are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large model security analysis, and in particular to a high-interaction honey spot design method and system based on GPT. Background Art

[0002] In recent years, active deception defense has become increasingly popular as a security strategy. Compared to traditional network security measures, active deception defense aims to protect networks and systems by misleading attackers. This approach involves creating false network resources, such as fake servers, services, and data. These resources are designed to appear valuable to attackers, attracting them into a controlled environment, when in fact they are traps used by security teams to monitor and analyze attacker behavior. The key advantage of active deception defense is that it not only detects attacks, but also collects intelligence about the attacker's methods, techniques, and intentions. This information is very valuable for strengthening defenses, preventing future attacks, and understanding attackers' motivations.

[0003] Currently, the main products for active deception defense are honeypots. The working principle of commercial high-interaction honeypots is to attract attackers and collect their attack behavior data by simulating real services and interacting with attackers. However, this type of design method has high deployment costs and there is a risk of honeypot escape.

[0004] Honeypots are virtual bait devices that interact directly with attackers and are based on traditional honeypots. Honeypots are deployed in the same external network segment as the company's real assets to attract attackers to interact with them. Attackers will be lured into the honeypot and perform operations and interactions similar to those of real systems, thereby exposing their attack behaviors. Due to the low cost of honeypot deployment, it is currently impossible to effectively respond to attackers' attack payloads.

[0005] Therefore, it is urgent to develop a solution to solve the above problems. Summary of the invention

[0006] The purpose of the present invention is to provide a high-interaction honey spot design method and system based on GPT, which improves the problem that the honey spot cannot effectively respond to attack instructions.

[0007] The present invention provides a high-interaction honey spot design method based on GPT, which adopts the following technical solution:

[0008] Analyze the scenario network to be deployed, crawl the network front-end scenario and service information based on the web crawler technology, simulate the scenario network interface based on the network front-end scenario and service information and deploy the vulnerable service, modify the shell program pointed to by the vulnerable service to the GPT interactive program, remove the service internal system call program, encapsulate the interface and service into a Docker container and deploy it to the honeyspot host;

[0009] The attack data of the attacker is received from the honey spot host, and the illegal instructions and malicious payloads in the attack data are identified based on the natural language processing pattern matching algorithm, and the corresponding error signal is returned. The attack instructions are recorded and combined with the historical attack database to analyze the attacker's intention to obtain the analysis results;

[0010] Based on the analysis results, a pre-designed prompt word template is embedded and input into the GPT response terminal, the RAG algorithm based on maximum approximation search searches the instruction database and generates response content, responds to the attack instruction based on the response content, and stores the current attack instruction and response result in the historical attack database.

[0011] Optionally, simulating a scenario network based on the network front-end scenario and service information and deploying a service with a vulnerability includes:

[0012] The network front-end scene is crawled through web crawler technology, the network of the scene to be deployed is highly simulated on the interface, and the service information of the network is crawled. Based on the service information, the service version with vulnerabilities is found from the vulnerability library, and the service with vulnerabilities is deployed. The service has complete input and output and system call processes.

[0013] Optionally, the process of removing the system call program inside the service includes: going deep into the service and removing the functional modules that directly interact with the local system, including those involving file reading and writing, and instruction transmission, and only retaining the user interface and the communication functions required to interact with the attacker.

[0014] Optionally, identify illegal instructions and malicious payloads in attack data based on natural language processing pattern matching algorithms, and return corresponding error signals, including:

[0015] Based on the natural language processing pattern matching method, the instructions in the current attack data that do not conform to the Linux syntax are identified. For the instructions that cannot be identified, an error signal is returned to the GPT response terminal. After receiving the error signal, the GPT response terminal sends the error information to the attacker;

[0016] After filtering out the correct instructions, the instructions are further analyzed through GPT to screen whether there are malicious payloads. When GPT analyzes malicious payloads, it will combine the version information of the current service and the host information of the current simulated host, and return an error signal to the GPT response terminal, and assist the GPT response terminal to respond based on the error signal.

[0017] Optionally, the process of recording the attack instructions and analyzing the attacker's intentions in combination with the historical attack database to obtain the analysis results includes:

[0018] When an attack instruction from an attacker is received, the attack instruction is recorded and stored in the historical attack database. The attack instruction is analyzed in real time based on the GPT analysis terminal and the accumulated historical attack data to obtain the attacker's behavior pattern and attack intention.

[0019] Optionally, the process of embedding a pre-designed prompt word template based on the analysis result and inputting it into the GPT response terminal includes:

[0020] The received attack data is analyzed in detail to extract key behavioral features and context information. Based on the attacker's behavior pattern and attack intention, the behavioral features and context information are combined with a pre-designed prompt word template to form a complete prompt word, which is then input into the GPT response terminal.

[0021] Optionally, the process of searching the instruction database and generating response content based on the RAG algorithm of maximum approximation search includes:

[0022] The attacker's input instructions are analyzed in detail to extract the key elements and intentions. The RAG algorithm based on maximum approximation search searches the instruction vector database to find the existing instructions and response patterns that are closest to the key elements and intentions. Based on the best matching result found, combined with the current service environment and the information of the simulated host, the final response content is generated.

[0023] In a second aspect of the present invention, a high-interaction honey spot design system based on GPT is provided, comprising a service simulator module, a GPT analysis terminal module and a GPT response terminal module, wherein:

[0024] Service simulator module: configured to analyze the network of the scenario to be deployed, crawl the network front-end scenario and service information based on the network crawler technology, simulate the scenario network interface based on the network front-end scenario and service information and deploy the service with vulnerabilities, modify the shell program pointed to by the service with vulnerabilities to the GPT interactive program, remove the service internal system call program, encapsulate the interface and service into a Docker container and deploy it to the honey spot host;

[0025] GPT analysis terminal module: configured to receive the attacker's attack data based on the honeypot host, identify illegal instructions and malicious payloads in the attack data based on the natural language processing pattern matching algorithm, and return the corresponding error signal, record the attack instructions and analyze the attacker's intentions in combination with the historical attack database to obtain the analysis results;

[0026] GPT response terminal module: configured to embed a pre-designed prompt word template based on the analysis result and input it into the GPT response terminal, search the instruction database based on the RAG algorithm of maximum approximation search and generate response content, respond to the attack instruction based on the response content, and store the current attack instruction and response result in the historical attack database.

[0027] The beneficial effects of a high-interaction honey spot design method and system based on GPT provided by the present invention are: utilizing GPT's powerful natural language processing and understanding generalization capabilities, the front-end honey spot is only used for interface display, and the back-end server deployed with GPT completes autonomous response, which significantly reduces the deployment cost of the honey spot. During the deployment process, only the front-end interface and data forwarding need to be simulated, and resource consumption is small; and since after the vulnerability in the honey spot is exploited, the incoming attack data will not be actually executed, but the large language model terminal will perform simulated response, and since all the vulnerabilities of the honey spot are simulated and generated, the security risk is significantly reduced; by designing the GPT analysis terminal, malicious data threatening the GPT response terminal is eliminated, and the high simulation of the GPT response terminal during the response process is guaranteed, which greatly avoids the attacker's detection; during the deployment and maintenance process, the page and data transmission of the deployment service are mainly simulated, and the server deployed with GPT implements the response and response feedback, which greatly reduces the deployment difficulty and maintenance cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 A flow chart of the GPT-based high-interaction honey spot design method provided by the present invention;

[0029] Figure 2 This is an example diagram of the GPT-based high-interaction honey spot design method provided by the present invention. DETAILED DESCRIPTION

[0030] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein should be understood by people with general skills in the field to which the present invention belongs. "Including" and similar words used in this article mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects.

[0031] The embodiment of the present invention provides a high-interaction honey spot design method based on GPT, see Figure 1 ,include:

[0032] S1. Analyze the scenario network to be deployed, crawl the network front-end scenario and service information based on the network crawler technology, simulate the scenario network interface based on the network front-end scenario and service information and deploy the service with vulnerabilities, modify the shell program pointed to by the service with vulnerabilities to the GPT interactive program, remove the service internal system call program, encapsulate the interface and service into a Docker container and deploy it to the honeyspot host;

[0033] S2, receiving the attacker's attack data based on the honey spot host, identifying illegal instructions and malicious payloads in the attack data based on the natural language processing pattern matching algorithm, and returning the corresponding error signal, recording the attack instructions and analyzing the attacker's intentions in combination with the historical attack database to obtain the analysis results;

[0034] S3. Based on the analysis result, a pre-designed prompt word template is embedded and input into the GPT response terminal. The RAG algorithm based on the maximum approximation search searches the instruction database and generates response content. Based on the response content, the attack instruction is responded to, and the current attack instruction and response result are stored in the historical attack database.

[0035] In some embodiments, the process of executing step S1 includes:

[0036] S1.1. Analyze the network to be deployed;

[0037] S1.2. Crawl network front-end scenarios and service information based on web crawler technology;

[0038] S1.3, based on the network front-end scenario and service information, simulate the scenario network interface and deploy the service with vulnerabilities;

[0039] S1.4. Modify the shell program pointed to by the service with the vulnerability to the GPT interactive program;

[0040] S1.5. Eliminate the internal system call program of the service, encapsulate the interface and service into a Docker container and deploy it to the honeyspot host.

[0041] Specifically, in the process of executing step S1.1, analyzing the scenario network to be deployed, it includes: clarifying the deployment goals and requirements: obtaining the specific services supported by the scenario network, determining the scale of terminal users accessing the network, and the area that the network needs to cover.

[0042] Further, execute step S1.2, crawl the network front-end scene and service information based on the web crawler technology, analyze the HTML structure, CSS style and JavaScript code of the target network front-end scene and service information based on the browser's developer tools, and write a crawler script based on the website structure to store the crawled data.

[0043] Specifically, in step S1.3, based on the network front-end scenario and service information, the scenario network is simulated and the service with the vulnerability is deployed, including:

[0044] The network front-end scene is crawled through web crawler technology, the network of the scene to be deployed is highly simulated on the interface, and the service information of the network is crawled. Based on the service information, the service version with vulnerabilities is found from the vulnerability library, and the service with vulnerabilities is deployed. The service has complete input and output and system call processes.

[0045] Specifically, in executing step S1.4, the process of modifying the shell program pointed to by the service with the vulnerability to the GPT interactive program includes: determining the input information received by the GPT model, integrating the GPT API, and modifying the shell program pointed to by the service to the new GPT interactive program.

[0046] In fact, the GPT interactive program is responsible for receiving the attacker's input instructions and communicating with the GPT response terminal deployed at the back end, thereby achieving a simulated response to the attacker's input without actually executing any system commands or operations. This step ensures that even if the attacker attempts to exploit a vulnerability in the service, no actual damage will be caused to the honeypot host.

[0047] Specifically, in the process of executing step S1.5, removing the internal system call program of the service, encapsulating the interface and service into a Docker container and deploying it to the honeyspot host, it includes: going deep into the service, removing the functional modules that directly interact with the local system, including those involving file reading and writing, and instruction transmission, retaining only the user interface and the communication functions required for interacting with the attacker, and encapsulating the interface and service into a Docker container and deploying it to the honeyspot host.

[0048] In fact, by eliminating the internal system call program of the service, encapsulating the interface and service into a Docker container and deploying it to the honeypot host, not only can the security risk of the honeypot host caused by the service being attacked be avoided, but also the resource overhead can be reduced.

[0049] In some embodiments, the process of executing step S2 includes:

[0050] S2.1, receiving the attacker's attack data based on the honey spot host;

[0051] S2.2, Identify illegal instructions and malicious payloads in attack data based on natural language processing pattern matching algorithm, and return corresponding error signals;

[0052] S2.3. Record the attack instructions and analyze the attacker's intentions in combination with the historical attack database to obtain analysis results.

[0053] Specifically, the process of executing step S2.2 includes:

[0054] Based on the natural language processing pattern matching method, the instructions in the current attack data that do not conform to the Linux syntax are identified. For the instructions that cannot be identified, an error signal is returned to the GPT response terminal. After receiving the error signal, the GPT response terminal sends the error information to the attacker;

[0055] After filtering out the correct instructions, the instructions are further analyzed through GPT to screen whether there are malicious payloads. When GPT analyzes malicious payloads, it will combine the version information of the current service and the host information of the current simulated host, and return an error signal to the GPT response terminal, and assist the GPT response terminal to respond based on the error signal.

[0056] Specifically, the process of executing step S2.3 includes:

[0057] When an attack instruction from an attacker is received, the attack instruction is recorded and stored in the historical attack database. The attack instruction is analyzed in real time based on the GPT analysis terminal and the accumulated historical attack data to obtain the attacker's behavior pattern and attack intention.

[0058] In some embodiments, the process of executing step S3 includes:

[0059] S3.1, embedding a pre-designed prompt word template based on the analysis results and inputting it into the GPT response terminal;

[0060] S3.2, the RAG algorithm based on maximum approximation search searches the instruction database and generates response content;

[0061] S3.3. Respond to the attack instruction based on the response result.

[0062] Specifically, in the process of executing step S3.1, embedding the pre-designed prompt word template based on the analysis result and inputting it into the GPT response terminal includes:

[0063] The received attack data is analyzed in detail to extract key behavioral features and context information. Based on the attacker's behavior pattern and attack intention, the behavioral features and context information are combined with a pre-designed prompt word template to form a complete prompt word, which is then input into the GPT response terminal.

[0064] In fact, the prompt word template contains specific context information for different types of attack instructions, ensuring that the GPT response terminal can generate highly simulated and realistic simulated response content, and by inputting the constructed prompt words into the GPT response terminal, using its powerful natural language processing capabilities to generate the final response, while ensuring the consistency and coherence of the context throughout the process to improve the deception effect on attackers.

[0065] Specifically, in the process of executing step S3.2, searching the instruction database and generating the response content based on the RAG algorithm of maximum approximation search, the following steps are included:

[0066] The attacker's input instructions are analyzed in detail to extract the key elements and intentions. The RAG algorithm based on maximum approximation search searches the instruction vector database to find the existing instructions and response patterns that are closest to the key elements and intentions. Based on the best matching result found, combined with the current service environment and the information of the simulated host, the final response content is generated.

[0067] In fact, the instruction database contains instructions for common operating systems, services, and attack tools and their corresponding response examples. The matching through the maximum approximation search method not only considers the specific content of the attack instructions, but also combines historical attack records and contextual information to ensure that the generated responses are highly authentic and consistent.

[0068] Further, step S3.3 is executed to feed back the generated response content to the attacker through the GPT response terminal and store it in the historical attack database for subsequent analysis and improvement, thereby improving the deception ability and security of the honey spot system.

[0069] In other embodiments, see Figure 2 ,include:

[0070] Based on the service simulator, the network front-end scenario and service information are simulated and deployed to the honey spot host. The attack data of the attacker is received based on the honey spot host, and the attack data is transmitted to the GPT analysis terminal to obtain the analysis results and return them to the user. The GPT response terminal generates false response content based on the data interacted with the GPT analysis terminal, and returns the response content to the honey spot host, and responds to the attacker based on the response content.

[0071] The embodiment of the present invention provides a high-interaction honey spot design system based on GPT, including a service simulator module, a GPT analysis terminal module and a GPT response terminal module, wherein:

[0072] Service simulator module: configured to analyze the network of the scenario to be deployed, crawl the network front-end scenario and service information based on the network crawler technology, simulate the scenario network interface based on the network front-end scenario and service information and deploy the service with vulnerabilities, modify the shell program pointed to by the service with vulnerabilities to the GPT interactive program, remove the service internal system call program, encapsulate the interface and service into a Docker container and deploy it to the honey spot host;

[0073] GPT analysis terminal module: configured to receive the attacker's attack data based on the honeypot host, identify illegal instructions and malicious payloads in the attack data based on the natural language processing pattern matching algorithm, and return the corresponding error signal, record the attack instructions and analyze the attacker's intentions in combination with the historical attack database to obtain the analysis results;

[0074] GPT response terminal module: configured to embed a pre-designed prompt word template based on the analysis result and input it into the GPT response terminal, search the instruction database based on the RAG algorithm of maximum approximation search and generate response content, respond to the attack instruction based on the response content, and store the current attack instruction and response result in the historical attack database.

[0075] Although the embodiments of the present invention are described in detail above, it is obvious to those skilled in the art that various modifications and variations can be made to these embodiments. However, it should be understood that such modifications and variations are within the scope and spirit of the present invention as described in the claims. Moreover, the present invention described herein may have other embodiments and may be implemented or realized in a variety of ways.

Claims

1. A high-interaction honey spot design method based on GPT, characterized in that: include: Analyze the scenario network to be deployed, crawl the network front-end scenario and service information based on the web crawler technology, simulate the scenario network interface based on the network front-end scenario and service information and deploy the vulnerable service, modify the shell program pointed to by the vulnerable service to the GPT interactive program, remove the service internal system call program, encapsulate the interface and service into a Docker container and deploy it to the honeyspot host; The attack data of the attacker is received from the honey spot host, and the illegal instructions and malicious payloads in the attack data are identified based on the natural language processing pattern matching algorithm, and the corresponding error signal is returned. The attack instructions are recorded and combined with the historical attack database to analyze the attacker's intention to obtain the analysis results; Based on the analysis results, a pre-designed prompt word template is embedded and input into the GPT response terminal, the RAG algorithm based on maximum approximation search searches the instruction database and generates response content, responds to the attack instruction based on the response content, and stores the current attack instruction and response result in the historical attack database.

2. A high-interaction honey spot design method based on GPT according to claim 1, characterized in that: Based on the network front-end scenario and service information, simulate the scenario network and deploy services with vulnerabilities, including: The network front-end scene is crawled through web crawler technology, the network of the scene to be deployed is highly simulated on the interface, and the service information of the network is crawled. Based on the service information, the service version with vulnerabilities is found from the vulnerability library, and the service with vulnerabilities is deployed. The service has complete input and output and system call processes.

3. A high-interaction honey spot design method based on GPT according to claim 1, characterized in that: The process of removing the system call program inside the service includes: going deep into the service and removing the functional modules that directly interact with the local system, including those involving file reading and writing, and instruction transmission, and only retaining the user interface and the communication functions required to interact with the attacker.

4. The high-interaction honey spot design method based on GPT according to claim 1, characterized in that: Identify illegal instructions and malicious payloads in attack data based on natural language processing pattern matching algorithms, and return corresponding error signals, including: Based on the natural language processing pattern matching method, the instructions in the current attack data that do not conform to the Linux syntax are identified. For the instructions that cannot be identified, an error signal is returned to the GPT response terminal. After receiving the error signal, the GPT response terminal sends the error information to the attacker; After filtering out the correct instructions, the instructions are further analyzed through GPT to screen whether there are malicious payloads. When GPT analyzes malicious payloads, it will combine the version information of the current service and the host information of the current simulated host, and return an error signal to the GPT response terminal, and assist the GPT response terminal to respond based on the error signal.

5. The high-interaction honey spot design method based on GPT according to claim 1, characterized in that: The process of recording the attack instructions and analyzing the attacker's intentions in combination with the historical attack database to obtain the analysis results includes: When an attack instruction from an attacker is received, the attack instruction is recorded and stored in the historical attack database. The attack instruction is analyzed in real time based on the GPT analysis terminal and the accumulated historical attack data to obtain the attacker's behavior pattern and attack intention.

6. A high-interaction honey spot design method based on GPT according to claim 5, characterized in that: The process of embedding the pre-designed prompt word template based on the analysis results and inputting it into the GPT response terminal includes: The received attack data is analyzed in detail to extract key behavioral features and context information. Based on the attacker's behavior pattern and attack intention, the behavioral features and context information are combined with a pre-designed prompt word template to form a complete prompt word, which is then input into the GPT response terminal.

7. The high-interaction honey spot design method based on GPT according to claim 1, characterized in that: The process of searching the instruction database and generating response content based on the RAG algorithm based on maximum approximation search includes: The attacker's input instructions are analyzed in detail to extract the key elements and intentions. The RAG algorithm based on maximum approximation search searches the instruction vector database to find the existing instructions and response patterns that are closest to the key elements and intentions. Based on the best matching result found, combined with the current service environment and the information of the simulated host, the final response content is generated.

8. A high-interaction honey spot design system based on GPT, characterized in that: It includes a service simulator module, a GPT analysis terminal module and a GPT response terminal module, wherein: Service simulator module: configured to analyze the network of the scenario to be deployed, crawl the network front-end scenario and service information based on the network crawler technology, simulate the scenario network interface based on the network front-end scenario and service information and deploy the service with vulnerabilities, modify the shell program pointed to by the service with vulnerabilities to the GPT interactive program, remove the service internal system call program, encapsulate the interface and service into a Docker container and deploy it to the honey spot host; GPT analysis terminal module: configured to receive the attacker's attack data based on the honeypot host, identify illegal instructions and malicious payloads in the attack data based on the natural language processing pattern matching algorithm, and return the corresponding error signal, record the attack instructions and analyze the attacker's intentions in combination with the historical attack database to obtain the analysis results; GPT response terminal module: configured to embed a pre-designed prompt word template based on the analysis result and input it into the GPT response terminal, search the instruction database based on the RAG algorithm of maximum approximation search and generate response content, respond to the attack instruction based on the response content, and store the current attack instruction and response result in the historical attack database.

Citation Information

Patent Citations

  • ChatGPT-based internal network honey point generation system and method

    CN117155683A

  • Honeycomb vulnerability generation method based on large language model

    CN117610026A

  • Automatic generation method of attack graph interaction rule for honey point deployment

    CN118101346A

  • Malware detection system and method

    US20090222920A1

  • Dynamic Honeypot System

    US20170134405A1