A honey point-based bait deployment method and deployment system
By performing text and webpage similarity analysis on the original data of honeypots and combining it with a large language model to generate highly deceptive baits, the problems of insufficient user data security, traceability and countermeasure capabilities of honeypots and honeypots have been solved. Adaptive camouflage and intelligent deployment have been achieved, improving the defense effect of honeypots.
Patent Information
- Application Number
- CN202511483518.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-10-17
AI Technical Summary
In existing active deception defense technologies, honeypots are insufficient to guarantee user data security, honey spots lack traceability and countermeasure capabilities, and honey baits lack adaptive camouflage capabilities and automated delivery mechanisms.
By acquiring native data from honey spots, text and webpage similarity analysis is performed using webpage parsing libraries and business scenario tag libraries. Combined with a large language model, highly realistic and sweet bait is generated and dynamically deployed in honey spots to determine the bait placement location, achieving adaptive camouflage and intelligent deployment.
It enhances the deception effect and deterrent power of honeypots, strengthens the attacker's ability to trace and counterattack, ensures user data security, and does not affect the normal operation of business.
Smart Images

Figure CN120956536B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer network security, and in particular to a honey bait deployment method and system based on honey spots. Background Technology
[0002] Proactive deception defense technology is a network security protection method whose core objective is to detect threats in real time and take corresponding defensive measures when attackers attempt to deceive or intrude into a system. This technology effectively confuses, delays, or blocks potential attack behaviors by deploying fictitious information and setting traps.
[0003] Common proactive deception defense techniques include honeypots, honey baits, and honey points. However, existing technologies have several shortcomings: traditional honeypot systems typically rely heavily on manual configuration and maintenance, making it difficult to dynamically adjust service content and posing a risk of being exploited by attackers. Furthermore, achieving highly realistic honeypots requires collecting large amounts of user traffic and interaction data, raising privacy risks. Honey point technology employs a passive detection mechanism, lacking the ability to track and counterattack the source of attacks. Traditional honey baits are usually static, singular, and fixed deception units, lacking adaptive camouflage capabilities and automated deployment mechanisms. Therefore, there is an urgent need to provide a solution to address these issues. Summary of the Invention
[0004] The purpose of this invention is to provide a honeypot-based bait deployment method that can improve the problems in existing active deception defense technologies, such as the difficulty of honeypots in ensuring user data security, the lack of traceability and countermeasure capabilities of honeypots, and the lack of adaptive camouflage capabilities and automated deployment mechanisms for honeypots.
[0005] In a first aspect, the present invention provides a method for deploying honey bait based on honey spots, comprising:
[0006] Acquire native data from Midian, obtain Midian webpage text based on webpage parsing library, extract keywords from the Midian webpage text to obtain webpage text keywords, and perform text similarity analysis on the webpage text keywords based on business scenario tag library to obtain text similarity;
[0007] Construct a honey spot scene tag library, and perform webpage similarity analysis on the honey spot native data based on the honey spot scene tag library to obtain webpage similarity;
[0008] The text similarity and the webpage similarity are fused to obtain the business scenario result for scene recognition;
[0009] The DOM tree of the original honey spot data is divided into nodes or node combinations. The weight value of the node or node combination is determined based on a predetermined rule. The node or node combination is sorted based on the weight value to determine the honey bait placement position.
[0010] Using a large language model, based on the business scenario results and the bait placement location, a mimicry of high-sweetness bait is obtained; a new honey spot is generated based on the original honey spot data and the mimicry of high-sweetness bait.
[0011] This invention provides a method for deploying bait based on honeypots. It involves parsing native honeypot data to obtain webpage text content and extracting keywords, calculating text similarity, analyzing the native honeypot data using a constructed honeypot scenario tag library, calculating webpage similarity, and weighted fusion of the two similarities to obtain the business scenario result. The DOM tree of the native honeypot data is divided into nodes or node combinations, weighted and sorted according to predetermined rules to determine the bait deployment location. A large language model is used to obtain mimicking high-sweetness bait based on the business scenario result and the bait deployment location, generating new honeypots containing mimicking high-sweetness bait, thus achieving proactive deception defense through high-sweetness induction, adaptive mimicry, and intelligent deployment.
[0012] Optionally, obtaining native data from Honeypoint includes: obtaining native data from Honeypoint through web crawlers, wherein the native data from Honeypoint includes HTML text, CSS styles, and JavaScript code.
[0013] Optionally, when obtaining the honeypot webpage text based on the webpage parsing library, the process includes: obtaining the webpage text content from the honeypot's native data based on the webpage parsing library, and preprocessing the webpage text content, wherein the preprocessing includes filtering stop words, filtering punctuation marks, and filtering word segmentation.
[0014] Optionally, when extracting keywords from the honeypot webpage text to obtain text similarity, the process includes: extracting keywords from the honeypot webpage text content using a keyword extraction method, performing text similarity analysis between the keywords and a business scenario tag library to obtain text similarity; the keyword extraction method includes a word frequency-inverse document frequency algorithm and a machine learning model; the business scenario tag library includes an open-source tag system, an industry keyword library, a pre-trained word vector model, a text classification model, an open-source knowledge graph, and a custom tag library.
[0015] Optionally, a honeypot scene tag library is constructed, and webpage similarity analysis is performed on the honeypot native data based on the honeypot scene tag library to obtain webpage similarity. This includes: associating business scene tags with the honeypot native data to obtain honeypot scene tags, constructing a honeypot scene tag library based on the honeypot scene tags, and performing webpage similarity analysis between the honeypot native data and the honeypot scene tags to obtain webpage similarity.
[0016] Optionally, when fusing the text similarity and the webpage similarity to obtain the business scenario result, the method includes: performing weighted fusion of the text similarity and the webpage similarity based on a fusion scenario analyzer to obtain the business scenario result.
[0017] Optionally, the predetermined rules include: nodes or node combinations at the top of the DOM tree have higher weights; nodes or node combinations at a shallower level in the DOM tree have higher weights; nodes or node combinations with semantic information have higher weights; nodes or node combinations containing links have higher weights; and nodes or node combinations with user interaction have higher weights.
[0018] Optionally, when using a large language model to obtain the mimicry of high-sweetness bait based on the business scenario results and the bait placement location, the process includes:
[0019] Input the business scenario results and the bait placement location into the prompt word template of the large language model;
[0020] Based on the business scenario results and the original data of the Honey Point, generate Honey Bait elements that are consistent with the appearance and function of the Honey Point webpage; generate isomorphic nodes or node combinations based on the nodes or node combinations corresponding to the Honey Point placement locations; construct a real file structure and implant fictitious information based on the basic templates in the Honey Bait library and the business scenario results.
[0021] Based on the honey bait elements, the isomorphic nodes or node combinations, and the real file structure, the large language model outputs a mimicry high-sweetness bait.
[0022] The real file structure includes file header information, metadata, and body content; the fictitious information includes sensitive data of real users and key enterprise information; and the basic template includes Excel documents, PDF reports, and configuration files.
[0023] Optionally, when generating a new honeypot based on the original honeypot data and the simulated high-sweetness bait, the process includes: loading the DOM tree of the original honeypot data based on a web page parsing library, locating the node or node combination corresponding to the honeypot placement position and dynamically inserting the simulated high-sweetness bait, and performing style and behavior consistency verification on the inserted simulated high-sweetness bait; serializing the modified DOM tree into HTML code and merging CSS and JavaScript resources to generate a new honeypot; deploying the new honeypot to a preset network environment and starting a monitoring mechanism to record all access behaviors to the simulated high-sweetness bait; the dynamic insertion includes replacing the original node, inserting child nodes within the original node, and inserting sibling nodes after the original node.
[0024] Secondly, the present invention provides a honey bait deployment system based on honey spots, comprising:
[0025] The environment perception module is used to acquire native data from the honeypot, obtain honeypot webpage text based on a webpage parsing library, extract keywords from the honeypot webpage text to obtain webpage text keywords, perform text similarity analysis on the webpage text keywords based on a business scenario tag library to obtain text similarity; construct a honeypot scene tag library, perform webpage similarity analysis on the honeypot native data based on the honeypot scene tag library to obtain webpage similarity; and fuse the text similarity and the webpage similarity to obtain business scenario results for scene recognition.
[0026] The mimicry and camouflage module is used to obtain mimicry high-sweetness bait by using a large language model based on the business scenario results of the environment perception module and the bait delivery location of the intelligent deployment module.
[0027] The intelligent deployment module is used to divide the DOM tree of the original honey spot data into nodes or node combinations, determine the weight value of the nodes or node combinations based on predetermined rules, sort the nodes or node combinations based on the weight values, and determine the honey bait placement position; and generate new honey spots based on the original honey spot data of the environmental perception module and the mimicry high-sweetness bait of the mimicry camouflage module. Attached Figure Description
[0028] Figure 1 A flowchart of a honey bait deployment method based on honey spots provided in an embodiment of the present invention;
[0029] Figure 2 A structural diagram of a honey bait deployment system based on honey spots provided in an embodiment of the present invention;
[0030] Figure 3 A flowchart of an environmental perception module in a honey bait deployment system based on honey spots, provided as an embodiment of the present invention;
[0031] Figure 4 A flowchart of a mimicry camouflage module in a honey bait deployment system based on honey spots, provided as an embodiment of the present invention;
[0032] Figure 5 This is a flowchart of an intelligent deployment module in a honey bait deployment system based on honey spots, provided as an embodiment of the present invention. Detailed Implementation
[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Unless otherwise defined, the technical or scientific terms used herein should have the ordinary meaning understood by those skilled in the art to which this invention pertains.
[0034] See Figure 1 This invention provides a method for deploying honey bait based on honey spots, comprising the following steps:
[0035] S1. Obtain the original data of Midian, obtain the Midian web page text based on the web page parsing library, extract the keywords of the Midian web page text to obtain the web page text keywords, and perform text similarity analysis on the web page text keywords based on the business scenario tag library to obtain the text similarity;
[0036] S2. Construct a Honey Point scene tag library, and perform webpage similarity analysis on Honey Point's native data based on the Honey Point scene tag library to obtain webpage similarity;
[0037] S3. Merge text similarity and webpage similarity to identify business scenarios and obtain business scenario results;
[0038] S4. Divide the DOM tree of the original honey spot data into nodes or node combinations, determine the weight value of the nodes or node combinations based on predetermined rules, sort the nodes or node combinations in descending order based on the weight values, and determine the honey bait placement position.
[0039] S5. Using a large language model, based on business scenario results and bait placement locations, obtain simulated high-sweetness bait; generate new honey spots based on original honey spot data and simulated high-sweetness bait.
[0040] In fact, the honey bait deployment method provided by this invention first acquires the original data of the honey spots, obtains the web page text of the honey spots based on the web page parsing library, extracts keywords from the web page text, and performs similarity analysis based on the business scenario tag library to obtain text similarity; then, it constructs a honey spot scenario tag library, analyzes the original data of the honey spots, and obtains web page similarity; it then weights and fuses the two similarity results to obtain the business scenario result; next, it divides the DOM tree of the original data of the honey spots into nodes or node organizations, assigns weights and sorts them according to predetermined rules, and determines the honey bait placement position; finally, it uses a large model language to obtain mimicry high-sweetness bait based on the business scenario result and the honey bait placement position, thus generating new honey spots containing mimicry high-sweetness bait, achieving the purpose of proactive deception defense with high sweetness induction, adaptive mimicry, and intelligent deployment.
[0041] In some embodiments, in step S1, firstly, raw data of the honeypot is collected through a web crawler; then, the web page text content is extracted using a web page parsing library, and the extracted web page text content is preprocessed; next, keywords of the web page text content are extracted using a keyword extraction method; finally, the extracted web page text keywords are compared with the text similarity of mainstream open-source business scenario tag libraries to obtain the similarity result, i.e., the text similarity. The original data for Honeypoint can be HTML text, CSS styles, and JavaScript code; preprocessing can include filtering words, filtering punctuation marks, and filtering word segmentation; keyword extraction methods can be the Term Frequency-Inverse Document Frequency (TF-IDF) algorithm and methods based on machine learning models.
[0042] Specifically, web page parsing libraries are mainly used to extract structured text content from raw data such as HTML, CSS, and JavaScript, and to analyze DOM nodes in malicious web pages. Web page parsing libraries can include BeautifulSoup, lxml, Jsoup, HtmlAgilityPack, Cheerio, Selenium, and Playwright / Puppeteer. BeautifulSoup is an example. In Python, BeautifulSoup is commonly used due to its concise syntax and support for multiple parsers (such as lxml and html.parser). For higher performance or XPath support, lxml, which offers faster parsing speed and is suitable for handling large HTML documents, is often chosen. In Java, Jsoup is frequently used for optimized HTML parsing and supports CSS selectors. On the .NET platform, HtmlAgilityPack is commonly used, as it supports XPath. In Node.js, Cheerio excels at server-side HTML parsing with its jQuery-like syntax, while Playwright / Puppeteer are modern browser automation tools that can be used to obtain the complete rendered DOM structure and accurately extract dynamic content. Selenium can simulate browser behavior and parse dynamically loaded web page content in multilingual environments.
[0043] Specifically, business scenario tag libraries are used for scenario classification or similarity calculation of web page text content. These libraries can be open-source tag systems, industry keyword libraries, pre-trained word vector models, text classification models, open-source knowledge graphs, and custom tag libraries. Open-source tag systems include the MITRE ATT&CK framework, which provides attack and tactical classification tags suitable for security scenario identification; industry keyword libraries include custom industry thesaurus (such as finance, healthcare, education, etc.), which are sets of keywords built according to different business domains; pre-trained word vector models include Word2Vec, GloVe, and FastText, which support semantic similarity calculation through word vector representation; text classification models include BERT, RoBERTa, and Sentence-BERT, which are based on the Transformer architecture and suitable for fine-grained text classification and similarity calculation; open-source knowledge graphs include Wikidata and DBpedia, which enhance scenario understanding capabilities through entity links and semantic tags; and custom tag libraries built according to actual business needs, such as self-built business scenario tag libraries containing tags like "login page," "payment interface," and "user management," to improve classification accuracy and interpretability.
[0044] The combination of a webpage parsing library and a business scenario tag library enables intelligent understanding of honeypot webpages and precise delivery of honey bait.
[0045] In some embodiments, in step S2, the publicly available raw data of Midian is first manually labeled with business scenario tags to obtain corresponding Midian scenario tags, and a Midian scenario tag library is constructed based on these tags. Then, a webpage similarity analysis is performed between the raw Midian data and the Midian webpages labeled in the Midian scenario tags to calculate the similarity between the webpages, i.e., the webpage similarity score. .
[0046] In some embodiments, when obtaining the business scenario result in step S3, the final business scenario result is calculated by a fusion scenario analyzer. This analyzer performs text similarity analysis. Similarity to web pages Assign preset weights respectively and Then, a weighted fusion calculation is performed to obtain the result for the business scenario. .
[0047] Honeypoint is a security protection system built on simulation and trap technologies, designed to monitor and protect target systems. Its core feature is that it does not require complete reproduction of all the functions of the protected system, but can effectively block attackers from intruding by detecting anomalies and suspicious attack behaviors, while ensuring that legitimate users' normal use is not affected. Honeypoint technology has low resource consumption, supports large-scale batch deployment, and is a highly efficient and lightweight security solution. Unlike honeypot technology, honeypoint only uses a passive detection mechanism and does not actively induce attackers, thus having an extremely low false positive rate.
[0048] However, honeypot technology has a significant shortcoming in current applications: a lack of ability to trace and counter attackers. Although honeypots can effectively block further intrusion into protected systems, they do not provide sufficient mechanisms to trace the source of attacks or implement countermeasures, thus limiting their effectiveness in proactive defense.
[0049] To enhance the deceptive effectiveness and deterrent power of honeypots, this invention proposes deploying honey bait—highly attractive data or files—within honeypots to lure attackers into accessing or downloading them. Once an attacker triggers the honey bait, their identity and behavior are exposed, leading to capture by the defender. However, the current challenge lies in the lack of effective technology to link honeypots and honey baits, specifically how to automatically deploy matching honey baits based on the characteristics of different honeypots to enhance the targeting and success rate of deception, ultimately achieving the goals of attack tracing and countermeasures. Furthermore, honeypots are diverse, potentially simulating different systems or applications with varying functions and characteristics. The design and deployment of honey baits must be tailored to specific contexts to ensure their rationality and consistency, reducing the risk of being detected by attackers. To address this issue, this invention calculates the importance of honeypot DOM tree nodes to determine the optimal placement of honey baits and embeds them into the honeypots, thereby achieving the goal of attack tracing and countermeasures based on honeypots.
[0050] In some embodiments, in step S4, the DOM tree of the original honeypot data is first divided into several nodes or node combinations. Then, the weight value of each node or node combination in the DOM tree is determined according to predetermined rules. Finally, the nodes or node combinations are sorted according to their weight values to determine the optimal honey bait placement position. The predetermined rules are as follows:
[0051] Node position: The position of a node or combination of nodes in the DOM tree affects its importance. Generally, elements located higher in the DOM tree (such as near the root node) have higher weight.
[0052] Node depth: The depth of a node in the DOM tree is negatively correlated with its weight. A greater depth generally indicates lower importance, and therefore the weight decreases as depth increases.
[0053] Node type: Predefined weights based on the semantic importance of the node labels. For example, <header> 、 <nav> 、 <article>Equal semantic labels are generally considered more valuable and thus assigned a higher weight;
[0054] Node linking: nodes or combinations of nodes that contain links (especially links to key pages) are assigned a higher weight;
[0055] User interaction: if a node has user interaction behavior (such as clicking, hovering, etc.), it is assigned a higher weight.
[0056] Bait is an extension of the honeypot technology. It is not only an information resource, but also an entity or resource carrier that actively lures illegal intruders. It usually contains bait-like digital data that can be used to track the behavior of attackers, such as fake email addresses, user accounts, database information, and fake programs. Since bait is a resource that normal users will not access, any access to it can be considered as a potential attack activity.
[0057] Existing bait generation methods are mostly designed for traditional honeypots and usually rely on a variety of fake files prepared in advance and widely deployed on all types of hosts. However, these methods have poor applicability in actual applications and are prone to mismatch between bait and actual environment, even causing logical conflicts, thereby reducing security protection performance. Therefore, the present application generates bait based on business scenario results and bait placement positions combined with large language model prompt word templates. This method can achieve the adaptive camouflage ability of bait for business scenarios, effectively improving its deception and concealment.
[0058] In some embodiments, when the mimetic high-sweet bait is obtained in step S5, the business scenario results and bait placement positions are input into a preset large language model prompt word template, which guides the model to perform the following generation logic:
[0059] Based on the business scenario results and the original data of the honeypot (such as HTML structure, CSS style, and JavaScript logic), the appearance and function of the bait are designed to be consistent with other elements of the honeypot webpage. For example, if the honeypot is an e-commerce page, the bait may be a "special offer link" or a "coupon popup"; if the honeypot is an office system, the bait may be an "employee directory" or a "financial report link";
[0060] Based on the node or node combination information to which the bait placement position belongs, the same node or node combination is generated. For example, if the placement position is If the tag is used, a realistic download link will be generated; if the placement is... <form> Then a false login form is generated;
[0061] Based on the base template selected from the bait library and the business scenario result, a real file structure is constructed, and fictitious information is added therein for simulating sensitive data of real users or enterprise key information. The real file structure includes file header information, metadata, and text content, the fictitious information includes sensitive data of real users and enterprise key information, and the base template includes an Excel document, a PDF report, and a configuration file;
[0062] The file header information is used for simulating a MIME type, a creation time, an author, and the like of a real file; the metadata can be document attributes, version information, and editing history; the false sensitive data (such as employee accounts, internal IPs, API keys, and the like) is implanted in the text content; and the enterprise key information can be a fake company internal document, a project plan, and a contract draft;
[0063] Finally, the content output by the large language model is the mimic high-sweet bait, and the bait content includes HTML / CSS / JS code required for front-end display; file entities (such as false documents and configuration files) required for a back end; metadata description; and bait triggering logic. The mimic high-sweet bait generation scheme proposed in the present application not only does not affect normal business operation, but also does not need to use real user data traffic, effectively guarantees user data security, and the generated bait has high inducibility to attackers, that is, it is "high-sweet"; at the same time, the bait has self-adaptive camouflage capability, can be dynamically adjusted according to a business scenario, and is naturally integrated into a honey point, so that a "mimic" effect is achieved.
[0064] Specifically, the large language model (LLM; Large Language Models, LLMs) involved in the present application refers to a large-scale natural language processing model based on deep learning, which usually has a parameter quantity of hundreds of billions or even more and is trained on a large amount of text data. The core goal of such models (such as GPT-3, PaLM, LLaMA, etc.) is to understand and generate natural language, and they can predict subsequent words or generate semantically coherent text based on context. The GPT series (Generative Pre-trained Transformer) proposed by OpenAI is a typical representative among them, especially the currently recognized most powerful GPT-4 model, which has a training corpus size of tens of billions of words. From the perspective of practical application, large language models are widely used in question answering, text generation, programming assistance, and machine translation. From the perspective of technical principles, LLMs are generally built on the Transformer architecture and achieve performance breakthroughs through substantial expansion of model size, pre-training data volume, and computing resources. Through the pre-training process, the model learns language patterns such as grammar, syntax, semantic ambiguity, and masters certain common sense and world knowledge. Overall, LLMs rely on the self-attention mechanism in Transformer, combine pre-training and generation strategies, and achieve understanding and synthesis of human natural language.
[0065] Specifically, the prompt is a text input used to guide the large language model to perform specific tasks or generate required content, which can include questions, instructions, context, input data, or output requirements, etc. to clarify the behavior and generation direction of the model. The method of designing and optimizing the prompt is called "prompt engineering", and this emerging field is committed to improving the task execution effect and safety of large language models.
[0066] The prompt template involved in the present application is a structured prompt construction method that provides a pre-set text format containing several placeholders that will be replaced by actual parameters in specific applications. The template can effectively organize the structure and variables of the prompt, guiding the template to generate results that meet the expected output format and content.
[0067] In some embodiments, when generating a new honeypot based on the honeypot native data and the metamorphic high-sweet bait in step S5, first, load the DOM tree of the honeypot native data based on a web page parsing library (such as BeautifulSoup, jsdom, etc.), and locate to the node or node combination corresponding to the bait deployment position; second, dynamically insert the metamorphic high-sweet bait (such as the generated HTML fragment, file download link, form, etc.) into the specified deployment position in the DOM tree, and the dynamic insertion includes replacing the original node, inserting a child node in the original node, and inserting a sibling node after the original node; then, perform style and behavior consistency checking on the inserted metamorphic high-sweet bait to ensure that the inserted bait is consistent with the original page in style (CSS), interactive behavior (JavaScript), and semantic structure, and avoid being recognized by attackers due to inconsistent styles; after that, serialize the modified DOM tree into HTML code again, merge CSS resources and JavaScript resources, and generate a new honeypot; finally, deploy the new honeypot to the preset network environment, and start the monitoring mechanism to record all access behaviors to the metamorphic high-sweet bait for subsequent attack tracing and countermeasure analysis.
[0068] In summary, the honeypot-based bait deployment method proposed by the present application, on the basis of a lightweight honeypot simulation system, deeply integrates bait technology that can be adaptively deployed, significantly enhancing the ability to trace and counterattack attack behavior. This scheme does not affect the normal operation of user business, nor does it need to provide user data traffic, fundamentally guaranteeing the security of user data. By combining business scenarios to dynamically adjust bait strategies, adaptive camouflage and automated deployment are achieved, effectively improving the pertinence and concealment of deception defense. Active deception defense with high-sweet induction, self-adaptive metamorphism, and intelligent deployment is achieved.
[0069] Referring to Figure 2 The application provides a honey point-based bait deployment system, comprising the following steps:
[0070] An environment perception module is configured to obtain honey point native data, obtain honey point webpage text based on a webpage parsing library, extract keywords of the honey point webpage text to obtain webpage text keywords, perform text similarity analysis on the webpage text keywords based on a business scenario label library to obtain a text similarity, construct a honey point scenario label library, perform webpage similarity analysis on the honey point native data based on the honey point scenario label library to obtain a webpage similarity, and perform fusion scenario recognition on the text similarity and the webpage similarity to obtain a business scenario result.
[0071] A mimicry camouflage module is configured to obtain mimicry high-sweet bait based on the business scenario result of the environment perception module and the bait deployment position of the intelligent deployment module by using a large language model.
[0072] An intelligent deployment module is configured to divide a DOM tree of honey point native data into nodes or node combinations, determine weight values of the nodes or node combinations based on predetermined rules, sort the nodes or node combinations based on the weight values, determine a bait deployment position, and generate a new honey point based on the honey point native data of the environment perception module and the mimicry high-sweet bait of the mimicry camouflage module.
[0073] Referring to Figure 3 The application embodiment provides a flowchart of an environment perception module in a honey point-based bait deployment system, and specifically comprises the following steps:
[0074] Honey point native data is obtained through a network crawler, webpage text content is obtained by parsing the honey point native data through a webpage parsing library, and the webpage text content is preprocessed (such as filtering words, filtering punctuation marks and filtering segmented words), keyword extraction is performed on the webpage text content, and text similarity analysis is performed on the business scenario label library to obtain a text similarity; a honey point scenario label library is constructed, webpage similarity analysis is performed on the honey point native data to obtain a webpage similarity; and finally, fusion scenario recognition is performed on the text similarity and the webpage similarity to calculate a business scenario result.
[0075] Referring to Figure 4 The application embodiment provides a flowchart of a mimicry camouflage module in a honey point-based bait deployment system, and specifically comprises the following steps:
[0076] The business scenario result and the bait deployment position are received, the business scenario result and the bait deployment position are filled into a large language model prompt word template, the large language model generates fictitious information according to the content of the prompt word template and a bait library and simulates a real file, and finally the result output by the model is a mimicry high-sweet bait.
[0077] Referring to Figure 5 A flow chart of an intelligent deployment module in a honeypot-based bait deployment system provided by an embodiment of the present application, specifically comprising:
[0078] Receiving honeypot native data and mimic high-sweet bait, splitting the DOM tree of the honeypot native data into nodes or node combinations, then determining the weight value of the nodes or node combinations in the DOM tree according to predetermined rules, then sorting the importance of the nodes or node combinations according to the weight value, determining the bait deployment position, and finally embedding the mimic high-sweet bait into the DOM tree of the honeypot webpage to generate a new honeypot.
[0079] It should be understood that the size of the sequence number of each process described above does not mean the order of execution in various embodiments of the present application. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0080] Although the embodiments of the present application have been described in detail above, it is obvious to those skilled in the art that various modifications and changes can be made to these embodiments. However, it should be understood that such modifications and changes all fall within the scope and spirit of the present application described in the claims. Moreover, the present application described herein can have other embodiments and can be implemented or realized in various ways.< / form> < / article> < / nav> < / header>
Claims
1. A method for deploying honey bait based on honey spots, characterized in that, include: Acquire native data from Midian, obtain Midian webpage text based on webpage parsing library, extract keywords from the Midian webpage text to obtain webpage text keywords, and perform text similarity analysis on the webpage text keywords based on business scenario tag library to obtain text similarity; Construct a honey spot scene tag library, and perform webpage similarity analysis on the honey spot native data based on the honey spot scene tag library to obtain webpage similarity; The text similarity and the webpage similarity are fused to obtain the business scenario result for scene recognition; The DOM tree of the original honey spot data is divided into nodes or node combinations. The weight value of the node or node combination is determined based on a predetermined rule. The node or node combination is sorted based on the weight value to determine the honey bait placement position. Based on the business scenario results and the bait placement location, a large language model is used to obtain a mimicry of high-sweetness bait; The process of generating new honey spots based on the original honey spot data and the simulated high-sweetness bait includes: inputting the business scenario results and the honey spot placement location into a large language model prompt word template; generating honey spot elements with the same appearance and function as the honey spot webpage based on the business scenario results and the original honey spot data; generating isomorphic nodes or node combinations based on the nodes or node combinations corresponding to the honey spot placement location; constructing a real file structure and embedding fictitious information based on the basic templates in the honey spot library and the business scenario results; and outputting simulated high-sweetness bait based on the honey spot elements, the isomorphic nodes or node combinations, and the real file structure. The real file structure includes file header information, metadata, and body content; the fictitious information includes sensitive data of real users and key enterprise information; and the basic templates include Excel documents, PDF reports, and configuration files.
2. The deployment method as described in claim 1, characterized in that, Obtaining native data from Honeypoint includes: obtaining native data from Honeypoint through web crawlers, wherein the native data from Honeypoint includes HTML text, CSS styles, and JavaScript code.
3. The deployment method as described in claim 1, characterized in that, Obtaining the text of a honeypot webpage based on a webpage parsing library includes: obtaining the text content of the webpage from the honeypot's native data based on the webpage parsing library, and preprocessing the text content of the webpage; the preprocessing includes filtering stop words, filtering punctuation marks, and filtering word segmentation.
4. The deployment method as described in claim 1, characterized in that, Obtaining text similarity includes: extracting keywords from the honeypot webpage text using keyword extraction methods, performing text similarity analysis between the keywords and a business scenario tag library, and obtaining text similarity; the keyword extraction method includes a word frequency-inverse document frequency algorithm and a machine learning model; the business scenario tag library includes an open-source tag system, an industry keyword library, a pre-trained word vector model, a text classification model, an open-source knowledge graph, and a custom tag library.
5. The deployment method as described in claim 1, characterized in that, Construct a Honey Point scenario tag library, and perform webpage similarity analysis on the Honey Point native data based on the Honey Point scenario tag library to obtain webpage similarity. This includes: associating business scenario tags with the Honey Point native data to obtain Honey Point scenario tags, constructing a Honey Point scenario tag library based on the Honey Point scenario tags, and performing webpage similarity analysis on the Honey Point native data and the Honey Point scenario tags to obtain webpage similarity.
6. The deployment method as described in claim 1, characterized in that, The process of fusing the text similarity and the webpage similarity to identify the business scenario results includes: performing weighted fusion of the text similarity and the webpage similarity based on a fusion scenario analyzer to obtain the business scenario results.
7. The deployment method as described in claim 1, characterized in that, The predetermined rules include: nodes or node combinations at the top of the DOM tree have higher weights; nodes or node combinations at a shallower level in the DOM tree have higher weights; nodes or node combinations with semantic information have higher weights; nodes or node combinations containing links have higher weights; and nodes or node combinations with user interaction have higher weights.
8. The deployment method as described in claim 1, characterized in that, The process of generating a new honeypot based on the original honeypot data and the simulated high-sweetness bait includes: loading the DOM tree of the original honeypot data based on a web page parsing library, locating the node or node combination corresponding to the honeypot placement position and dynamically inserting the simulated high-sweetness bait, and performing style and behavior consistency verification on the inserted simulated high-sweetness bait; serializing the modified DOM tree into HTML code and merging CSS and JavaScript resources to generate a new honeypot; deploying the new honeypot to a preset network environment and starting a monitoring mechanism to record all access behaviors to the simulated high-sweetness bait; the dynamic insertion includes replacing the original node, inserting child nodes within the original node, and inserting sibling nodes after the original node.
9. A honey bait deployment system based on honey spots, used to implement the method as described in any one of claims 1-8, characterized in that, include: The environment perception module is used to acquire native data from the honeypot, obtain honeypot webpage text based on a webpage parsing library, extract keywords from the honeypot webpage text to obtain webpage text keywords, perform text similarity analysis on the webpage text keywords based on a business scenario tag library to obtain text similarity; construct a honeypot scene tag library, perform webpage similarity analysis on the honeypot native data based on the honeypot scene tag library to obtain webpage similarity; and fuse the text similarity and the webpage similarity to obtain business scenario results for scene recognition. The mimicry and camouflage module is used to obtain mimicry high-sweetness bait by using a large language model based on the business scenario results of the environment perception module and the bait delivery location of the intelligent deployment module. The intelligent deployment module is used to divide the DOM tree of the honey spot native data into nodes or node combinations, determine the weight value of the node or node combination based on predetermined rules, sort the node or node combination based on the weight value, and determine the honey bait placement position. New honey spots are generated based on the original honey spot data from the environmental perception module and the mimicry high-sweetness bait from the mimicry camouflage module.
Citation Information
Patent Citations
Honeycomb vulnerability generation method based on large language model
CN117610026A
Method and system for automatically deploying honey spots based on service scene matching
CN120128391A