Honeypot protection method, apparatus, device, and medium
By desensitizing and fine-tuning the large language model, highly realistic response traffic is generated, solving the problems of high honeypot construction costs and easy detection, and achieving low-cost, highly realistic and efficient honeypot protection.
Patent Information
- Application Number
- CN202311459528.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-03
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2043-11-03
AI Technical Summary
Existing security defense honeypots require a lot of manual configuration during the setup process, resulting in high deployment costs and easy vulnerability to attackers.
By using a large language model to anonymize and fine-tune real network traffic data, highly realistic response traffic is generated to simulate normal business responses, attracting attackers to launch in-depth attacks.
It reduces the labor costs of honeypot construction, improves simulation effects, enhances the ability to lure attackers, and has high versatility and continuous update capabilities.
Smart Images

Figure CN119324789B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of artificial intelligence and information security, and more particularly to a honeypot protection method, device, equipment, medium and program product. BACKGROUND
[0002] In today's digital age, network security has become an integral part of organizations and businesses. As network threats continue to evolve, security experts must take innovative approaches to protect critical systems and data. Network security defense honeypots are virtual or physical security traps designed to simulate real systems, services, or applications to attract potential attackers. Their design purpose is to collect valuable information about attack techniques, help organizations better understand the threat landscape, detect potential attacks in advance, and improve security defense strategies.
[0003] The inventors found that the existing security defense honeypot has the following defects in the process of implementing the present inventive concept: the honeypot needs a lot of manual planning and development and configuration to achieve a high-business-simulation honeypot system during the construction process, the deployment cost is extremely high, and the manual configuration process will bring unnatural system behavior patterns, so the honeypot is also easy to be recognized by attackers. SUMMARY
[0004] In view of the above problems, the present disclosure provides a honeypot protection method, device, equipment, medium and program product that can generate high-simulation response traffic using a language large model.
[0005] In a first aspect of the embodiments of the present disclosure, a honeypot protection method is provided. The method comprises: when an attack request to a target website is listened to, determining the type of a resource requested by the attack request to obtain a target resource type, wherein the type of the resource is divided into a text type and a non-text type; when the target resource type is the text type, extracting target index information of the requested resource from the attack request; querying a prompt word from a prompt word library based on the target index information to obtain a target prompt word; providing a target question sentence obtained based on the target prompt word to a language large model for honeypot protection, and obtaining generation data of the language large model for honeypot protection, wherein the language large model for honeypot protection is obtained by fine-tuning a general language large model using desensitization-processed data of real network traffic data in the target website; and generating a response request to the attack request based on the generation data.
[0006] According to an embodiment of the present disclosure, the determining the type of the requested resource of the attack request comprises: classifying the attack request into one of three categories of code request, data request and other request, wherein when the attack request is classified into the code request or the data request, the target resource type is determined as a text type. The code request requests a code resource, and the data request requests a data resource in a natural language.
[0007] According to an embodiment of the present disclosure, the extracting the target index information of the requested resource from the attack request when the target resource type is the text type comprises: when the attack request is classified into the code request, extracting a code resource name from the attack request based on keyword matching, wherein the code resource name is the target index information.
[0008] According to an embodiment of the present disclosure, the querying a prompt word from a prompt word library based on the target index information to obtain a target prompt word comprises: querying a prompt word from a code prompt word database preset for the code request to obtain the target prompt word, wherein the prompt words in the code prompt word database are prompt words constructed based on code resources in the target website after desensitization processing.
[0009] According to an embodiment of the present disclosure, the extracting the target index information of the requested resource from the attack request when the target resource type is the text type comprises: when the attack request is classified into the data request, extracting filtering condition information of a data resource requested to be returned from the attack request; and determining a category label of the data resource requested to be returned in the attack request by matching the filtering condition information with a storage category label of a data resource in the target website to obtain a target category label; wherein the target index information comprises the target category label.
[0010] According to an embodiment of the present disclosure, the querying a prompt word from a prompt word library based on the target index information to obtain a target prompt word comprises: querying a prompt word from the data prompt word database based on the target category label to obtain the target prompt word; wherein the prompt words in the data prompt word database are constructed based on the storage category label.
[0011] According to an embodiment of the present disclosure, the determining the type of the requested resource of the attack request when the attack request is monitored comprises: classifying the attack request by using a trained classifier; or determining the type of the requested resource of the attack request by matching a preset keyword list with message information of the attack request.
[0012] According to an embodiment of the present disclosure, the training process of the language large model for honeypot protection is as follows: obtaining the real network traffic data in the target website; screening the to-be-analyzed traffic data for transmitting resources of a text type from the real network traffic data; obtaining desensitized text by performing desensitization processing on the resources transmitted in the to-be-analyzed traffic data; constructing a prompt word based on the desensitized text; and performing fine-tuning training on the general language large model by taking the prompt word and the desensitized text as training data.
[0013] According to an embodiment of the present disclosure, the to-be-analyzed traffic data is divided into code request traffic and data request traffic; wherein the resources transmitted by the code request traffic are code resources, and the resources transmitted by the data request traffic are data resources represented in a natural language. Wherein, the constructing a prompt word based on the desensitized text includes: for the code request traffic and the data request traffic, the way of constructing a prompt word is different.
[0014] According to an embodiment of the present disclosure, the constructing a prompt word based on the desensitized text includes: for the data request traffic, when constructing a prompt word, the data represented in a natural language in the desensitized text is classified according to the storage classification label of the data resources in the target website to obtain a classification result; and a corresponding prompt word is constructed based on the classification result.
[0015] According to an embodiment of the present disclosure, the constructing a prompt word based on the desensitized text includes: for the code request traffic, when constructing a prompt word, a prompt word is extracted from the code of the desensitized text.
[0016] A second aspect of the embodiments of the present disclosure provides a honeypot protection device. The honeypot protection device includes a request message processing unit, a query processing unit, a large model data generation unit, and a return message rendering unit. The request message processing unit is configured to determine the type of a resource requested by an attack request when the attack request to a target website is detected, to obtain a target resource type, and to extract target index information of the requested resource from the attack request when the target resource type is a text type; wherein the type of the resource is divided into a text type and a non-text type. The query processing unit is configured to query a prompt word from a prompt word library based on the target index information to obtain a target prompt word. The large model data generation unit is configured to provide a target question sentence obtained based on the target prompt word to a language large model for honeypot protection, and to obtain generation data of the language large model for honeypot protection; wherein the language large model for honeypot protection is data obtained by performing desensitization processing on real network traffic data in the target website, and is obtained by fine-tuning training on a general language large model. The return message rendering unit is configured to generate a response request to the attack request based on the generation data.
[0017] In a third aspect, an electronic device is provided. The electronic device includes one or more processors and memory. The memory is configured to store one or more programs. The one or more programs, when executed by the one or more processors, cause the one or more processors to perform the method described above.
[0018] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores executable instructions, which, when executed by a processor, cause the processor to perform the method described above.
[0019] In a fifth aspect, a computer program product is provided. The computer program product includes a computer program, which, when executed by a processor, implements the method described above.
[0020] The one or more embodiments described above have the following advantages or beneficial effects: the language large model can be fine-tuned and trained using the desensitized real network traffic data, and then the language large model can be used to simulate normal business responses to attack request simulation normal business responses to input request text resources, to generate false response data, thereby attracting attackers to further attack. The excellent text data processing capability of the language large model is combined with the actual application scenario of the honeypot protection, to provide high simulation response data for network attacks, and to improve the ability of the honeypot to lure attackers. BRIEF DESCRIPTION OF DRAWINGS
[0021] The above and other objects, features and advantages of the present disclosure will become more apparent from the following description when taken in conjunction with the accompanying drawings, in which:
[0022] Figure 1 An application scenario diagram of a honeypot protection method, apparatus, device, medium and program product according to an embodiment of the present disclosure is schematically shown;
[0023] Figure 2 A flowchart of a honeypot protection method according to an embodiment of the present disclosure is schematically shown;
[0024] Figure 3 A flowchart of a honeypot protection method according to another embodiment of the present disclosure is schematically shown;
[0025] Figure 4 A flowchart of fine-tuning and training a general language large model in a honeypot protection method according to an embodiment of the present disclosure is schematically shown;
[0026] Figure 5 A fine-tuning and training process of a general language large model according to another embodiment of the present disclosure is schematically shown;
[0027] Figure 6A structural block diagram of a honeypot protection apparatus according to an embodiment of the disclosure is schematically shown; and
[0028] Figure 7 A block diagram of an electronic device suitable for implementing a honeypot protection method according to an embodiment of the disclosure is schematically shown. DETAILED DESCRIPTION
[0029] Hereinafter, embodiments of the disclosure will be described with reference to the accompanying drawings. It should be understood, however, that the description which follows is merely illustrative and is not intended to limit the scope of the disclosure. In the following detailed description of embodiments of the disclosure, numerous specific details are set forth in order to provide a thorough understanding of the embodiments of the disclosure. However, it would be apparent to one skilled in the art that the embodiments of the disclosure can be practiced without these specific details. In other instances, well-known structures and functions have not been described in detail in order to avoid obscuring the concepts of the disclosure.
[0030] The terms used herein are merely used to describe specific embodiments and are not intended to limit the disclosure. The terms "include" and "have" and the like used herein indicate the presence of the described features, steps, operations, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, or components.
[0031] All terms used herein, including technical and scientific terms, have the same meanings as those generally understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having meanings consistent with the context of the present description, and should not be interpreted in an idealized or overly formal manner.
[0032] In the case of using expressions similar to "at least one of A, B, and C, etc.", it should generally be interpreted to include at least one of each of the items, unless otherwise defined (for example, "a system having at least one of A, B, and C" should include but not be limited to a system having A alone, a system having B alone, a system having C alone, a system having both A and B, a system having both A and C, a system having both B and C, and / or a system having A, B, and C, etc.).
[0033] The embodiment of the present disclosure provides a honeypot protection method and device based on a language large model, equipment, a medium and a program. According to the embodiment of the present disclosure, the language large model can be fine-tuned and trained by using desensitized real network traffic data, and then the language large model is used to simulate normal business responses to input attack requests to generate false response data, thereby attracting attackers to further attack, delaying the attack process of the attacker and recording the attack behavior for subsequent tracing and attack method research. In this way, when building a honeypot, the artificial cost is very small, the language large model can simulate the business logic without omission after fine-tuning and training, and the generated message is real and reliable. Moreover, the language large model can simulate different business scenarios after fine-tuning, and there is no need to repeat the development work when building various honeypots, and the business scenarios can be updated continuously, which has high universality.
[0034] In the technical solution of the present application, the user information (including but not limited to user personal information, user image information, user equipment information such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved are information and data authorized by the user or authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of related data comply with relevant laws, regulations and standards of relevant countries and regions, necessary security measures are taken, public order and good customs are not violated, and appropriate operation portals are provided for users to choose authorization or refusal.
[0035] Figure 1 An application scenario diagram of a honeypot protection method, device, equipment, medium and program product according to an embodiment of the present disclosure is schematically shown.
[0036] As shown in Figure 1 The application scenario includes a user terminal 11, a large model honeypot server 101, a security supervision server 103 and a traffic data server 104.
[0037] The large model honeypot server 101 is deployed with a large model honeypot module 102, and the large model honeypot module 102 can be provided with a language large model for honeypot protection. The large model honeypot server 101 can also be deployed with a honeypot website. The honeypot website is a mirror website of the target website that needs to be protected in the real network environment (such as an enterprise intranet environment), and when it is found that the user in the user terminal 11 has a suspicion of attacking the target website, the attack request sent by the user terminal 11 is routed to the honeypot website in the large model honeypot server 101.
[0038] The honeypot protection method according to the embodiments of the present disclosure can be executed by the large model honeypot server 101. Wherein, after the large model honeypot server 101 receives the request message of the user terminal 11 accessing the honeypot website, the language large model in the large model honeypot module 102 can generate high simulation response data and return the response message to the user terminal 11.
[0039] In this application scenario, the large model honeypot server 101 can also record the request log of the interaction with the user terminal 11, and send the structured request log to the security supervision server 103. The security supervision server 103 can be located in a real network environment, and can perform user behavior security analysis and requester identity tracing based on the request log.
[0040] Further, the large model honeypot server 101 can also accept the latest real network traffic data sent by the traffic data server 104 at regular intervals or irregular intervals. Wherein, the traffic data server 104 can collect real network traffic data of the target website, and send it to the large model honeypot server 101 at regular intervals or irregular intervals. The large model honeypot server 101 can use the desensitized real network traffic data to fine-tune the training of the language large model in the large model honeypot module 102 to improve the high simulation capability. In one embodiment, the traffic data server 104 can perform data cleaning, screening, desensitization processing (such as adding false information), and structured conversion on the collected real network traffic data, and then send it to the large model honeypot server 101 for fine-tuning training of the language large model. Of course, the traffic data server 104 can also directly provide the collected real network traffic data to the large model honeypot server 101, and then perform screening, desensitization processing, etc. by the large model honeypot server 101, and then used for training of the language large model.
[0041] It can be seen that in this application scenario, the large model honeypot server 101 can be isolated from the real network environment. The data interaction between the large model honeypot server 101 and the real network environment is reflected in that the large model honeypot server 101 transmits the structured request log to the security supervision server 103, and the large model honeypot server 101 obtains the real network traffic data of the target website or the traffic data after desensitization processing from the traffic data server 104. Therefore, the embodiments of the present disclosure are completely isolated from the real network environment (for example, the intranet) when deploying the honeypot protection, which can effectively avoid the risk of being attacked and penetrated due to unclear network boundary during the process of building the honeypot.
[0042] The honeypot protection apparatus, device, medium and program product according to the embodiments of the present disclosure can be deployed in the large model honeypot server 101, for example, in the form of a large model honeypot module 102. Of course, the honeypot protection apparatus, device, medium and program product according to the embodiments of the present disclosure can also be partially deployed in the large model honeypot server 101 and partially deployed in the traffic data server 104, wherein the part deployed in the traffic data server 104 is used to construct training data for fine-tuning training of a language large model.
[0043] It should be understood that Figure 1 The number of terminal devices, networks and servers in the above system is only illustrative. According to the implementation needs, there can be any number of terminal devices, networks and servers.
[0044] It should be noted that the honeypot protection method, apparatus, device, medium and program product determined by the embodiments of the present disclosure can be used in the financial field, and can also be used in any field other than the financial field. The present disclosure does not limit the application field.
[0045] The honeypot protection method and apparatus according to the embodiments of the present disclosure will be described in detail below based on the system architecture Figure 1 described above. It should be noted that the serial numbers of the operations in the following method are only used to represent the operations for description, and should not be regarded as representing the execution sequence of the operations. Unless explicitly stated, the method does not need to be executed in the order shown.
[0046] Figure 2 The flowchart of the honeypot protection method according to an embodiment of the present disclosure is schematically shown.
[0047] As shown in Figure 2 , the honeypot protection method according to the embodiment can include operation S201 to operation S205.
[0048] In operation S201, when an attack request to a target website is listened to, the type of resource requested by the attack request is determined to obtain a target resource type.
[0049] The type of resource can be divided into a text type and a non-text type. The text type may, for example, be data expressed in a natural language and / or program language code. The non-text type may, for example, be picture, audio, video and the like.
[0050] The attack request can be classified by using a trained classifier, or the type of resource requested by the attack request can be determined by matching a preset keyword list and the message information of the attack request.
[0051] In operation S202, when the target resource type is a text type, target index information of the requested resource is extracted from the attack request. For example, information such as a keyword can be extracted according to a preset index extraction rule.
[0052] In operation S203, a prompt word is queried from a prompt word library based on the target index information, and a target prompt word is obtained.
[0053] The target prompt word is used to generate a target question sentence provided to the language large model. For example, the question sentence can be constructed by a preset sentence template, for example, the target prompt word is added to the position reserved in the preset sentence template, so as to generate a complete target question sentence.
[0054] In operation S204, the target question sentence obtained based on the target prompt word is provided to the language large model for honeypot protection, and generation data of the language large model for honeypot protection is obtained. The language large model for honeypot protection is obtained by fine-tuning a general language large model using desensitized data of real network traffic data in the target website.
[0055] The language large model for honeypot protection is a fine-tuned language large model (such as CodeBERT, LLaMA2, Code LLaMA, etc.) for generating response text (such as natural language description text data or code). In fine-tuning the language large model, traffic data for transmitting resources of a text type is selected from real network traffic data to obtain training data.
[0056] The disclosed embodiments use desensitized data of real network traffic data to fine-tune a general language large model to obtain a language large model for honeypot protection of a target website. Then, the target question sentence generated based on the target prompt word is used as the input of the language large model for honeypot protection, and the language large model generates highly simulated business return data or code file response text data, which can achieve a simulation deception effect.
[0057] In operation S205, a response request to the attack request is generated based on the generation data. The generation data is processed and rendered in the format of a return message, for example, according to different request types, the corresponding return header is added to form a return message, and then the return message is returned to the user terminal 11.
[0058] The language large model can be used to simulate normal business responses to attack request simulation of input request text resources, generate false response data, and thus attract attackers to further attack after the language large model is fine-tuned and trained by using desensitized real network flow data. In this way, the excellent text data processing capability of the language large model is combined with the actual application scenario of the honeypot protection, high-simulation response data is provided for network attacks, and the ability of the honeypot to lure attackers is improved. Moreover, text type resources usually account for a considerable part of resources provided by most websites (except video websites or audio websites), and for attack requests for text type resources, responding to the generated data of the fine-tuned language large model can effectively cope with the large number of requests of attackers, and has wide applicability.
[0059] Figure 3 A flowchart of a honeypot protection method according to another embodiment of the present disclosure is schematically shown.
[0060] As Figure 3 shown, the honeypot protection method according to this embodiment can include operation S301, operation S311~operation S312 or operation S321~operation S323, operation S303 and operation S204~operation S205.
[0061] In operation S301, the attack request is classified into one of three categories of code request, data request and other request.
[0062] Specifically, after receiving the request message, the structured request data in the request message is read, and then the request is classified, for example, into three categories of code request, data request and other request (for example, picture request, audio / video request, etc.), and a category label is attached.
[0063] The resource requested by the code request is a code resource, for example, an Html, Js, CSS code file required for request browser rendering. Code requests are mostly seen when the user terminal 11 establishes a connection with the target website or opens a new web page.
[0064] The resource requested by the data request is a data resource represented in natural language, for example, data presented in a web page for data transmission in JSON or XML. Data requests are mostly seen in the user interaction process.
[0065] In other requests, the picture request is a request for necessary picture files such as jpg, gif, ico, etc. for decorating a web page.
[0066] When the attack request is classified as a code request or a data request, the target resource type is determined as a text type. In website access, code requests and data requests are very common and have a large amount of access. Using a large language model to generate a response message for these two types of requests can enable the large language model to respond to most attack requests with high simulation.
[0067] If the attack request is classified as a code request, the target prompt is obtained through operation S311 to operation S312.
[0068] Specifically, in operation S311, when the attack request is classified as a code request, the code resource name (for example, code file name, code module name, etc.) is extracted from the attack request based on keyword matching, wherein the code resource name is the target index information. In the website database, code resources are often stored in the form of code files, code modules, code hierarchical folders, etc., and corresponding code resource names are provided for search queries. Therefore, the code resource name can be extracted as the target index information.
[0069] Then in operation S312, the prompt is queried from the code prompt database preset for the code request to obtain the target prompt. The prompts in the code prompt database are prompts constructed based on the desensitization processing of the code resources in the target website. For example, the code resource name (such as the code file name) of the request is used as input to query the corresponding prompt in the code prompt database, which can contain detailed description information for guiding the large language model to generate the corresponding code text.
[0070] If the attack request is classified as a data request, the target prompt is obtained through operation S321 to operation S323.
[0071] Specifically, in operation S321, when the attack request is classified as a data request, the filtering condition information of the data resource returned by the request is extracted from the attack request. For example, when the user requests what one-year financial products are, the filtering condition information for this request can include "one year" and "financial products". According to the filtering condition information, it can be determined what data needs to be returned by the data request.
[0072] Next, in operation S322, the class label of the data resource requested to be returned in the attack request is determined by matching the filtering condition information with the storage classification label of the data resource in the target website to obtain the target class label, wherein the target index information includes the target class label.
[0073] The data resource can be stored in the website by storing the classification label, so the category label can be extracted as index information. For example, in the mobile bank background, the bank's financial products are classified and stored according to the business categories such as loans, funds, and gold. Each business category is further divided into categories such as term, currency, yield, and risk level. In this way, the filtering conditions extracted from the attack request can be matched with the category labels to determine the target category label. For example, for "requesting a one-year financial product", the target category label matched can include one or more category labels with a term of "one year" in various category products matched with "financial products".
[0074] Next, in operation S323, based on the target category label, the prompt word is queried from the data prompt word database to obtain the target prompt word for data generation, wherein the prompt words in the data prompt word database are constructed based on the storage classification label, so that the queried prompt word is highly related or consistent with the classification label of the real data resource in the target website, so that the attacker is difficult to distinguish the authenticity of the data in the response message obtained based on the prompt word, and the purpose of luring the attacker is achieved.
[0075] In the embodiments of the present disclosure, considering that the storage and query methods of code resources and data resources in the website database are different, the corresponding code prompt word database and data prompt word database are constructed respectively, so as to facilitate the targeted query of prompt words for code requests and data requests.
[0076] After obtaining the target prompt word through the above operation S312 or operation S323, in operation S303, the target question sentence is generated based on the target prompt word.
[0077] Then in operation S204, the target question sentence is provided to the language large model for honeypot protection, and the generated data of the language large model for honeypot protection is obtained.
[0078] And in operation S205, a response request to the attack request is generated based on the generated data.
[0079] The embodiments of the present disclosure can use the language large model to generate response data for data requests and code requests. Accordingly, when fine-tuning the language large model, the data request traffic and the code request traffic in the real network traffic data can be used to train the language large model respectively. The language large model has the ability to understand natural language and text context. At the same time, the language large model pre-trained by code data also has considerable code interpretation and logical reasoning ability, and can perform the tasks of code generation and data generation.
[0080] Figure 4A flowchart of the process of fine-tuning a general language large model in the honeypot protection method of the embodiments of the present disclosure is shown.
[0081] As shown in Figure 4 The process of fine-tuning a general language large model according to the embodiments includes operations S401-S405.
[0082] In operation S401, real network traffic data in the target website is obtained. For example, the traffic data server 104 can send the real network traffic data collected in the target website to the large model honeypot module 102 at a regular time or upon request of the large honeypot server 101. The traffic data in the traffic data server 104 is artificially or automatically collected traffic data generated by accessing the target website in a real network environment.
[0083] In operation S402, the to-be-analyzed traffic data for transmitting resources of a text type is filtered from the real network traffic data.
[0084] The type of the resource requested by each request in the real network traffic data can be analyzed in a similar manner to the filtering of the resource type of the request in operation S201. The request message and the response message of the request for the resource of the text type are filtered, and the to-be-analyzed traffic data is obtained.
[0085] In one embodiment, the to-be-analyzed traffic data can be divided into code request traffic and data request traffic. For example, the requests in the real network traffic data can be classified into code requests, data requests, and other requests in the manner of the foregoing operation S301. The request message and the response message of the request classified into the code requests and the data requests constitute the to-be-analyzed traffic data. The traffic data classified into the code requests is referred to as code request traffic herein, and the traffic data classified into the data requests is referred to as data request traffic herein.
[0086] In operation S403, desensitization text is obtained by performing desensitization processing on the resources transmitted in the to-be-analyzed traffic data. For example, for the code request traffic and the data request traffic, respective regular expressions can be set to identify characters or strings that need to be replaced or covered, and desensitization processing is performed.
[0087] In operation S404, a prompt word is constructed based on the desensitization text.
[0088] In operation S405, the general language large model is fine-tuned using the training data composed of the prompt word and the desensitization text.
[0089] In some embodiments, the server background stores code resources and data resources in different ways, for example, code resources are stored and queried in the form of code files, code modules, code hierarchical structure, etc., and data resources are stored in the form of storage classification tags, which facilitate quick access and query through category tags, so that the way of constructing prompt words is different for code request traffic and data request traffic, thereby constructing different training data.
[0090] For example, for data request traffic, when constructing prompt words, the data expressed in natural language in the desensitized text is classified according to the storage classification tags of the data resources in the target website to obtain a classification result, and then based on the classification result, the corresponding prompt words are constructed. For example, the category tags in the classification result are used as prompt words, and the training data is composed of prompt words and data under each category tag to fine-tune the data generation capability of the language large model.
[0091] For example, for code request traffic, when constructing prompt words, prompt words are extracted from the code of the desensitized text. For example, the logic and process of the code can be analyzed and combed by using scripts or machine learning models, and prompt words that can describe the behavior and structure of the code are extracted from the code text of the response message, and then the training data is composed of the prompt words and the code text from which the prompt words are extracted to fine-tune the code generation capability of the language large model.
[0092] Figure 5 An exemplary fine-tuning training process of a general language large model according to another embodiment of the present disclosure is shown.
[0093] As shown in Figure 5 The fine-tuning training process can include steps S1-S6.
[0094] Firstly in S1, the cleaning and desensitization of the obtained real network traffic data of the target website can be completed by the traffic data server 104 or the large model honeypot server 101. Specifically, the original request message and response message in the collected real network traffic data are modified to remove some sensitive information (for example, unnecessary device parameter information, etc.), while adding some false information. At the same time, the traffic data is screened to select the code request traffic and data request traffic required for language large model training.
[0095] Next in S2, training data is constructed. For the code request traffic and data request flow obtained by processing in S1, corresponding training data is constructed respectively.
[0096] Next in S3, the training loss function of the model is set. For the training data, a loss function is designed to measure the difference between the answer generated by the language large model and the standard answer in the training data.
[0097] In one embodiment, the MSE (Mean Squared Error) is used as the loss function to measure the distance between the data generated by the model and the target data.
[0098] The MSE loss function formula is as follows:
[0099]
[0100] where xn and yn represent the generated response data sequence vector and the real data sequence vector. The formula calculates the square of the difference between the two vectors, and the greater the square of the difference, the greater the distance between the two data, and the greater the difference between the generated data and the training data. L represents the loss set of all training data. By calculating the average loss or the sum of the loss of the loss set, and calculating the gradient after back propagation, the model parameters can be updated.
[0101] Next, in S4, the model is iteratively trained. Specifically, a parameter efficient fine-tuning method (Low-Rank Adaptation of Large Language Models, LoRA) that consumes less computing resources can be used to periodically fine-tune the language large model to improve the simulation capability and broaden the diversity of response business data.
[0102] Then, in S5, the model performance is verified. The reserved part of the training data is used to verify the fine-tuning effect of the language large model, to verify whether the performance of the language large model meets the standard, and if it meets the standard, the iterative training is stopped.
[0103] Finally, in S6, after the training is completed, the model parameters of the language large model are updated to obtain a language large model for honeypot protection. The language large model for honeypot protection can generate corresponding code text or fake data according to the different types of prompt words when receiving the target question sentence based on the target prompt word in the foregoing operation S204.
[0104] Based on the honeypot protection method of the above embodiments, the embodiments of the present disclosure also provide a honeypot protection device. The following will be described in detail Figure 6 The device will be described in detail.
[0105] Figure 6 The structure block diagram of the honeypot protection device 600 according to the embodiments of the present disclosure is schematically shown.
[0106] As Figure 6As shown, according to some embodiments of the present disclosure, the apparatus 600 can include a request message processing unit 601, a query processing unit 602, a large model data generation unit 603, and a return message rendering unit 604. According to some other embodiments of the present disclosure, the apparatus 600 can further include a training data extraction unit 605 and a model training unit 606. The apparatus 600 can perform the operations described in the foregoing method embodiments with reference to FIGS. 2 and 3. Figures 2-5 The honeypot protection method described above.
[0107] The request message processing unit 601 is configured to, when an attack request to a target website is detected, determine the type of the requested resource of the attack request to obtain a target resource type, and when the target resource type is a text type, extract target index information of the requested resource from the attack request, wherein the type of the resource is divided into a text type and a non-text type. In an embodiment, the request message processing unit 601 can perform the operations S201 and S202 described above.
[0108] The query processing unit 602 is configured to query a prompt word from a prompt word library based on the target index information to obtain a target prompt word. In an embodiment, the query processing unit 602 can perform the operation S203 described above. Alternatively, in another embodiment, the query processing unit 602 can perform the operations S311-S312 or the operations S321-S323 described above.
[0109] In an embodiment, the request message processing unit 601 is responsible for receiving the request message sent by the user terminal 11, forming structured request data, and then sending the request data to the query processing unit 602.
[0110] For example, the request message processing unit 601 can first classify the received message, such as into three categories of code request, data request, and other requests (e.g., picture request), and label them. Then, according to the category of the request, extract the key information in the message, such as the requested file name, picture name, or requested data category, etc. Next, form the extracted request key information and request category into structured (e.g., JSON format) request data. For example, if the request is for a JavaScript code, the JSON format request data formed is as follows: {Type: “JS”, Content: { RequestURL: https: / / www.example.com / func.js, …}, Time: “2023-08-03”, …}. Finally, the structured request data can be sent to the query processing unit 602. In addition, the structured request data can also be sent to the security supervision server 103 as a security log.
[0111] The query processing unit 602 can parse the received structured request data, and use different databases to query the prompt words according to different categories in the request data.
[0112] For example, for a code request, the code file name in the request can be used as input to query the corresponding prompt word in the code prompt word database, which contains detailed description information for guiding the code large model to generate the corresponding code text. For example, for a data request, the data required in the request can be requested first. The required data is classified according to the data storage classification label in the target website backend server to obtain a target category label, and then the target category label is used to retrieve the data prompt word database to obtain a prompt word for data generation.
[0113] When the prompt word corresponding to the code request or the data request is obtained, the prompt word is used to construct a question sentence of the language large model, and the question sentence is sent to the large model data generation unit 603. For other requests, the language large model is not used, and the response information can be formed in a manner corresponding to the other requests, such as for a picture request, the picture database can be used to query the request file name and return the picture file, and then sent to the return message rendering unit 604.
[0114] The large model data generation unit 603 is configured to provide the target question sentence obtained based on the target prompt word to the language large model for honeypot protection, and obtain the generated data of the language large model for honeypot protection. The language large model for honeypot protection is obtained by fine-tuning a general language large model using desensitized data obtained by desensitizing real network traffic data in the target website, and can generate high-simulation business response information. The language large model has the ability to understand natural language and text context, and the language large model pre-trained by code data also has considerable code explanation and logical reasoning ability, and can perform code generation and data generation. In one embodiment, the large model data generation unit 603 can perform the operation S204 described above.
[0115] The return message rendering unit 604 is configured to generate a response request to the attack request based on the generated data. The return message rendering unit 604 can receive the generated data sent by the large model data generation unit 603, or the return data corresponding to other requests sent by the query processing unit 602, and then add the corresponding return header according to the different request types, and send the return message to the requester. In one embodiment, the return message rendering unit 604 can perform the operation S205 described above.
[0116] According to an embodiment of the present disclosure, the device 600 can also fine-tune train the language large model through the training data extraction unit 605 and the model training unit 606. The specific fine-tune training process can refer to the foregoing Figures 4-5 introduction.
[0117] Specifically, the training data extraction unit 605 can be configured to: obtain real network traffic data in a target website, filter out to-be-analyzed traffic data for transmitting resources of a text type from the real network traffic data, obtain desensitized text by performing desensitization processing on the resources transmitted in the to-be-analyzed traffic data, construct a prompt word based on the desensitized text, and construct training data with the prompt word and the desensitized text.
[0118] The model training unit 606 can fine-tune train the general language large model by using the training data. For example, first, use MSE (Mean Squared Error) as a loss function to measure the distance between the data generated by the model and the target data. Then, periodically fine-tune train the large model by using a LoRA model fine-tune method which consumes less computing resources to improve the simulation capability and broaden the diversity of response business data. After training, the model parameters of the language large model in the large model data generation unit 603 are updated.
[0119] According to an embodiment of the present disclosure, any multiple modules in the request message processing unit 601, the query processing unit 602, the large model data generation unit 603, the return message rendering unit 604, the training data extraction unit 605, and the model training unit 606 can be combined in one module, or any one of the modules can be split into multiple modules. Alternatively, at least part of the function of one or more of the modules can be combined with at least part of the function of other modules, and implemented in one module. According to an embodiment of the present disclosure, at least one of the request message processing unit 601, the query processing unit 602, the large model data generation unit 603, the return message rendering unit 604, the training data extraction unit 605, and the model training unit 606 can be at least partially implemented as a hardware circuit, such as an FPGA (Field Programmable Gate Array), a PLA (Programmable Logic Array), a system on chip, a system on board, a system in package, an ASIC (Application Specific Integrated Circuit), or any other reasonable way of integrating or packaging a circuit, etc. hardware or firmware, or implemented in any one of software, hardware, and firmware or in a proper combination of any of them. Alternatively, at least one of the request message processing unit 601, the query processing unit 602, the large model data generation unit 603, the return message rendering unit 604, the training data extraction unit 605, and the model training unit 606 can be at least partially implemented as a computer program module which can perform corresponding functions when the computer program module is run.
[0120] Figure 7 A block diagram of an electronic device suitable for implementing a honeypot protection method according to embodiments of the present disclosure is shown schematically.
[0121] As shown in Figure 7 The electronic device 700 according to embodiments of the present disclosure includes a processor 701 that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 702 or loaded from a storage section 708 into a random access memory (RAM) 703. The processor 701 can include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor, and / or a related chipset, and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), and so on. The processor 701 can also include an on-board memory for cache use. The processor 701 can include a single processing unit or multiple processing units for performing the various actions of the method processes according to embodiments of the present disclosure.
[0122] In the RAM 703, various programs and data required for the operation of the electronic device 700 are stored. The processor 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. The processor 701 performs various operations of the method processes according to embodiments of the present disclosure by executing the programs in the ROM 702 and / or the RAM 703. Note that the programs can also be stored in one or more memories other than the ROM 702 and the RAM 703. The processor 701 can also perform various operations of the method processes according to embodiments of the present disclosure by executing the programs stored in the one or more memories.
[0123] According to embodiments of the present disclosure, the electronic device 700 can also include an input / output (I / O) interface 705 that is also connected to the bus 704. The electronic device 700 can also include one or more of the following components connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as necessary. A removable recording medium 711 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is attached to the drive 710 as necessary, so that a computer program read therefrom is installed into the storage section 708 as necessary.
[0124] The present disclosure also provides a computer readable storage medium, which can be included in the device / apparatus / system described in the above embodiments, or can exist independently without being assembled into the device / apparatus / system. The above computer readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of the present disclosure.
[0125] According to an embodiment of the present disclosure, the computer readable storage medium can be a non-volatile computer readable storage medium, which can include, but is not limited to, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any appropriate combination thereof. In the present disclosure, the computer readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in connection with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer readable storage medium can include one or more memories of the ROM 702 and / or the RAM 703 described above and / or one or more memories other than the ROM 702 and the RAM 703.
[0126] The embodiments of the present disclosure also include a computer program product, which includes a computer program containing program codes for executing the methods shown in the flowcharts. When the computer program product is run in a computer system, the program codes are used to make the computer system implement the honeypot protection method provided by the embodiments of the present disclosure.
[0127] The above functions defined in the system / apparatus of the embodiments of the present disclosure are performed when the computer program is executed by the processor 701. According to an embodiment of the present disclosure, the system, apparatus, module, unit, etc. described above can be implemented by computer program modules.
[0128] In one embodiment, the computer program can rely on a tangible storage medium such as an optical storage device, a magnetic storage device, etc. In another embodiment, the computer program can also be transmitted, distributed, and downloaded in the form of a signal via a network medium, and be downloaded and installed through the communication part 709 and / or installed from the detachable medium 711. The program codes contained in the computer program can be transmitted by any appropriate network medium, including but not limited to wireless, wired, etc., or any appropriate combination thereof.
[0129] In such embodiments, the computer program can be downloaded and installed from the network via the communication section 709, and / or installed from the removable media 711. When the computer program is executed by the processor 701, the above-described functions defined in the system of the embodiments of the present disclosure are executed. According to the embodiments of the present disclosure, the system, device, apparatus, module, unit, and the like described above can be implemented by the computer program modules.
[0130] According to the embodiments of the present disclosure, the program code for executing the computer program provided by the embodiments of the present disclosure can be written in any combination of one or more programming languages, and specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming language, and / or assembly / machine language. The programming language includes, but is not limited to, such as Java, C++, python, "C" language, or similar programming language. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case involving a remote computing device, the remote computing device can be connected to the user computing device through any kind of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, connected to the Internet through an Internet service provider).
[0131] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams can represent a module, a segment, or a portion of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in a different order than that shown in the figures. For example, two blocks noted in succession can actually be executed substantially concurrently, or they can sometimes be executed in reverse order, depending on the functionality involved. It should also be noted that each block in the flowcharts or block diagrams, and combinations of blocks in the flowcharts or block diagrams, can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0132] Those skilled in the art can understand that the features described in various embodiments of the present disclosure and / or claims can be combined or / and integrated, even if such combinations or integrations are not explicitly described in the present disclosure. In particular, the features described in various embodiments of the present disclosure and / or claims can be combined and / or integrated in various combinations, without departing from the spirit and teachings of the present disclosure. All such combinations and / or integrations are within the scope of the present disclosure.
[0133] The above described embodiments of the present disclosure. However, these embodiments are merely for illustrative purposes, and are not intended to limit the scope of the present disclosure. Although each embodiment is described above separately, this does not mean that the measures in each embodiment cannot be advantageously used in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Those skilled in the art can make various substitutions and modifications without departing from the scope of the present disclosure, and all such substitutions and modifications shall fall within the scope of the present disclosure.
Claims
1. A honeypot protection method, comprising: when an attack request to a target website is monitored, determining a type of a resource requested by the attack request to obtain a target resource type; wherein the type of the resource is divided into a text type and a non-text type; when the target resource type is the text type, extracting target index information of the requested resource from the attack request; querying a prompt word from a prompt word library based on the target index information to obtain a target prompt word; providing a target question sentence obtained based on the target prompt word to a language large model for honeypot protection, and obtaining generation data of the language large model for honeypot protection; wherein the language large model for honeypot protection is obtained by fine-tuning a general language large model using data obtained by desensitizing real network traffic data in the target website; and generating a response request to the attack request based on the generation data.
2. The method of claim 1, wherein, The determination of the type of the resource requested by the attack request comprises: classifying the attack request into one of three categories of code request, data request and other request, wherein when the attack request is classified into a code request or a data request, the target resource type is determined to be the text type; wherein the resource requested by the code request is a code resource, and the resource requested by the data request is a data resource represented in natural language.
3. The method of claim 2, wherein, When the target resource type is the text type, the target index information of the requested resource is extracted from the attack request, comprising: when the attack request is classified into the code request, a code resource name is extracted from the attack request based on keyword matching, wherein the code resource name is the target index information.
4. The method of claim 2 or 3, wherein, The target prompt word is obtained by querying a prompt word from a code prompt word database preset for the code request, wherein the prompt words in the code prompt word database are prompt words constructed based on code resources in the target website after desensitization processing. When the target resource type is the text type, the target index information of the requested resource is extracted from the attack request, comprising:
5. The method of claim 2, wherein, when the attack request is classified into the data request, filtering condition information of a data resource requested to be returned is extracted from the attack request; and by matching the filtering condition information with a storage classification label of the data resource in the target website, a classification label of the data resource requested to be returned in the attack request is determined to obtain a target classification label; wherein the target index information comprises the target classification label. The target prompt word is obtained by querying a prompt word from a data prompt word database based on the target classification label; 6. The method of claim 5, wherein, wherein the prompt words in the data prompt word database are constructed based on the storage classification label. When an attack request to a target website is monitored, the type of the resource requested by the attack request is determined, comprising: 7. The method of claim 1, wherein, classifying the attack request by using the trained classifier; or determining the type of resource requested by the attack request by matching a preset keyword list and message information of the attack request.
8. The method of claim 1, wherein, The training process of the language large model for honeypot protection is as follows: obtaining the real network traffic data in the target website; screening the to-be-analyzed traffic data for transmitting resources of a text type from the real network traffic data; obtaining desensitized text by desensitizing the resources transmitted in the to-be-analyzed traffic data; constructing prompt words based on the desensitized text; and composing training data with the prompt words and the desensitized text to fine-tune the general language large model. The to-be-analyzed traffic data is divided into code request traffic and data request traffic; wherein the resources transmitted by the code request traffic are code resources, and the resources transmitted by the data request traffic are data resources represented in natural language; wherein the constructing prompt words based on the desensitized text comprises:
9. The method of claim 8, wherein, The way of constructing prompt words is different for the code request traffic and the data request traffic. The constructing prompt words based on the desensitized text comprises: for the data request traffic, when constructing prompt words, 10. The method of claim 9, wherein, classifying the data represented in natural language in the desensitized text according to the storage classification label of the data resources in the target website to obtain a classification result; and constructing corresponding prompt words based on the classification result. The constructing prompt words based on the desensitized text comprises:
11. The method of claim 9, wherein, For the code request traffic, when constructing prompt words, the prompt words are extracted from the code of the desensitized text.
12. A honeypot protection device, comprising: a request message processing unit configured to, when an attack request to a target website is listened to, determine a type of resource requested by the attack request to obtain a target resource type, and when the target resource type is a text type, extract target index information of the requested resource from the attack request; wherein the type of resource is divided into a text type and a non-text type; a query processing unit configured to query prompt words from a prompt word library based on the target index information to obtain target prompt words; a large model data generation unit configured to provide a target question sentence obtained based on the target prompt words to a language large model for honeypot protection, and obtain generation data of the language large model for honeypot protection; wherein the language large model for honeypot protection is data obtained by desensitizing real network traffic data in the target website, and is obtained by fine-tuning a general language large model; and a return message rendering unit configured to generate a response request to the attack request based on the generation data.
13. An electronic device, comprising: one or more processors; a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors perform the method of any one of claims 1-11. 14. A computer-readable storage medium having stored thereon executable instructions that, when executed by a processor, cause the processor to perform the method of any one of claims 1-11.
Citation Information
Patent Citations
Honeypot protection method and device, storage medium and electronic equipment
CN116800525A
Trapping mirror image generation method and device, equipment and storage medium
CN116962060A