A defense counterattack method and system for large language model extraction attacks
By using user behavior analysis and dynamic adjustment methods, user groups are segmented and targeted processing is carried out, solving the defense problem against large language model extraction attacks and achieving effective defense against malicious users while ensuring a normal user experience.
Patent Information
- Application Number
- CN202511308790.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-09-15
AI Technical Summary
Existing technologies are insufficient to effectively defend against extraction attacks on large language models, especially when the normal user experience is not affected. Furthermore, existing methods often fail to accurately distinguish user types, resulting in a lack of targeted and proactive protection strategies.
By segmenting user groups through user behavior analysis, and employing non-replacement interference processing, probabilistic content replacement interference processing, and reverse attack processing, the user groups are dynamically adjusted. Combined with periodic detection processing, this ensures normal interaction for authorized users and effective defense against malicious users.
It improves the effectiveness and accuracy of attack defense and counterattack against large language model extraction, avoids misjudgment, enhances anti-interference and self-correction capabilities, and ensures that normal user experience is not affected.
Smart Images

Figure CN120834959B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of security protection and counterattack against large language model attacks, and in particular to a defensive counterattack method and system for large language model extraction attacks. Background Technology
[0002] With the rapid development of large language model technology, its application in various industries is becoming increasingly widespread. However, large language models face a variety of security threats, especially extraction attacks. Attackers obtain a large number of input-output pairs of large language models and analyze the model's structure, parameters, or sensitive information in the training data, posing a serious threat to the model's security and privacy. Once these input-output pairs are obtained and used by attackers, it may lead to serious consequences such as leakage of user privacy and leakage of confidential information. Therefore, protective measures against large language model extraction attacks are essential.
[0003] In the current technology, there is a relative lack of protective measures against large language model extraction attacks. Traditional protection methods mostly focus on access control or encrypted transmission, which are difficult to effectively deal with situations where attackers obtain input-output pairs through legitimate interactions. Moreover, existing methods either affect the user experience of normal users or fail to fundamentally prevent the input-output pairs from being exploited by attackers, making it difficult to strike a balance between protecting data security and ensuring user experience. In addition, some existing data protection technologies, such as simple content replacement or encryption, often disrupt the semantic coherence of input and output content, causing normal users to be unable to accurately understand the model's output results, further affecting the user experience. At the same time, existing technologies lack specificity, applying the same protection strategy to all users, which often brings unnecessary trouble to normal users. Furthermore, existing technologies are lacking in proactive response, focusing more on passive defense and not giving sufficient consideration to interfering with the attacker's extraction behavior.
[0004] Therefore, how to design a defense and counterattack method against large language model extraction attacks, so as to avoid excessive impact on normal user interaction and effectively resist extraction attacks, has become an urgent problem to be solved. Summary of the Invention
[0005] Based on this, the present invention proposes a defense and counterattack method and system for large language model extraction attacks. It segments user groups through user behavior analysis to accurately match response strategies, avoiding impact on the user experience of normal users. Then, through targeted content processing strategies, it ensures that trusted users receive complete and accurate model responses, guaranteeing their user experience is unaffected while effectively suppressing attack behavior from suspicious users and providing accurate and effective defense and counterattack against malicious users. Furthermore, it dynamically updates the user group to avoid misjudgments, enhancing anti-interference and self-correction capabilities. Finally, through continuous detection and protection throughout the entire cycle, it further improves the reliability of protection. The present invention improves the effectiveness and accuracy of defense and counterattack against large language model extraction attacks.
[0006] This invention proposes a defensive countermeasure method against large language model extraction attacks, comprising:
[0007] Collect the initial interaction behavior data of the current target user and build a user behavior analysis system to divide the user groups. The user behavior analysis system includes user behavior information, user input content, user behavior analysis functions and user behavior analysis results. The user groups include trusted users, suspicious users and malicious users.
[0008] User interaction content processing is performed based on the user group, including non-replacement interference processing, probabilistic content replacement interference processing, and reverse attack processing.
[0009] The user interaction behavior data after the user interaction content is processed is obtained, so as to dynamically adjust the processing of the user interaction content based on the user interaction behavior data. The dynamic adjustment includes updating the user group and replacing interference processing to maintain the user interaction content.
[0010] Update the historical behavior information of the current target user and perform intent and domain analysis to complete the periodic detection process.
[0011] In summary, this method for defending against large language model extraction attacks utilizes user behavior analysis to segment user groups and precisely match response strategies, avoiding impact on the user experience of legitimate users. Furthermore, targeted content processing strategies ensure that trusted users receive complete and accurate model responses, guaranteeing an unaffected user experience while effectively suppressing attacks from suspicious users and providing accurate and effective defense against malicious users. Dynamic adjustments to the user group update prevent misjudgments, enhancing anti-interference and self-correction capabilities. Finally, continuous detection and protection throughout the entire lifecycle further improves the reliability of the defense. This invention enhances the effectiveness and accuracy of defending against large language model extraction attacks. Specifically, the system collects initial user interaction data from the current target user and constructs a user behavior analysis system to segment user groups. This system includes user behavior information, user input content, user behavior analysis functions, and user behavior analysis results. The user groups include trusted users, suspicious users, and malicious users to accurately match response strategies and avoid impacting the user experience of normal users. User interaction content processing is performed based on the user groups, including non-replacement interference processing, probabilistic content replacement interference processing, and reverse attack processing. This ensures that trusted users receive complete and accurate model responses, guaranteeing their user experience is not affected. While responding, it effectively suppresses the attack behavior of suspicious users and provides accurate and effective defense and counterattack against malicious users. It obtains user interaction behavior data after processing user interaction content, and dynamically adjusts the processing of user interaction content based on the user interaction behavior data. The dynamic adjustment includes updating user groups and replacing interference processing to avoid misjudgment, enhances anti-interference ability and self-correction ability, updates the historical behavior information of the current target user and performs intent and domain analysis to complete periodic detection processing, further improving the reliability of protection. This invention improves the effectiveness and accuracy of defense and counterattack against attacks targeting large language model extraction.
[0012] Furthermore, the step of collecting the initial user interaction data of the current target user and constructing a user behavior analysis system to segment user groups specifically includes:
[0013] Obtain the initial user interaction behavior data of the current target user. This initial user interaction behavior data includes multiple user interaction behaviors before any replacement or interference processing. Construct a user behavior analysis system based on this initial user interaction behavior data. The specific algorithm for this user behavior analysis system is as follows:
[0014] ,
[0015] ,
[0016] ,
[0017] ,
[0018] ,
[0019] in, This refers to a user behavior analysis system. Represents user behavior information, This indicates the user's input. This represents a user behavior analysis function. This indicates the results of user behavior analysis. , This represents the total number of the user's initial interaction actions.
[0020] Each user input has a unique corresponding user behavior analysis result;
[0021] Based on the user behavior analysis results, user groups are segmented, when... When this happens, the current target user is determined to be a credit user. If so, the current target user is determined to be a suspicious user. If so, the current target user is determined to be a malicious user. Indicates the ordinal number of the user interaction behavior.
[0022] Furthermore, the step of processing user interaction content based on the user group specifically includes:
[0023] After obtaining the current user group, process the user interaction content;
[0024] If the current target user is determined to be a trusted user, non-replacement interference processing is performed; if the current target user is determined to be a suspicious user, probabilistic content replacement interference processing is performed; if the current target user is determined to be a malicious user, reverse attack processing is performed.
[0025] The non-replacement interference processing marks the current target user's user interaction behavior as a normal usage requirement, receives all user interaction behaviors of the current target user, and generates the original output content.
[0026] The probabilistic content replacement interference processing performs content replacement based on user behavior analysis results, user inquiry volume, and the proportion of large language model parameters. The specific algorithm for the probabilistic content replacement interference processing is as follows:
[0027] ,
[0028] ,
[0029] ,
[0030] in, Represents a probability function. The control parameter representing the rate of change of the probability function in the vertical direction. , The control parameter representing the rate of change of the probability function in the horizontal direction. Indicates the threshold for triggering conditions. This represents the user credibility evaluation function. The control parameter represents the range of the independent variable of the probability function. Indicates the first User behavior analysis results corresponding to each user interaction behavior. Indicates the preceding The number of queries per user interaction. , Indicates the weighting coefficient. This represents the total number of parameters in a large language model. ;
[0031] A random decision is made based on the probability function of the probabilistic content replacement interference processing to either return the original output content or perform content replacement, wherein the probability of returning the original output content is... The probability of the content being replaced is ;
[0032] The content replacement is based on a content replacement degree algorithm, which is as follows:
[0033] ,
[0034] in, Indicates the degree of content replacement. This indicates the content replacement degree adjustment coefficient. , Indicates the first User behavior analysis results corresponding to each user interaction behavior;
[0035] The content replacement degree is incorporated into the content replacement prompt word template, which is then input into the content replacement model. This model performs replacement processing on the output content to generate candidate processing results. A semantic consistency detection model then calculates the semantic consistency index between the candidate processing results and the original output content. If the semantic consistency index is less than a first consistency threshold, the replacement process is repeated. If the semantic consistency index is greater than or equal to the first consistency threshold, the candidate processing results are returned to the current target user. The first consistency threshold is... .
[0036] Furthermore, the step of performing reverse attack processing if the current target user is determined to be a malicious user specifically includes:
[0037] If the current target user is determined to be a malicious user, a reverse attack is performed. The reverse attack is based on the degree of forced content replacement, and the specific algorithm for the degree of forced content replacement is as follows:
[0038] ,
[0039] in, Indicates the degree of forced content replacement. This represents the adversarial sample strength coefficient, 0.6 ≤ ≤0.9, Indicates the first User behavior analysis results corresponding to each user interaction behavior;
[0040] The input and output content of the current target user are forcibly replaced. The degree of forced replacement is incorporated into the adversarial example generation prompt template. This template is then input into the content replacement model, which generates adversarial examples. The semantic consistency detection model calculates the semantic consistency index between the adversarial examples and the original output content. If the semantic consistency index is less than a second consistency threshold, the replacement process is repeated. If the semantic consistency index is greater than or equal to the second consistency threshold, the adversarial example is returned to the current target user. The second consistency threshold is... .
[0041] Furthermore, the step of acquiring user interaction behavior data after processing user interaction content, and dynamically adjusting the processing of user interaction content based on the user interaction behavior data, specifically includes:
[0042] Acquire user secondary interaction behavior data after processing user interaction content, and re-divide user groups based on the user secondary interaction behavior data;
[0043] If the current target user is reclassified as a trusted user, then the user interaction behavior of the current target user is marked as normal usage needs, all user interaction behaviors of the current target user are received and the original output content is generated.
[0044] If the current target user is reclassified as a suspicious user, then probabilistic content replacement interference processing will be carried out based on the user behavior analysis results, the number of user inquiries, and the proportion of large language model parameters.
[0045] Furthermore, the step of re-segmenting user groups based on the user's secondary interaction behavior data further includes:
[0046] If the current target user is reclassified as a malicious user, then user behavior analysis will no longer be performed on the current target user's user interaction behavior and the malicious user determination will be maintained until the current target user terminates the interaction.
[0047] While maintaining the malicious user determination, the system continuously performs reverse attack processing on the current target user, forcibly replacing the input and output content of the current target user, generating adversarial samples, and returning them to the current target user.
[0048] Furthermore, the step of updating the historical behavior information of the current target user and performing intent and domain analysis to complete the periodic detection process specifically includes:
[0049] If the current target user is determined to be a trusted user or a suspicious user, update the user behavior information, match historical interaction data and network device status information, perform intent and domain analysis on the input content of the current target user, and continuously update the user group for the current target user based on the intent and domain analysis results until the user terminates the interaction.
[0050] This invention proposes a defense and counterattack system against large language model extraction attacks, comprising:
[0051] The analysis module is used to collect the user's first interaction behavior data of the current target user and build a user behavior analysis system to divide user groups. The user behavior analysis system includes user behavior information, user input content, user behavior analysis functions and user behavior analysis results. The user groups include trusted users, suspicious users and malicious users.
[0052] The content processing module is used to process user interaction content according to the user group. The user interaction content processing includes non-replacement interference processing, probabilistic content replacement interference processing, and reverse attack processing.
[0053] The dynamic adjustment module is used to acquire user secondary interaction behavior data after user interaction content processing, so as to dynamically adjust the user interaction content processing according to the user secondary interaction behavior data. The dynamic adjustment includes user group updates and replacement interference processing to maintain.
[0054] The update module is used to update the historical behavior information of the current target user and perform intent and domain analysis to complete the periodic detection process.
[0055] The present invention also provides a storage medium that stores one or more programs, which, when executed by a processor, implement the defense and counterattack method against large language model extraction attacks as described above.
[0056] The present invention also provides a computer device, the computer device including a memory and a processor, wherein:
[0057] The memory is used to store computer programs;
[0058] When the processor executes the computer program stored in the memory, it implements the defense and counterattack method against large language model extraction attacks as described above. Attached Figure Description
[0059] Figure 1 The flowchart shows the defensive counterattack method against large language model extraction attacks proposed in the first embodiment of the present invention.
[0060] Figure 2 This is a schematic diagram of the defensive counterattack system against large language model extraction attacks proposed in the second embodiment of the present invention.
[0061] The following detailed description, in conjunction with the accompanying drawings, will further illustrate the present invention. Detailed Implementation
[0062] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Several embodiments of the invention are illustrated in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete.
[0063] It should be noted that when a component is said to be "fixed to" another component, it can be directly on the other component or there may be an intervening component. When a component is said to be "connected to" another component, it can be directly connected to the other component or there may be an intervening component. The terms "vertical," "horizontal," "left," "right," and similar expressions used in this document are for illustrative purposes only.
[0064] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0065] Please see Figure 1 The diagram shows a flowchart of a defense and counterattack method against large language model extraction attacks proposed in the first embodiment of the present invention. This defense and counterattack method against large language model extraction attacks includes steps S01 to S04, wherein:
[0066] Step S01: Collect the initial user interaction data of the current target user and build a user behavior analysis system to segment user groups;
[0067] It should be noted that in this embodiment, the user behavior analysis system includes user behavior information, user input content, user behavior analysis functions, and user behavior analysis results. The user group includes trusted users, suspicious users, and malicious users. The system acquires the initial interaction behavior data of the current target user, which includes multiple user interaction behaviors before any replacement or interference processing. The user behavior analysis system is constructed based on this initial interaction behavior data. The specific algorithm of the user behavior analysis system is as follows:
[0068] ,
[0069] ,
[0070] ,
[0071] ,
[0072] ,
[0073] in, This refers to a user behavior analysis system. Represents user behavior information, This indicates the user's input. This represents a user behavior analysis function. This indicates the results of user behavior analysis. , This represents the total number of the user's initial interaction actions.
[0074] Each user input has a unique corresponding user behavior analysis result;
[0075] Based on the user behavior analysis results, user groups are segmented, when... When this happens, the current target user is determined to be a credit user. If so, the current target user is determined to be a suspicious user. If so, the current target user is determined to be a malicious user. Indicates the ordinal number of the user interaction behavior.
[0076] Step S02: Process user interaction content according to user groups;
[0077] It should be noted that in this embodiment, the user interaction content processing includes non-replacement interference processing, probabilistic content replacement interference processing, and reverse attack processing. After obtaining the current user group, user interaction content processing is performed.
[0078] If the current target user is determined to be a trusted user, non-replacement interference processing is performed; if the current target user is determined to be a suspicious user, probabilistic content replacement interference processing is performed; if the current target user is determined to be a malicious user, reverse attack processing is performed.
[0079] The non-replacement interference processing marks the current target user's user interaction behavior as a normal usage requirement, receives all user interaction behaviors of the current target user, and generates the original output content.
[0080] The probabilistic content replacement interference processing performs content replacement based on user behavior analysis results, user inquiry volume, and the proportion of large language model parameters. The specific algorithm for the probabilistic content replacement interference processing is as follows:
[0081] ,
[0082] ,
[0083] ,
[0084] in, Represents a probability function. The control parameter representing the rate of change of the probability function in the vertical direction. , The control parameter representing the rate of change of the probability function in the horizontal direction. Indicates the threshold for triggering conditions. This represents the user credibility evaluation function. The control parameter represents the range of the independent variable of the probability function. Indicates the first User behavior analysis results corresponding to each user interaction behavior. Indicates the preceding The number of queries per user interaction. , Indicates the weighting coefficient. This represents the total number of parameters in a large language model. ;
[0085] A random decision is made based on the probability function of the probabilistic content replacement interference processing to either return the original output content or perform content replacement, wherein the probability of returning the original output content is... The probability of the content being replaced is ;
[0086] The content replacement is based on a content replacement degree algorithm, which is as follows:
[0087] ,
[0088] in, Indicates the degree of content replacement. This indicates the content replacement degree adjustment coefficient. , Indicates the first User behavior analysis results corresponding to each user interaction behavior;
[0089] The content replacement degree is incorporated into the content replacement prompt word template, which is then input into the content replacement model. This model performs replacement processing on the output content to generate candidate processing results. A semantic consistency detection model then calculates the semantic consistency index between the candidate processing results and the original output content. If the semantic consistency index is less than a first consistency threshold, the replacement process is repeated. If the semantic consistency index is greater than or equal to the first consistency threshold, the candidate processing results are returned to the current target user. The first consistency threshold is... ;
[0090] If the current target user is determined to be a malicious user, a reverse attack is performed. The reverse attack is based on the degree of forced content replacement, and the specific algorithm for the degree of forced content replacement is as follows:
[0091] ,
[0092] in, Indicates the degree of forced content replacement. This represents the adversarial sample strength coefficient, 0.6 ≤ ≤0.9, Indicates the first User behavior analysis results corresponding to each user interaction behavior;
[0093] The input and output content of the current target user are forcibly replaced. The degree of forced replacement is incorporated into the adversarial example generation prompt template. This template is then input into the content replacement model, which generates adversarial examples. The semantic consistency detection model calculates the semantic consistency index between the adversarial examples and the original output content. If the semantic consistency index is less than a second consistency threshold, the replacement process is repeated. If the semantic consistency index is greater than or equal to the second consistency threshold, the adversarial example is returned to the current target user. The second consistency threshold is... .
[0094] Step S03: Obtain user secondary interaction behavior data after processing user interaction content, so as to dynamically adjust the processing of user interaction content based on the user secondary interaction behavior data;
[0095] It should be noted that in this embodiment, the dynamic adjustment includes user group update and replacement interference processing maintenance, obtaining user secondary interaction behavior data after user interaction content processing, and re-dividing user groups based on the user secondary interaction behavior data;
[0096] If the current target user is reclassified as a trusted user, then the user interaction behavior of the current target user is marked as normal usage needs, all user interaction behaviors of the current target user are received and the original output content is generated.
[0097] If the current target user is reclassified as a suspicious user, then probabilistic content replacement interference processing will be carried out based on the user behavior analysis results, the number of user inquiries, and the proportion of parameters of the large language model.
[0098] If the current target user is reclassified as a malicious user, then user behavior analysis will no longer be performed on the current target user's user interaction behavior and the malicious user determination will be maintained until the current target user terminates the interaction.
[0099] While maintaining the malicious user determination, the system continuously performs reverse attack processing on the current target user, forcibly replacing the input and output content of the current target user, generating adversarial samples, and returning them to the current target user.
[0100] Step S04: Update the historical behavior information of the current target user and perform intent and domain analysis to complete the periodic detection process;
[0101] It should be noted that in this embodiment, if the current target user is determined to be a trusted user or a suspicious user, the user behavior information is updated, historical interaction data and network device status information are matched, the intent and domain analysis of the current target user's input content is performed, and the user group is continuously updated based on the intent and domain analysis results until the user terminates the interaction.
[0102] In summary, this method for defending against large language model extraction attacks utilizes user behavior analysis to segment user groups and precisely match response strategies, avoiding impact on the user experience of legitimate users. Furthermore, targeted content processing strategies ensure that trusted users receive complete and accurate model responses, guaranteeing an unaffected user experience while effectively suppressing attacks from suspicious users and providing accurate and effective defense against malicious users. Dynamic adjustments to the user group update prevent misjudgments, enhancing anti-interference and self-correction capabilities. Finally, continuous detection and protection throughout the entire lifecycle further improves the reliability of the defense. This invention enhances the effectiveness and accuracy of defending against large language model extraction attacks. Specifically, the system collects initial user interaction data from the current target user and constructs a user behavior analysis system to segment user groups. This system includes user behavior information, user input content, user behavior analysis functions, and user behavior analysis results. The user groups include trusted users, suspicious users, and malicious users to accurately match response strategies and avoid impacting the user experience of normal users. User interaction content processing is performed based on the user groups, including non-replacement interference processing, probabilistic content replacement interference processing, and reverse attack processing. This ensures that trusted users receive complete and accurate model responses, guaranteeing their user experience is not affected. While responding, it effectively suppresses the attack behavior of suspicious users and provides accurate and effective defense and counterattack against malicious users. It obtains user interaction behavior data after processing user interaction content, and dynamically adjusts the processing of user interaction content based on the user interaction behavior data. The dynamic adjustment includes updating user groups and replacing interference processing to avoid misjudgment, enhances anti-interference ability and self-correction ability, updates the historical behavior information of the current target user and performs intent and domain analysis to complete periodic detection processing, further improving the reliability of protection. This invention improves the effectiveness and accuracy of defense and counterattack against attacks targeting large language model extraction.
[0103] Please see Figure 2 The diagram shows a structural schematic of a defense and counterattack system against large language model extraction attacks proposed in the second embodiment of the present invention. The system includes:
[0104] Analysis module 10 is used to collect the user's first interaction behavior data of the current target user and build a user behavior analysis system to divide user groups. The user behavior analysis system includes user behavior information, user input content, user behavior analysis functions and user behavior analysis results. The user groups include trusted users, suspicious users and malicious users.
[0105] Content processing module 20 is used to process user interaction content according to the user group. The user interaction content processing includes non-replacement interference processing, probabilistic content replacement interference processing, and reverse attack processing.
[0106] The dynamic adjustment module 30 is used to acquire user secondary interaction behavior data after user interaction content processing, so as to dynamically adjust the user interaction content processing according to the user secondary interaction behavior data. The dynamic adjustment includes user group update and replacement interference processing maintenance.
[0107] The update module 40 is used to update the historical behavior information of the current target user and perform intent and domain analysis to complete the periodic detection process.
[0108] The present invention also proposes a computer storage medium storing one or more programs, which, when executed by a processor, implement the aforementioned defensive and counter-attack method against large language model extraction attacks.
[0109] The present invention also proposes a computer device, including a memory and a processor, wherein the memory is used to store computer programs, and the processor is used to execute the computer programs stored in the memory, so as to realize the above-mentioned defense and counterattack method against large language model extraction attacks.
[0110] Those skilled in the art will understand that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can mean any means that can contain stored, communicated, propagated, or transmitted programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0111] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0112] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0113] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0114] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
Claims
1. A defensive counterattack method against large language model extraction attacks, characterized in that, include: Collect the initial interaction behavior data of the current target user and build a user behavior analysis system to divide the user groups. The user behavior analysis system includes user behavior information, user input content, user behavior analysis functions and user behavior analysis results. The user groups include trusted users, suspicious users and malicious users. Each user input has a unique corresponding user behavior analysis result; User groups are segmented based on the user behavior analysis results; User interaction content processing is performed based on the user group, including non-replacement interference processing, probabilistic content replacement interference processing, and reverse attack processing. Specifically, if the current target user is determined to be a trusted user, non-replacement interference processing will be performed; If the current target user is determined to be a suspicious user, then a probabilistic content replacement interference process will be performed; The probabilistic content replacement interference processing performs content replacement based on user behavior analysis results, user inquiry volume, and the proportion of parameters in the large language model. The content replacement is based on a content replacement degree algorithm, incorporating the content replacement degree into the content replacement prompt word template. The content replacement prompt word template is then input into the large content replacement model. The large content replacement model performs replacement processing on the output content to generate candidate processing results. Then, the semantic consistency index between the candidate processing results and the original output content is calculated based on the semantic consistency detection model. If the semantic consistency index is less than a first consistency threshold, the replacement processing is repeated. If the semantic consistency index is greater than or equal to the first consistency threshold, the candidate processing result is returned to the current target user. If the current target user is determined to be a malicious user, then a reverse attack will be performed. The reverse attack processing is based on the degree of forced content replacement. The input and output content of the current target user are forcibly replaced. The degree of forced content replacement is incorporated into the adversarial sample generation prompt word template. The adversarial sample generation prompt word template is input into the content replacement model. The content replacement model generates adversarial samples. The semantic consistency detection model calculates the semantic consistency index between the adversarial sample and the original output content. If the semantic consistency index is less than the second consistency threshold, the replacement process is repeated. If the semantic consistency index is greater than or equal to the second consistency threshold, the adversarial sample is returned to the current target user. The user interaction behavior data after the user interaction content is processed is obtained, so as to dynamically adjust the processing of the user interaction content based on the user interaction behavior data. The dynamic adjustment includes updating the user group and replacing interference processing to maintain the user interaction content. Update the historical behavior information of the current target user and perform intent and domain analysis to complete the periodic detection process.
2. The defensive counterattack method against large language model extraction attacks according to claim 1, characterized in that, The steps of collecting the initial user interaction data of the current target user and constructing a user behavior analysis system to segment user groups specifically include: Obtain the initial user interaction behavior data of the current target user. This initial user interaction behavior data includes multiple user interaction behaviors before any replacement or interference processing. Construct a user behavior analysis system based on this initial user interaction behavior data. The specific algorithm for this user behavior analysis system is as follows: , , , , , in, This refers to a user behavior analysis system. Represents user behavior information, This indicates the user's input. This represents a user behavior analysis function. This indicates the results of user behavior analysis. , This represents the total number of the user's initial interaction actions. when When this happens, the current target user is determined to be a credit user. If so, the current target user is determined to be a suspicious user. If so, the current target user is determined to be a malicious user. Indicates the ordinal number of the user interaction behavior.
3. The defensive counterattack method against large language model extraction attacks according to claim 1, characterized in that, The step of processing user interaction content based on the user group specifically includes: After obtaining the current user group, process the user interaction content; The non-replacement interference processing marks the current target user's user interaction behavior as a normal usage requirement, receives all user interaction behaviors of the current target user, and generates the original output content. The specific algorithm for the probabilistic content replacement interference processing is as follows: , , , in, Represents a probability function. The control parameter representing the rate of change of the probability function in the vertical direction. , The control parameter representing the rate of change of the probability function in the horizontal direction. Indicates the threshold for triggering conditions. This represents the user credibility evaluation function. The control parameter represents the range of the independent variable of the probability function. Indicates the first User behavior analysis results corresponding to each user interaction behavior. Indicates the preceding The number of queries per user interaction. , Indicates the weighting coefficient. This represents the total number of parameters in a large language model. ; A random decision is made based on the probability function of the probabilistic content replacement interference processing to either return the original output content or perform content replacement, wherein the probability of returning the original output content is... The probability of the content being replaced is ; The algorithm for the degree of content replacement is as follows: , in, Indicates the degree of content replacement. This indicates the content replacement degree adjustment coefficient. , Indicates the first User behavior analysis results corresponding to each user interaction behavior; The first consistency threshold is .
4. The defensive counterattack method against large language model extraction attacks according to claim 1, characterized in that, The step of performing reverse attack processing if the current target user is determined to be a malicious user specifically includes: If the current target user is determined to be a malicious user, a reverse attack will be performed. The specific algorithm for the degree of forced content replacement is as follows: , in, Indicates the degree of forced content replacement. This represents the adversarial sample strength coefficient, 0.6 ≤ ≤0.9, Indicates the first User behavior analysis results corresponding to each user interaction behavior; The second consistency threshold is .
5. The defensive counterattack method against large language model extraction attacks according to claim 1, characterized in that, The step of acquiring user interaction behavior data after processing user interaction content, and dynamically adjusting the processing of user interaction content based on the user interaction behavior data, specifically includes: Acquire user secondary interaction behavior data after processing user interaction content, and re-divide user groups based on the user secondary interaction behavior data; If the current target user is reclassified as a trusted user, then the user interaction behavior of the current target user is marked as normal usage needs, all user interaction behaviors of the current target user are received and the original output content is generated. If the current target user is reclassified as a suspicious user, then probabilistic content replacement interference processing will be carried out based on the user behavior analysis results, the number of user inquiries, and the proportion of large language model parameters.
6. The defensive counterattack method against large language model extraction attacks according to claim 5, characterized in that, The step of re-segmenting user groups based on the user's secondary interaction behavior data further includes: If the current target user is reclassified as a malicious user, then user behavior analysis will no longer be performed on the current target user's user interaction behavior and the malicious user determination will be maintained until the current target user terminates the interaction. While maintaining the malicious user determination, the system continuously performs reverse attack processing on the current target user, forcibly replacing the input and output content of the current target user, generating adversarial samples, and returning them to the current target user.
7. The defensive counterattack method against large language model extraction attacks according to claim 1, characterized in that, The steps of updating the historical behavior information of the current target user and performing intent and domain analysis to complete the periodic detection process specifically include: If the current target user is determined to be a trusted user or a suspicious user, update the user behavior information, match historical interaction data and network device status information, perform intent and domain analysis on the input content of the current target user, and continuously update the user group for the current target user based on the intent and domain analysis results until the user terminates the interaction.
8. A defense and counterattack system against large language model extraction attacks, characterized in that, include: The analysis module is used to collect the user's first interaction behavior data of the current target user and build a user behavior analysis system to divide user groups. The user behavior analysis system includes user behavior information, user input content, user behavior analysis functions and user behavior analysis results. The user groups include trusted users, suspicious users and malicious users. Each user input has a unique corresponding user behavior analysis result; User groups are segmented based on the user behavior analysis results; The content processing module is used to process user interaction content according to the user group. The user interaction content processing includes non-replacement interference processing, probabilistic content replacement interference processing, and reverse attack processing. Specifically, if the current target user is determined to be a trusted user, non-replacement interference processing will be performed; If the current target user is determined to be a suspicious user, then a probabilistic content replacement interference process will be performed; The probabilistic content replacement interference processing performs content replacement based on user behavior analysis results, user inquiry volume, and the proportion of parameters in the large language model. The content replacement is based on a content replacement degree algorithm, incorporating the content replacement degree into the content replacement prompt word template. The content replacement prompt word template is then input into the large content replacement model. The large content replacement model performs replacement processing on the output content to generate candidate processing results. Then, the semantic consistency index between the candidate processing results and the original output content is calculated based on the semantic consistency detection model. If the semantic consistency index is less than a first consistency threshold, the replacement processing is repeated. If the semantic consistency index is greater than or equal to the first consistency threshold, the candidate processing result is returned to the current target user. If the current target user is determined to be a malicious user, then a reverse attack will be performed. The reverse attack processing is based on the degree of forced content replacement. The input and output content of the current target user are forcibly replaced. The degree of forced content replacement is incorporated into the adversarial sample generation prompt word template. The adversarial sample generation prompt word template is input into the content replacement model. The content replacement model generates adversarial samples. The semantic consistency detection model calculates the semantic consistency index between the adversarial sample and the original output content. If the semantic consistency index is less than the second consistency threshold, the replacement process is repeated. If the semantic consistency index is greater than or equal to the second consistency threshold, the adversarial sample is returned to the current target user. The dynamic adjustment module is used to acquire user secondary interaction behavior data after user interaction content processing, so as to dynamically adjust the user interaction content processing according to the user secondary interaction behavior data. The dynamic adjustment includes user group updates and replacement interference processing to maintain. The update module is used to update the historical behavior information of the current target user and perform intent and domain analysis to complete the periodic detection process.
9. A storage medium, characterized in that, The storage medium stores one or more programs that, when executed by a processor, implement the defensive counterattack method against large language model extraction attacks as described in any one of claims 1-7.
10. A computer device, characterized in that, The computer device includes a memory and a processor, wherein: The memory is used to store computer programs; When the processor executes the computer program stored in the memory, it implements the defense and counterattack method against large language model extraction attacks as described in any one of claims 1-7.
Citation Information
Patent Citations
APT attack detection method based on large language model
CN120389886A
Electric power honeypot construction method and system based on large language model
CN120498898A