Large model attack detection method and device, electronic equipment and storage medium
By performing multi-dimensional attack feature analysis on the prompt word sequence of the large model, the system identifies users' probing attack behaviors, solves the problem of insufficient security protection in the calling process of the large model, and ensures the security and stability of the model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI DOUXIANG INFORMATION TECH CO LTD
- Filing Date
- 2026-02-05
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies are unable to effectively identify and prevent risks such as jailbreak attacks and guessing attacks against large models, especially due to insufficient security protection in the model invocation process.
By acquiring the prompt word sequence of a large user input model, we perform similarity analysis, user request frequency analysis, prompt word instruction difference analysis, prompt word instruction repetition analysis, and prompt word instruction information entropy analysis to identify whether users are engaging in probing attack behavior.
It achieves accurate identification of probing attacks such as jailbreak attacks and guessing attacks, ensuring the security and stability of the large model's operation.
Smart Images

Figure CN121661366B_ABST
Abstract
Description
Large-scale attack detection methods, devices, electronic equipment, and storage media Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, electronic device, and storage medium for detecting large-scale attacks. Background Technology
[0002] With the deep integration and widespread application of large-scale models in fields such as natural language processing, image generation, and recommender systems, the cybersecurity field is facing new and severe challenges. Due to their complexity and immense computing power, large-scale models are easy targets for attackers. Attackers can precisely design input strategies to induce models to generate malicious outputs that deviate from their intended goals, severely damaging the model's credibility and stability, and potentially causing it to go out of control, posing significant data security risks to users and enterprises.
[0003] The relevant technologies mainly focus on the protection and monitoring of network boundaries, operating systems and application layers, but neglect the security protection of large model calling links, and cannot effectively identify the risks such as jailbreak attacks and guessing attacks faced by large models. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a method, apparatus, electronic device and storage medium for detecting large model attacks, so as to effectively identify probing attack behaviors such as jailbreak attacks and guessing attacks faced by large models.
[0005] In a first aspect, embodiments of the present invention provide a method for detecting large-scale model attacks. The method includes: acquiring a sequence of prompt words input by a user into a large-scale model; wherein the prompt word sequence includes several prompt word instructions arranged in chronological order of input; performing attack feature analysis on the prompt word sequence; wherein the attack feature analysis includes at least one of the following: similarity analysis of prompt word instructions, user request frequency analysis, difference analysis of prompt word instructions, repetition analysis of prompt word instructions, and information entropy analysis of prompt word instructions; and determining whether the user has engaged in probing attack behavior against the large-scale model based on the analysis results corresponding to at least one attack feature analysis.
[0006] Secondly, embodiments of the present invention provide a large model attack detection device, the device comprising: a first acquisition module, configured to acquire a prompt word sequence of a large model input by a user; wherein the prompt word sequence includes a plurality of prompt word instructions arranged in chronological order of input; a first analysis module, configured to perform attack feature analysis on the prompt word sequence; wherein the attack feature analysis includes at least one of the following: similarity analysis of prompt word instructions, user request frequency analysis, difference analysis of prompt word instructions, repetition analysis of prompt word instructions, and information entropy analysis of prompt word instructions; and a first determination module, configured to determine whether the user has engaged in probing attack behavior against the large model based on the analysis results corresponding to at least one attack feature analysis.
[0007] Thirdly, embodiments of the present invention provide an electronic device, including a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above-described large-scale attack detection method.
[0008] Fourthly, embodiments of the present invention provide a storage medium storing machine-executable instructions. When the machine-executable instructions are invoked and executed by a processor, the machine-executable instructions cause the processor to implement the aforementioned large-scale attack detection method.
[0009] The embodiments of the present invention bring the following beneficial effects:
[0010] The aforementioned large-scale model attack detection method, apparatus, electronic device, and storage medium include the following steps: acquiring a sequence of prompt words input by a user into a large-scale model; wherein the prompt word sequence includes several prompt word instructions arranged in chronological order of input; performing attack feature analysis on the prompt word sequence; wherein the attack feature analysis includes at least one of the following: similarity analysis of prompt word instructions, user request frequency analysis, difference analysis of prompt word instructions, repetition analysis of prompt word instructions, and information entropy analysis of prompt word instructions; and determining whether the user has engaged in probing attack behavior against the large-scale model based on the analysis results corresponding to at least one attack feature analysis.
[0011] This method acquires a sequence of prompt words from a large user-input model, consisting of several prompt words arranged chronologically. Targeting common attacker tactics such as batch probing, frequency bombing, semantic evasion, and redundant input, it flexibly utilizes at least one dimension—prompt word similarity analysis, user request frequency analysis, prompt word difference analysis, prompt word repetition analysis, and prompt word information entropy analysis—to conduct attack feature analysis, thereby determining whether the user is engaging in probing attacks against the large model. In practical applications, analysis dimensions can be selectively deployed according to the needs of the specific application scenario to accurately identify various probing attacks such as jailbreak attacks and guessing attacks, ensuring the operational security and stability of the large model.
[0012] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.
[0013] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0014] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0015] Figure 1 is a flowchart of a large-model attack detection method provided in an embodiment of the present invention;
[0016] Figure 2 is a flowchart of another large-model attack detection method provided in an embodiment of the present invention;
[0017] Figure 3 is a schematic diagram of a large-scale attack detection device provided in an embodiment of the present invention;
[0018] Figure 4 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] With the widespread application of large-scale models across various fields, particularly their deep integration into natural language processing, image generation, and recommender systems, the new challenges facing cybersecurity are becoming increasingly severe. While large-scale models possess enormous technological potential, their complexity and high computational power make them easy targets for attackers. Currently, one of the greatest security threats facing large-scale models is "jailbreak attacks," where attackers use clever input strategies to induce the model to produce malicious outputs that deviate from its original design goals. Such attacks not only severely jeopardize the model's credibility and stability but can also lead to model loss of control, posing significant data security risks to users and businesses.
[0021] With the continuous advancement of technology, hackers have become capable of manipulating large models through various methods for malicious inferences and improper use. For example, attackers may exploit model vulnerabilities to guide the model to output dangerous or inaccurate results through subtle changes in input, or even force the model to perform tasks it shouldn't, such as generating malicious content or leaking sensitive information. Therefore, how to effectively detect and prevent guessing attacks on large models has become a crucial issue that urgently needs to be addressed in the field of cybersecurity.
[0022] Among related technologies, attack protection measures for large models are relatively simple, focusing mainly on the protection of network boundaries, operating systems and application layers. However, during the use of large models, the security of the model itself is often overlooked, and it is impossible to effectively identify the risks such as jailbreak attacks and guessing attacks faced by large models.
[0023] Based on this, embodiments of the present invention provide a method, device, electronic device, and storage medium for detecting large model attacks. This technology can be applied to monitoring scenarios involving large model attacks such as jailbreak attacks and guessing attacks.
[0024] To facilitate understanding of this embodiment, a detailed description of a large-scale attack detection method disclosed in this invention will be provided first, as shown in Figure 1. This method includes the following steps:
[0025] Step S102: Obtain the prompt word sequence of the user input model; wherein, the prompt word sequence includes several prompt word instructions arranged in the order of input time.
[0026] The aforementioned large model can be a natural language processing model or a large language model, such as the GPT series models or the LLaMA series models. The aforementioned prompt word sequence refers to an ordered set of prompt word instructions submitted by the user to the large model, arranged in chronological order of input. Each prompt word instruction in this sequence corresponds to an independent user request. The aforementioned prompt word instructions refer to the text instructions actively input by the user to trigger the large model to execute a task. Their content features and input time are the core objects for subsequent attack feature detection, such as similarity analysis, difference analysis, and information entropy analysis.
[0027] In practice, user-submitted prompt commands can be collected at preset time intervals, and a complete prompt command sequence can be generated by combining the user's historical prompt command commands. The data dimensions of this prompt command sequence can include the prompt command, the command input time, and the feedback content of the large model in response to the prompt command. The prompt command commands in the prompt command sequence are ordered in chronological order of input time.
[0028] A user's prompt sequence includes the following:
[0029] ['Hello', '2025 / 04 / 30 10:00', 'Hello, how can I help you?']
[0030] ['Hi', '2025 / 04 / 30 10:01', 'Hello, how can I help you?']
[0031] ['hello', '2025 / 04 / 30 10:03', 'Hello, do you need help?']
[0032] Step S104: Perform attack feature analysis on the prompt word sequence; wherein, the attack feature analysis includes at least one of the following: similarity analysis of prompt word instructions, user request frequency analysis, difference analysis of prompt word instructions, repetition analysis of prompt word instructions, and information entropy analysis of prompt word instructions.
[0033] The similarity analysis of the aforementioned prompt words is used to identify whether multiple prompt words input by the user have highly similar characteristics, in order to capture malicious input behaviors such as batch probing and homogenization attacks. User request frequency analysis is used to monitor the frequency of user-input prompt words and the time interval characteristics of adjacent requests, promptly detecting attack behaviors that exhaust model resources, such as frequency bombardment and high-frequency requests within a short period. The difference analysis of prompt words is used to compare the vocabulary and character differences between other prompt words and a reference prompt word, identifying covert attacks implemented through abnormal character insertion, meaningless word additions or subtractions, or malicious word additions or subtractions. The reference prompt word is usually selected from prompt word sequences with typical characteristics, such as first-time input commands or suspected attack commands. The repetition analysis of prompt words is used to detect the degree of repetition and redundancy within a single prompt word, determining whether there are attack behaviors that interfere with the model's normal response through meaningless repetition. The information entropy analysis of prompt words measures the semantic richness and fluctuation characteristics of user prompt words, identifying abnormal behavior with excessively small fluctuations in prompt word entropy values, and locating regular malicious inputs such as guessing attacks and jailbreaking attacks.
[0034] It should be noted that attack signature analysis can be any single analysis dimension or a combination of multiple analysis dimensions mentioned above. The specific combination of analysis dimensions can be flexibly selected according to the detection requirements of the actual application scenario. Each analysis dimension executes its process independently and outputs analysis results independently, without interference between dimensions.
[0035] Step S106: Based on the analysis results corresponding to at least one attack feature, determine whether the user has engaged in probing attack behavior against the large model.
[0036] The above-mentioned attack feature analysis corresponds to one attack feature analysis dimension.
[0037] In other words, it can be determined directly based on the results of a single attack feature analysis dimension, or it can be determined comprehensively by combining the results of multiple attack feature analysis dimensions to determine whether a user has engaged in probing attack behavior against a large model.
[0038] In one implementation, the analysis results of attack feature analysis for each dimension are divided into two categories: abnormal results and non-abnormal results. Abnormal results indicate that user behavior matching the preset attack features of the corresponding dimension was detected in the analysis of the corresponding dimension; non-abnormal results indicate that user behavior matching the preset attack features of the corresponding dimension was not detected in the analysis of the corresponding dimension.
[0039] In one implementation, if only one dimension of attack feature analysis, such as user request frequency analysis, is deployed, when the output of the attack feature analysis of that dimension is abnormal, it can be determined that the user has engaged in probing attack behavior.
[0040] In one implementation, if multiple analysis dimensions are combined and deployed, that is, attack feature analysis is performed on the prompt word sequence in multiple dimensions such as prompt word command similarity analysis, user request frequency analysis, and prompt word command difference analysis, then the number of abnormal results corresponding to each dimension can be counted. When the number reaches the abnormal number threshold, it is determined that the user has engaged in probing attack behavior; otherwise, if the analysis results of each dimension are normal, or the number of abnormal results does not reach the threshold, it is determined that the user has not engaged in probing attack behavior against the large model.
[0041] The aforementioned large-scale model attack detection method includes: acquiring a sequence of prompt words input by the user into the large-scale model; wherein the prompt word sequence includes several prompt word instructions arranged in chronological order of input; performing attack feature analysis on the prompt word sequence; wherein the attack feature analysis includes at least one of the following: similarity analysis of prompt word instructions, user request frequency analysis, difference analysis of prompt word instructions, repetition analysis of prompt word instructions, and information entropy analysis of prompt word instructions; and determining whether the user has engaged in probing attack behavior against the large-scale model based on the analysis results corresponding to at least one attack feature analysis.
[0042] This method acquires a sequence of prompt words from a large user-input model, consisting of several prompt words arranged chronologically. Targeting common attacker tactics such as batch probing, frequency bombing, semantic evasion, and redundant input, it flexibly utilizes at least one dimension—prompt word similarity analysis, user request frequency analysis, prompt word difference analysis, prompt word repetition analysis, and prompt word information entropy analysis—to conduct attack feature analysis, thereby determining whether the user is engaging in probing attacks against the large model. In practical applications, analysis dimensions can be selectively deployed according to the needs of the specific application scenario to accurately identify various probing attacks such as jailbreak attacks and guessing attacks, ensuring the operational security and stability of the large model.
[0043] The following examples provide a specific implementation method for attack feature analysis of prompt word sequences.
[0044] In one approach, if the attack feature analysis of the prompt word sequence includes a similarity analysis of the prompt word instructions, then the analysis process corresponding to the similarity analysis of the prompt word instructions is as follows:
[0045] Each prompt word instruction in the prompt word sequence is segmented to obtain a vocabulary set corresponding to each prompt word instruction; the intersection words in vocabulary sets A and B corresponding to any two prompt word instructions whose semantic features satisfy a preset correlation degree are determined; the first ratio of the number of intersection words to the number of words in the interaction vocabulary set is calculated respectively; wherein, the interaction vocabulary set includes set A and set B; prompt word instructions corresponding to target vocabulary sets whose first ratio exceeds a first ratio threshold are determined as abnormal data; if the number of abnormal data in the prompt word sequence exceeds a preset number threshold, the analysis result corresponding to the similarity analysis of the prompt word instructions is determined as an abnormal result.
[0046] The vocabulary set mentioned above contains words or phrases; the word segmentation process can adopt a word segmentation algorithm adapted to natural language semantic understanding, and can adaptively adjust the word segmentation granularity according to the language type of the prompt word instruction to ensure that the segmented vocabulary set can accurately reflect the core semantics of the prompt word instruction.
[0047] The aforementioned semantic features that satisfy the preset correlation degree of the intersection words refer to words that come from two interactive word sets and have the same or similar semantics. The determination method is as follows: extract the feature vectors of each word in the two word sets respectively, calculate the feature vector distance between the two words belonging to different sets, and if the feature vector distance is less than or equal to the preset distance threshold, then the two words are determined as the intersection words that satisfy the preset correlation degree.
[0048] In real time, firstly, each prompt word instruction in the prompt word sequence is processed by word segmentation to obtain a vocabulary set corresponding to each prompt word instruction. The vocabulary set contains words or phrases, and the word segmentation process can adaptively adjust the splitting granularity according to the prompt word instruction to ensure that the vocabulary set can accurately reflect the core semantics of the prompt word instruction.
[0049] Furthermore, any two prompt words are selected from the multiple prompt word instructions, and their corresponding vocabulary sets A and B are obtained as interaction sets. The intersection words with similar or identical meanings in these two vocabulary sets are then determined.
[0050] Then, the number of intersecting words is counted, and the ratio of the number of intersecting words to the total number of words in word set A and the ratio of the number of intersecting words to the total number of words in word set B are calculated respectively. If any ratio exceeds the preset first ratio threshold, the prompt word instruction to which the word set corresponding to the ratio belongs is judged as abnormal data and marked as abnormal.
[0051] If the total number of prompt words marked as anomalous exceeds a preset threshold, it may indicate that the user is probing the model's vulnerabilities by making multiple minor modifications to the input, thus determining that the analysis result corresponding to the prompt word similarity analysis is anomaly; conversely, if the total number exceeds the threshold, the corresponding analysis result is determined to be normal. Here, the preset threshold can be dynamically configured based on the total number of prompt words acquired, for example, set to 50% of the total number of prompt words.
[0052] In one approach, if the attack feature analysis of the prompt word sequence includes user request frequency analysis, then the analysis process corresponding to the user request frequency analysis is as follows:
[0053] Based on the input time of each prompt word command, determine the number of times the user requests the large model within a specified time window and the interval between each two adjacent requests; if the number of requests exceeds the preset threshold and the maximum value of the interval is lower than the specified duration, determine that the analysis result corresponding to the user request frequency analysis is an abnormal result.
[0054] The specified time window is a time interval shorter than a preset duration threshold, typically set to a shorter duration to accurately capture high-frequency access behaviors within a short period. The number of requests refers to the total number of times a user submits prompt word commands to the large model within the specified time window. The interval between each two adjacent requests refers to the difference in input time between two adjacent prompt word commands.
[0055] When performing user request frequency analysis, based on the input time of the prompt word command, the total number of times the user submits the prompt word command to the large model within a specified time window is counted, and the difference between the input times of two adjacent prompt word commands is calculated.
[0056] If the number of requests counted exceeds a preset threshold, and the maximum interval between two consecutive requests is less than a specified duration—meaning the user exhibits a high request frequency with excessively short intervals within a short period—then the user is deemed to be engaging in malicious behavior such as guessing attacks, and the analysis result corresponding to the user request frequency analysis is determined to be an anomaly. Conversely, if the number of requests exceeds a preset threshold, the corresponding analysis result is determined to be normal.
[0057] For example, if a specified time window is set to 1 minute, a preset number of times threshold is 10, and a specified duration is 2 seconds, if a user sends 10 or more prompt word commands to the large model within 1 minute, and the maximum difference between the input times of two adjacent prompt word commands is less than 2 seconds, then malicious behavior may exist, and the analysis result corresponding to the user request frequency analysis is determined to be an abnormal result.
[0058] When a user requests a frequency analysis result that is abnormal, the system will automatically trigger a security warning and take further protective measures according to preset policies, such as limiting the user's subsequent access frequency, requiring authentication, or temporarily blocking access, in order to avoid high-frequency attacks causing the exhaustion of large model resources or response delays, and to ensure the stability of the large model service.
[0059] In one approach, if the attack feature analysis of the prompt word sequence includes: a difference analysis of the prompt word instructions, then the analysis process corresponding to the difference analysis of the prompt word instructions is as follows:
[0060] Identify the reference prompt word instruction in the prompt word sequence, and calculate the lexical differences between the non-reference prompt word instruction and the reference prompt word instruction; wherein, the data differences include: semantic lexical differences and non-semantic lexical differences; determine the number of lexical units of a specified format in the non-semantic lexical differences; if the number of lexical units of a specified format exceeds the difference number threshold, and there are additions or subtractions of words of a preset type in the non-reference prompt word instruction, then the analysis result corresponding to the difference analysis of the prompt word instruction is determined to be an abnormal result.
[0061] The aforementioned reference prompt words refer to the prompt word instructions selected from the prompt word sequence as the benchmark for difference comparison. These are typically selected from instructions with typical characteristics in the sequence, such as initial input instructions or suspected attack instructions. The aforementioned non-reference prompt word instructions refer to all prompt word instructions in the prompt word sequence other than the reference prompt word instructions. The aforementioned lexical differences refer to the differences between two prompt word instructions at the level of their smallest constituent units. The aforementioned semantic lexical differences refer to differences in lexical units that carry semantic meaning, such as the substitution, addition, or deletion of verbs and nouns. The aforementioned non-semantic lexical differences refer to differences in lexical units that do not carry semantic meaning, such as changes to spaces, punctuation marks, and formatting tags. The aforementioned number of specified format lexical units refers to special symbols, spaces, or abnormal characters (such as #¥%&). The number of lexical units (etc.). The above-mentioned preset types of vocabulary refer to words that have no actual semantic meaning, only serve as grammatical connectors or auxiliary functions of tone, or malicious words, etc.
[0062] In practice, suspected attack commands can be used as reference prompt words. Specifically, commands that match the preset attack characteristics can be selected from the prompt word sequence as reference prompt word commands, such as prompt word commands that contain sensitive evasion words, abnormal format combinations, or malicious guiding statements.
[0063] Then, the lexical differences between each non-reference prompt word instruction and the reference prompt word instruction in the prompt word sequence are calculated sequentially. The number of special symbols, spaces, or abnormal characters in the non-semantic lexical differences corresponding to each non-reference prompt word instruction is determined. If this number exceeds a preset difference threshold, and there are additions or subtractions of the aforementioned preset types of words in the non-reference prompt word instruction, then the analysis result corresponding to the difference analysis of the prompt word instruction is determined to be an abnormal result. Otherwise, the corresponding analysis result is determined to be an abnormal result.
[0064] In one approach, if the attack feature analysis of the prompt sequence includes a repeatability analysis of the prompt instructions, then the analysis process corresponding to the repeatability analysis of the prompt instructions is as follows:
[0065] Each prompt word instruction is segmented to obtain a vocabulary set corresponding to each prompt word instruction; the vocabulary sets are deduplicated to obtain a vocabulary processing set corresponding to each vocabulary set; the first vocabulary count of each vocabulary set and the second vocabulary count of the corresponding vocabulary processing set are determined; for each prompt word instruction corresponding to the first vocabulary set, a second ratio of the second vocabulary count to the first vocabulary count is calculated; the mean of the second ratio is determined, and if the mean is lower than a preset mean threshold, the analysis result corresponding to the repetition analysis of the prompt word instruction is determined to be an abnormal result.
[0066] The vocabulary set mentioned above contains words or phrases; the word segmentation process can adopt a word segmentation algorithm adapted to natural language semantic understanding, and can adaptively adjust the word segmentation granularity according to the language type of the prompt word instruction to ensure that the segmented vocabulary set can accurately reflect the core semantics of the prompt word instruction.
[0067] Specifically, each prompt word instruction is first processed by word segmentation to obtain a vocabulary set corresponding to each prompt word instruction. The vocabulary set contains words or phrases.
[0068] Then, the vocabulary set is deduplicated to remove duplicate words, so that each word is retained only once, resulting in the vocabulary processing set corresponding to each vocabulary set.
[0069] Then, count the number of first words in each word set and the number of second words in the word set after deduplication. For each prompt word instruction, calculate the ratio of the number of second words to the number of first words in the word set, and record it as the second ratio. This yields the second ratio for each prompt word instruction in the prompt word sequence.
[0070] Then, the mean of all second ratios is calculated. If the mean is lower than a preset mean threshold, such as 80%, it indicates that the prompt word command has a high degree of word repetition. This feature is usually a typical manifestation of malicious behaviors such as guessing attacks, and thus the analysis result corresponding to the repetition analysis of the prompt word command is determined to be an abnormal result. Conversely, if the mean is higher than the threshold, the corresponding analysis result is determined to be an abnormal result.
[0071] In one approach, if the attack feature analysis of the prompt word sequence includes information entropy analysis of the prompt word instruction, then the analysis process corresponding to the information entropy analysis of the prompt word instruction is as follows:
[0072] Each prompt word instruction is segmented to obtain the vocabulary set corresponding to each prompt word instruction; the information entropy of each vocabulary set is calculated; multiple target prompt word instructions are identified, and the information entropy values of the target vocabulary sets corresponding to the multiple target prompt word instructions are calculated respectively; the dispersion of the information entropy of the multiple target vocabulary sets is determined, wherein the multiple target prompt word instructions are arranged consecutively in the prompt word sequence; if the dispersion is lower than the preset dispersion threshold, the analysis result corresponding to the information entropy analysis of the prompt word instruction is determined to be an abnormal result.
[0073] The consecutive arrangement of multiple target prompt words in the prompt word sequence means that these multiple target prompt words refer to prompt word commands entered consecutively by the user. The information entropy mentioned above is a quantitative indicator describing the semantic complexity and richness of the vocabulary set corresponding to the prompt word command. A higher information entropy value indicates a richer semantic meaning and stronger randomness of the content, typically corresponding to the natural interaction behavior of normal users. A lower information entropy value indicates a more concentrated vocabulary distribution, less semantic variation, and stronger regularity of the content, potentially indicating malicious input characteristics. The preset duration range can be set as needed. The dispersion of the information entropy value can be calculated by taking the difference between the maximum and minimum values among multiple information entropy values, or by calculating the standard deviation of multiple information entropy values, using the standard deviation as the dispersion.
[0074] In one practical approach, the specific implementation steps are as follows:
[0075] First, each prompt word instruction is processed by word segmentation to obtain a vocabulary set corresponding to each prompt word instruction. The vocabulary set contains words or phrases.
[0076] Then, the following formula for information entropy is provided:
[0077]
[0078] in, The number of words in the vocabulary set; For the first word in this vocabulary set The frequency of each word.
[0079] For each vocabulary set, the information entropy of that vocabulary set is calculated using the information entropy formula described above.
[0080] Then, determine the multiple prompt words input by the user in succession, calculate the information entropy of the vocabulary set corresponding to the multiple prompt words, obtain the standard deviation of the multiple information entropies, and use the standard deviation as the dispersion of the information entropy.
[0081] If the dispersion is lower than the preset dispersion threshold, it indicates that the semantic complexity of the prompt words input by the user tends to be consistent, and there is a possibility of multiple similar guessing attacks, jailbreaking attacks, or other malicious behaviors. In this case, the information entropy analysis result corresponding to the prompt word command is determined to be an abnormal result. Conversely, if the dispersion is higher than the preset threshold, the corresponding analysis result is determined to be an abnormal result.
[0082] When the information entropy analysis result corresponding to the prompt word command is abnormal, the system will automatically trigger a security warning and take further protective measures according to the preset strategy to avoid high-frequency attacks causing large model resources to be exhausted or response delays, and to ensure the stability of large model services.
[0083] The above data analysis, by acquiring the prompt word sequence of user access to the large model within a preset time period, can conduct multi-dimensional collaborative data analysis based on the input time and prompt word instructions. Targeting typical attack methods commonly used by attackers, such as batch probing, frequency bombing, semantic evasion, and repetitive input, it can accurately match corresponding analysis dimensions such as similarity, request frequency, difference, repetition, and information entropy. It can comprehensively monitor user behavior from multiple levels such as semantic features, time features, and data structure features, and achieve accurate identification and prevention of various malicious attacks.
[0084] Please refer to Figure 2. This solution also provides an exemplary embodiment to illustrate the large-model attack detection method in more detail. As shown in Figure 2, this method includes the following steps:
[0085] Step S202: Obtain the prompt word sequence of the user input model; wherein, the prompt word sequence includes several prompt word instructions arranged in the order of input time.
[0086] This step is the same as step S102, and will not be described again.
[0087] Step S204: Perform attack feature analysis on the prompt word sequence; wherein, the attack feature analysis includes at least one of the following: similarity analysis of prompt word instructions, user request frequency analysis, difference analysis of prompt word instructions, repetition analysis of prompt word instructions, and information entropy analysis of prompt word instructions.
[0088] This step is the same as step S104, and will not be described again.
[0089] Step S206: Count the number of abnormal results in the analysis results corresponding to various attack characteristics; if the number of abnormal results exceeds the preset abnormal number threshold, it is determined that the user has engaged in probing attack behavior against the large model; if the number of abnormal results does not exceed the abnormal number threshold, it is determined that the user has not engaged in probing attack behavior against the large model.
[0090] In this step, the prompt word sequence is subjected to attack feature analysis in multiple dimensions. The analysis results for each dimension of attack feature analysis include abnormal results or no abnormal results.
[0091] Here, by analyzing attack characteristics from multiple dimensions, determining the number of abnormal results, and comparing them with a preset threshold for the number of abnormal results, it is possible to determine whether a user has engaged in probing attack behavior. Specifically, if the number of abnormal results reaches or exceeds the preset threshold, it is determined that the user has engaged in probing attack behavior against a large model; if the number of abnormal results does not reach the threshold, it is determined that the user has not engaged in probing attack behavior against a large model.
[0092] For example, if a user's prompt word sequence is subjected to attack feature analysis in five dimensions, including similarity analysis of prompt word commands, frequency analysis of user requests, difference analysis of prompt word commands, repetition analysis of prompt word commands, and information entropy analysis of prompt word commands, the preset threshold for the number of anomalies is 2.
[0093] If two or more of the analysis results corresponding to the five dimensions are judged as abnormal, such as the analysis results corresponding to the similarity analysis of prompt words and the repetition analysis of prompt words are both abnormal, then it is determined that the user has engaged in probing attack behavior against the large model.
[0094] If only one or zero items are identified as abnormal results, it is determined that the user has not engaged in any probing attacks against the large model.
[0095] This approach acquires the sequence of prompt words from user access to the large model within a preset time period. Multi-dimensional data analysis is then performed based on various information, including input time, prompt instructions, and model output. For common attacker tactics such as batch probing, frequency bombing, semantic evasion, and redundant input, user behavior is monitored from multiple analytical dimensions, including similarity, request frequency, difference, repetition, and information entropy. Multiple detection results are used to determine whether a user is engaging in malicious behavior. This multi-dimensional fusion detection mode avoids misjudgments caused by single-dimensional anomalies and ensures effective identification of attacks such as jailbreaks and guessing attacks, thus safeguarding the security of the large model.
[0096] The large model attack detection device of the present invention is described below. Please refer to Figure 3. One embodiment of the large model attack detection device of the present invention includes:
[0097] The first acquisition module 302 is used to acquire the prompt word sequence of the user input large model; wherein, the prompt word sequence includes several prompt word instructions arranged in the order of input time;
[0098] The first analysis module 304 is used to perform attack feature analysis on the prompt word sequence; wherein the attack feature analysis includes at least one of the following: similarity analysis of prompt word instructions, user request frequency analysis, difference analysis of prompt word instructions, repetition analysis of prompt word instructions, and information entropy analysis of prompt word instructions;
[0099] The first determining module 306 is used to determine whether the user has engaged in any probing attack behavior against the large model based on the analysis results corresponding to at least one of the attack features.
[0100] This method acquires a sequence of prompt words from a large user-input model, consisting of several prompt words arranged chronologically. Targeting common attacker tactics such as batch probing, frequency bombing, semantic evasion, and redundant input, it flexibly utilizes at least one dimension—prompt word similarity analysis, user request frequency analysis, prompt word difference analysis, prompt word repetition analysis, and prompt word information entropy analysis—to conduct attack feature analysis, thereby determining whether the user is engaging in probing attacks against the large model. In practical applications, analysis dimensions can be selectively deployed according to the needs of the specific application scenario to accurately identify various probing attacks such as jailbreak attacks and guessing attacks, ensuring the operational security and stability of the large model.
[0101] The above data analysis includes: similarity analysis of prompt word instructions; the first analysis module is used to perform word segmentation on each prompt word instruction in the prompt word sequence to obtain a vocabulary set corresponding to each prompt word instruction; determine the intersection of vocabulary sets A and B corresponding to any two prompt word instructions whose semantic features satisfy a preset correlation degree; calculate a first ratio between the number of intersection words and the number of words in the interactive vocabulary set; wherein, the interactive vocabulary set includes set A and set B; determine the prompt word instructions corresponding to the target vocabulary set whose first ratio exceeds a first ratio threshold as abnormal data; if the number of abnormal data in the prompt word sequence exceeds a preset number threshold, determine the analysis result corresponding to the similarity analysis of the prompt word instructions as an abnormal result.
[0102] The above data analysis includes: user request frequency analysis; the first analysis module is used to determine the number of times the user requests the large model within a specified time window and the interval between each adjacent request based on the input time of each of the prompt words; if the number of requests exceeds a preset threshold and the maximum value of the interval is lower than a specified duration, the analysis result corresponding to the user request frequency analysis is determined to be an abnormal result.
[0103] The above data analysis includes: difference analysis of the prompt word instructions; the first analysis module is used to determine the reference prompt word instructions in the prompt word sequence and calculate the lexical differences between the non-reference prompt word instructions and the reference prompt word instructions; wherein, the lexical differences include: semantic lexical differences and non-semantic lexical differences; determining the number of lexical units of a specified format in the non-semantic lexical differences; if the number of lexical units of the specified format exceeds the difference number threshold, and the non-reference prompt word instructions have additions or subtractions of words of a preset type, then the analysis result corresponding to the difference analysis of the prompt word instructions is determined to be an abnormal result.
[0104] The above data analysis includes: repetition analysis of the prompt word instructions; the first analysis module is used to perform word segmentation processing on each prompt word instruction to obtain a vocabulary set corresponding to each prompt word instruction; perform deduplication processing on the vocabulary set to obtain a vocabulary processing set corresponding to each vocabulary set; determine the first number of words in the vocabulary set and the second number of words in the vocabulary processing set corresponding to the vocabulary set; for each prompt word instruction corresponding to the first vocabulary set, calculate a second ratio of the second number of words to the first number of words; determine the mean of the second ratio, and if the mean is lower than a preset mean threshold, determine that the analysis result corresponding to the repetition analysis of the prompt word instructions is an abnormal result.
[0105] The above data analysis includes: information entropy analysis of the prompt word instructions; the first analysis module is used to perform word segmentation processing on each prompt word instruction to obtain the vocabulary set corresponding to each prompt word instruction, and calculate the information entropy of each vocabulary set; determine multiple target prompt word instructions, calculate the information entropy of the target vocabulary sets corresponding to the multiple target prompt word instructions respectively, and determine the dispersion of the information entropy of the multiple target vocabulary sets; wherein, the multiple target prompt word instructions are arranged consecutively in the prompt word sequence; if the dispersion is lower than a preset dispersion threshold, the analysis result corresponding to the information entropy analysis of the prompt word instruction is determined to be an abnormal result.
[0106] The aforementioned prompt word sequence is subjected to multiple attack feature analyses; the analysis results corresponding to each attack feature analysis include abnormal results or no abnormal results; the aforementioned first determining module is used to count the number of abnormal results among the analysis results corresponding to the multiple attack feature analyses; if the number of abnormal results exceeds a preset abnormal number threshold, it is determined that the user has engaged in probing attack behavior against the large model; if the number of abnormal results does not exceed the abnormal number threshold, it is determined that the user has not engaged in probing attack behavior against the large model.
[0107] This embodiment also provides an electronic device, including a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above-described large-scale attack detection method.
[0108] Referring to Figure 4, the electronic device includes a processor 100 and a memory 101. The memory 101 stores machine-executable instructions that can be executed by the processor 100. The processor 100 executes the machine-executable instructions to implement the large-scale attack detection method described above.
[0109] Furthermore, the electronic device shown in Figure 4 also includes a bus 102 and a communication interface 103, with the processor 100, the communication interface 103, and the memory 101 connected via the bus 102.
[0110] The memory 101 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 103 (which can be wired or wireless), such as the Internet, wide area network, local area network, or metropolitan area network. The bus 102 may be an ISA bus, PCI bus, or EISA bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only a single bidirectional arrow is used in Figure 4, but this does not indicate that there is only one bus or one type of bus.
[0111] Processor 100 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 100 or by instructions in software form. Processor 100 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a readily available storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 101, and the processor 100 reads the information from memory 101 and, in conjunction with its hardware, completes the steps of the method described in the foregoing embodiments.
[0112] This embodiment also provides a storage medium storing machine-executable instructions. When the machine-executable instructions are called and executed by the processor, the machine-executable instructions cause the processor to implement the above-mentioned large-scale attack detection method.
[0113] The computer program products of the large-scale attack detection method, apparatus, electronic device, and storage medium provided in the embodiments of the present invention include a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the preceding method embodiments. For specific implementation, please refer to the method embodiments, which will not be repeated here.
[0114] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system and apparatus described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0115] Furthermore, in the description of the embodiments of the present invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.
[0116] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0117] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0118] Finally, it should be noted that the above embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for detecting large-scale model attacks, characterized in that, The method includes: acquiring a prompt word sequence of a large model input by a user; wherein the prompt word sequence includes several prompt word instructions arranged in chronological order of input; performing attack feature analysis on the prompt word sequence; wherein the attack feature analysis includes similarity analysis of prompt word instructions, and at least one of the following: user request frequency analysis, difference analysis of prompt word instructions, repetition analysis of prompt word instructions, and information entropy analysis of prompt word instructions; wherein the user request frequency analysis is used to monitor the frequency of user input prompt word instructions and the time interval characteristics of adjacent requests; based on the analysis results corresponding to at least one of the attack feature analyses, determining whether the user has engaged in probing attack behavior against the large model; and performing similarity analysis of the prompt word instructions on the prompt word sequence. The method includes: performing word segmentation on each prompt word instruction in the prompt word sequence to obtain a vocabulary set corresponding to each prompt word instruction; wherein the vocabulary set contains words or phrases; determining the intersection of vocabulary sets A and B corresponding to any two prompt word instructions whose semantic features satisfy a preset correlation degree; calculating a first ratio between the number of intersection words and the number of words in the interactive vocabulary set; wherein the interactive vocabulary set includes vocabulary set A and vocabulary set B; identifying the prompt word instructions corresponding to the interactive vocabulary set whose first ratio exceeds a first ratio threshold as abnormal data; and determining the analysis result corresponding to the similarity analysis of the prompt word instructions as an abnormal result if the number of abnormal data in the prompt word sequence exceeds a preset number threshold.
2. The method according to claim 1, characterized in that, The attack feature analysis includes: user request frequency analysis; the attack feature analysis of the prompt word sequence includes: based on the input time of each prompt word instruction, determining the number of times the user requests the large model within a specified time window and the interval between each two adjacent requests; if the number of requests exceeds a preset threshold and the maximum value of the interval is lower than a specified duration, the analysis result corresponding to the user request frequency analysis is determined to be an abnormal result.
3. The method according to claim 1, characterized in that, The attack feature analysis includes: difference analysis of the prompt word instructions; the attack feature analysis of the prompt word sequence includes: determining the reference prompt word instructions in the prompt word sequence, and calculating the lexical differences between the non-reference prompt word instructions and the reference prompt word instructions; wherein, the lexical differences include: semantic lexical differences and non-semantic lexical differences; determining the number of lexical units of a specified format in the non-semantic lexical differences; if the number of lexical units of the specified format exceeds the difference number threshold, and the non-reference prompt word instructions have additions or subtractions of words of a preset type, then the analysis result corresponding to the difference analysis of the prompt word instructions is determined to be an abnormal result.
4. The method according to claim 1, characterized in that, The attack feature analysis includes: the repetition analysis of the prompt word instructions; the attack feature analysis of the prompt word sequence includes: performing word segmentation on each prompt word instruction to obtain a vocabulary set corresponding to each prompt word instruction; performing deduplication on the vocabulary set to obtain a vocabulary processing set corresponding to each vocabulary set; determining the first number of words in the vocabulary set and the second number of words in the vocabulary processing set corresponding to the vocabulary set; for each prompt word instruction corresponding to the first vocabulary set, calculating a second ratio of the second number of words to the first number of words; determining the mean of the second ratio, and if the mean is lower than a preset mean threshold, determining the analysis result corresponding to the repetition analysis of the prompt word instructions as an abnormal result.
5. The method according to claim 1, characterized in that, The attack feature analysis includes: information entropy analysis of the prompt word instructions; the attack feature analysis of the prompt word sequence includes: performing word segmentation on each prompt word instruction to obtain a vocabulary set corresponding to each prompt word instruction, and calculating the information entropy of each vocabulary set; identifying multiple target prompt word instructions, calculating the information entropy of the target vocabulary sets corresponding to the multiple target prompt word instructions respectively, and determining the dispersion of the information entropy of the multiple target vocabulary sets; wherein, the multiple target prompt word instructions are arranged consecutively in the prompt word sequence; if the dispersion is lower than a preset dispersion threshold, the analysis result corresponding to the information entropy analysis of the prompt word instructions is determined to be an abnormal result.
6. The method according to claim 1, characterized in that, The prompt word sequence is subjected to multiple attack feature analyses; the analysis results corresponding to each attack feature analysis include abnormal results or no abnormal results; determining whether the user has engaged in probing attack behavior against the large model based on the analysis results corresponding to at least one of the attack feature analyses includes: counting the number of abnormal results among the analysis results corresponding to the multiple attack feature analyses; if the number of abnormal results exceeds a preset abnormal number threshold, it is determined that the user has engaged in probing attack behavior against the large model; if the number of abnormal results does not exceed the abnormal number threshold, it is determined that the user has not engaged in probing attack behavior against the large model.
7. A large-scale attack detection device, characterized in that, The device includes: a first acquisition module, configured to acquire a sequence of prompt words input by a user into a large model; wherein the sequence of prompt words includes several prompt word instructions arranged in chronological order of input; a first analysis module, configured to perform attack feature analysis on the prompt word sequence; wherein the attack feature analysis includes similarity analysis of prompt word instructions, and at least one of the following: user request frequency analysis, difference analysis of prompt word instructions, repetition analysis of prompt word instructions, and information entropy analysis of prompt word instructions; wherein the user request frequency analysis is used to monitor the frequency of user input of prompt word instructions and the time interval characteristics of adjacent requests; a first determination module, configured to determine whether the user has engaged in probing attack behavior against the large model based on the analysis results corresponding to at least one of the attack feature analyses; An analysis module is used to perform word segmentation on each prompt word instruction in the prompt word sequence to obtain a vocabulary set corresponding to each prompt word instruction; wherein, the vocabulary set contains words or phrases; determine the intersection words in vocabulary sets A and B corresponding to any two prompt word instructions whose semantic features satisfy a preset correlation degree; calculate a first ratio between the number of intersection words and the number of words in the interaction vocabulary set; wherein, the interaction vocabulary set includes vocabulary set A and vocabulary set B; determine the prompt word instructions corresponding to the interaction vocabulary set whose first ratio exceeds a first ratio threshold as abnormal data; if the number of abnormal data in the prompt word sequence exceeds a preset number threshold, determine the analysis result corresponding to the similarity analysis of the prompt word instructions as an abnormal result.
8. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing machine-executable instructions that can be executed by the processor, the processor executing the machine-executable instructions to implement the large-scale attack detection method according to any one of claims 1-6.
9. A storage medium, characterized in that, The storage medium stores machine-executable instructions, which, when invoked and executed by a processor, cause the processor to implement the large-scale attack detection method according to any one of claims 1-6.
Citation Information
Patent Citations
Attack detection method and device, electronic equipment and nonvolatile storage medium
CN117201171A
Prompt word attack detection method and device for large language model
CN118445815A