Method, apparatus, device and storage medium for detecting stealing attack behavior
By estimating the topic and logical continuity likelihood estimation of the conversation data between users and the large language model and detecting theft attack behavior, the theft attack problem faced by the large language model is solved and the security of the model is enhanced.
Patent Information
- Application Number
- CN202411900875.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-12-23
AI Technical Summary
Large language models face stealing attacks, and malicious users infer the model structure and functions through repeated interactions, resulting in leakage of technical secrets and economic losses.
Theft attack behavior is detected by performing topic correlation likelihood estimation and logical continuity likelihood estimation on historical dialogue data between the user and the large language model.
Effectively identify and prevent stealing and attacks, enhance the security of large language models, and prevent technical secrets from being leaked and economic losses.
Smart Images

Figure CN119357952B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of natural language processing technology, and in particular to a method, device, equipment and storage medium for detecting stealing attack behaviors. Background Art
[0002] Large language models are deep learning models trained on huge datasets to understand human language. However, when malicious users interact with the model repeatedly and ask different questions, they can gradually infer the structure and function of the model, thereby obtaining the core intellectual property rights of the model. Such stealing attack behaviors not only threaten the technological secrets of enterprises, but may also lead to the leakage of intellectual property rights and even significant economic losses.
[0003] In related technologies, the defense mechanisms for large language models mainly focus on traditional access control and encryption technologies. However, malicious users can simulate the conversation behaviors of ordinary users, continuously ask questions to the model, and use the characteristics of model conversation generation to obtain information and infer the internal mechanism of the model. For example, they can disguise themselves through daily template conversations to evade the monitoring of traditional detection mechanisms, greatly affecting the security of large language models. Summary of the Invention
[0004] The purpose of this application is to provide a method, device, equipment and storage medium for detecting stealing attack behaviors, which can detect stealing attack behaviors according to the theme and logic of the conversation content, and enhance the security of large language models.
[0005] An embodiment of this application provides a method for detecting stealing attack behaviors, including:
[0006] Obtain historical conversation data between a user and a large language model; the historical conversation data includes multi-round conversation content data;
[0007] Perform a likelihood estimation of topic relevance on the conversation content data to obtain a topic likelihood estimation result representing the topic correlation strength between the conversation content data;
[0008] Perform a likelihood estimation of logical continuity on the conversation content data to obtain a logical likelihood estimation result representing the logical continuity strength between the conversation content data;
[0009] Perform a joint estimation on the topic likelihood estimation result and the logical likelihood estimation result to obtain a joint estimation result representing the topic correlation strength and logical continuity strength between the conversation content data;
[0010] Analyze the joint estimation result to detect stealing attack behaviors.
[0011] In some embodiments, the method for detecting stealing attack behaviors further includes:
[0012] Using a sliding window method, update the historical conversation data to add the latest conversation data to the historical conversation data and remove some of the earliest conversation data.
[0013] In some embodiments, the likelihood estimation of the topic relevance of the conversation content data to obtain a topic likelihood estimation result representing the topic relevance strength between the conversation content data includes:
[0014] Concatenate the historical conversation data and a preset first prompt word to obtain a first concatenated data;
[0015] Input the first concatenated data into the large language model to perform likelihood estimation of the topic relevance of the conversation content data, and obtain the topic likelihood estimation probability distribution of the large language model;
[0016] According to the topic likelihood estimation probability distribution, calculate the probability that the topic relevance strength between each conversation content data is less than a preset topic relevance strength threshold to obtain the topic likelihood estimation result.
[0017] In some embodiments, the likelihood estimation of the logical continuity of the conversation content data to obtain a logical likelihood estimation result representing the logical continuity strength between the conversation content data includes:
[0018] Concatenate the historical conversation data and a preset second prompt word to obtain a second concatenated data;
[0019] Input the second concatenated data into the large language model to perform likelihood estimation of the logical continuity of the conversation content data, and obtain the logical continuity estimation probability distribution of the large language model;
[0020] According to the logical continuity estimation probability distribution, the probability that the logical continuity strength between each conversation content data is less than a preset logical continuity strength threshold to obtain the logical likelihood estimation result.
[0021] In some embodiments, the joint estimation of the topic likelihood estimation result and the logical likelihood estimation result to obtain a joint estimation result representing the topic relevance strength and logical continuity strength between the historical conversation data includes:
[0022] According to the topic likelihood estimation result and the logical likelihood estimation result, calculate the probability that the topic relevance strength between the conversation content data is less than a preset topic relevance strength threshold and the logical continuity strength between each conversation content data is less than a preset logical continuity strength threshold to obtain the joint estimation result.
[0023] In some embodiments, analyzing the joint estimation result to detect stealing attack behavior includes:
[0024] According to the joint estimation result, determining whether the joint estimation probability exceeds a preset probability threshold; the joint estimation probability is the probability that the topic correlation strength between the conversation content data is less than a preset topic correlation strength threshold and the logical continuity strength between each conversation content data is less than a preset logical continuity strength threshold;
[0025] If it exceeds, generating a prompt message indicating that a stealing attack behavior is detected;
[0026] If it does not exceed, generating a prompt message indicating that no stealing attack behavior is detected.
[0027] In some embodiments, analyzing the joint estimation result to detect stealing attack behavior further includes:
[0028] Obtaining prompt messages generated at multiple different times within a preset time period;
[0029] When the number of prompt messages indicating that a stealing attack behavior is detected exceeds a preset number threshold, determining that there is an attack behavior.
[0030] An embodiment of the present application further provides a stealing attack behavior detection device, including:
[0031] A first module, configured to obtain historical conversation data between a user and a large language model; the historical conversation data includes multiple rounds of conversation content data;
[0032] A second module, configured to perform a likelihood estimation of topic relevance on the conversation content data to obtain a topic likelihood estimation result representing the topic correlation strength between the conversation content data;
[0033] A third module, configured to perform a likelihood estimation of logical continuity on the conversation content data to obtain a logical likelihood estimation result representing the logical continuity strength between the conversation content data;
[0034] A fourth module, configured to perform a joint estimation on the topic likelihood estimation result and the logical likelihood estimation result to obtain a joint estimation result representing the topic correlation strength and the logical continuity strength between the conversation content data;
[0035] A fifth module, configured to analyze the joint estimation result to detect stealing attack behavior.
[0036] An embodiment of the present application further provides an electronic device, where the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the above-mentioned stealing attack behavior detection method is implemented.
[0037] The embodiment of the present application also provides a computer-readable storage medium storing a computer program, characterized in that when the computer program is executed by a processor, the above-mentioned stealing attack behavior detection method is implemented.
[0038] Advantages of the present application: By performing likelihood estimation of topic relevance and likelihood estimation of logical continuity on the historical conversation data between the user and the large language model, a topic likelihood estimation result representing the topic correlation strength between the conversation content data and a logical likelihood estimation result representing the logical continuity strength between the conversation content data are obtained, and through joint estimation of the topic likelihood estimation result and the logical likelihood estimation result, a joint estimation result representing the topic correlation strength and logical continuity strength between the conversation content data is obtained, and then the stealing attack behavior is detected according to the joint estimation result. Since the topic likelihood estimation result and the logical likelihood estimation result between the content data of each round in the historical conversation data are used as the basis for detecting the stealing attack behavior, it is possible to accurately identify whether the user's behavior has malicious stealing attacks and enhance the security of the large language model. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a flowchart of the stealing attack behavior detection method provided by the embodiment of the present application.
[0040] Figure 2 It is a flowchart of the specific method of step S102 provided by the embodiment of the present application.
[0041] Figure 3 It is a flowchart of the specific method of step S103 provided by the embodiment of the present application.
[0042] Figure 4 It is a flowchart of the specific method of step S105 provided by the embodiment of the present application.
[0043] Figure 5 It is a schematic structural diagram of the stealing attack behavior detection device provided by the embodiment of the present application.
[0044] Figure 6 It is a schematic hardware structure diagram of the electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0046] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown can be executed in a different module division from that in the device or a different order from that in the flowchart. Terms such as "first" and "second" in the specification, claims and drawings are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0047] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0048] Figure 1 is a flowchart of the stealing attack behavior detection method provided by an embodiment of this application. Please refer to Figure 1 , in some embodiments, the method may include but is not limited to steps S101 to S105.
[0049] It should be noted that the execution subject of the stealing attack behavior detection method provided in the embodiments of this specification can be applied to end-side devices or cloud-side devices, and this embodiment does not make any limitation in this regard.
[0050] Step S101, obtain the historical conversation data between the user and the large language model.
[0051] The execution subject can obtain the historical conversation data between the user and the large language model.
[0052] The historical conversation data includes the multi-round conversation content data between the user and the large language model. Each round of conversation content data consists of the question data input by the user and the answer data replied by the large language model, and the historical conversation data is updated with the interaction between the user and the large language model.
[0053] In some embodiments, the stealing attack behavior detection method further includes the following steps: adopt a sliding window method to update the historical conversation data, so as to add the latest conversation data and eliminate some of the earliest conversation data in the historical conversation data.
[0054] By adopting a sliding window method to update the historical conversation data, the timeliness of the historical conversation data can be ensured. In specific implementation, when new conversation content data is generated, the earliest conversation turn will be removed, so that the number of conversation content data in the historical conversation data remains constant, thereby ensuring that the analyzed historical conversation data can accurately reflect the user's recent interaction behavior. More specifically, the number of conversation content data in the historical conversation data has an important impact on the accuracy and efficiency of detection. It can be set that the historical conversation data includes 20 rounds of conversation content data to achieve a balance between retaining sufficient context information and controlling the computational complexity. To optimize the storage and processing efficiency, the execution entity can only retain the currently obtained fixed historical conversation data, and the previously obtained historical conversation data will be automatically cleared to ensure the analysis of the latest user behavior.
[0055] Step S102, perform a likelihood estimation of the topic relevance of the conversation content data to obtain a topic likelihood estimation result representing the topic correlation strength between the conversation content data.
[0056] In specific implementation, it can be to first extract the topic words of each round of conversation content data in the historical conversation data, then determine the similarity between the topic words of each round of conversation content data in pairs, obtain a distribution of the topic correlation strength representing the relationship between each conversation content data according to the similarity between the topic words of each round of conversation content data in pairs, and then obtain the corresponding topic likelihood estimation result according to the topic correlation strength distribution.
[0057] Figure 2 It is a flowchart of the specific method of step S102 provided by the embodiments of the present application. Please refer to Figure 2 , in some embodiments, the method includes but is not limited to steps S201 to S203.
[0058] Step S201, splice the historical conversation data and a preset first prompt word to obtain a first spliced data.
[0059] In specific implementation, the historical conversation data and the preset first prompt word are spliced according to a preset format to form a corresponding data sequence, and the first spliced data is obtained.
[0060] Exemplarily, it can be spliced in the order that the historical conversation data is in the front and the first prompt word is in the back. The first prompt word can be "Please read the above historical conversation data and judge whether the topics of the conversation content data in each round are relevant. If the topics of most conversation turns are related to each other, please answer'related'; if the topics of the conversation turns are constantly changing and irrelevant, please answer 'irrelevant'."
[0061] Step S202: Input the first spliced data into the large language model to perform a likelihood estimation of topic relevance for the conversation content data, and obtain the topic likelihood estimation probability distribution of the large language model.
[0062] In specific implementation, the large language model obtains the first spliced data, extracts the topics of each round of conversation content data in the historical conversation data according to the indication of the first prompt word, and evaluates the similarity of the extracted topics of each conversation content data to obtain the topic similarity evaluation result. Then, according to the topic similarity evaluation result, determine the topic similarity distribution between each pair of conversation content data, that is, the topic correlation intensity distribution, and further determine the topic likelihood estimation probability distribution of the large language model according to the topic correlation intensity distribution.
[0063] It can be understood that the topic likelihood estimation probability distribution of the large language model includes the similarity probability between the topics of each conversation content data. The probability value of the similarity probability is positively correlated with the topic correlation intensity between the conversation content data. The large language model processes the topic likelihood estimation probability distribution and outputs the corresponding topic similarity evaluation result, which may be the topic similarity evaluation result related to the topics of each conversation content data or the topic similarity evaluation result unrelated to the topics of each conversation content data.
[0064] Step S203: According to the topic likelihood estimation probability distribution, calculate the probability that the topic correlation intensity between each pair of conversation content data is less than the preset topic correlation intensity threshold to obtain the topic likelihood estimation result.
[0065] In specific implementation, introduce a likelihood estimation function to estimate the topic likelihood estimation probability distribution of the large language model, so as to calculate the probability that the large language model determines that the topics of each conversation content data are relevant and the probability that the large language model determines that the topics of each conversation content data are irrelevant in the topic likelihood estimation probability distribution of the large language model. Among them, the probability that the large language model determines that the topics of each conversation content data are relevant refers to the probability that the topic correlation intensity between each pair of conversation content data is not less than the preset topic correlation intensity threshold, and the probability that the large language model determines that the topics of each conversation content data are irrelevant refers to the probability that the topic correlation intensity between each pair of conversation content data is less than the preset topic correlation intensity threshold. According to the probability that the topic correlation intensity between each pair of conversation content data is less than the preset topic correlation intensity threshold, obtain the topic likelihood estimation result.
[0066] Step S103: Perform a likelihood estimation of logical continuity for the conversation content data to obtain a logical likelihood estimation result representing the logical continuity intensity between the conversation content data.
[0067] In a specific implementation, the context semantics of each round of conversation content data in the historical conversation data may be first extracted, and then the correlation strength between the context semantic features of each round of conversation content data is determined. According to the correlation strength between the context semantic features of each round of conversation content data, the logical continuity strength distribution characterizing the conversation content data is obtained, and then the corresponding logical likelihood estimation result is obtained according to the logical continuity strength distribution.
[0068] Figure 3 is a flowchart of the specific method of step S103 provided in the embodiment of the present application. Figure 3 In some embodiments, the method includes but is not limited to steps S301 to S303.
[0069] Step S301, splicing the historical conversation data and the preset second prompt word to obtain second spliced data.
[0070] In a specific implementation, the historical conversation data and the preset second prompt word are spliced according to a preset format to form a corresponding data sequence to obtain the second spliced data.
[0071] Exemplarily, the historical conversation data may be spliced in the order of first and second prompt words in the second order, and the second prompt words may be "Please read the following historical conversation data to determine whether there is logical continuity between the content data of each round of conversation. That is, determine whether each round of conversation can carry forward the content of the previous round and start a discussion, rather than just repeating questions around the same topic. If there is a relationship of continuity between each round of conversation, please answer "there is continuity"; if there is no relationship of continuity between each round of conversation, please answer "there is no continuity.".
[0072] Step S302: input the second concatenated data into the large language model to perform logical continuity likelihood estimation on the conversation content data, and obtain the logical continuity estimation probability distribution of the large language model.
[0073] In a specific implementation, the large language model obtains the second spliced data, performs context semantic extraction on the content data of each round of conversation in the historical conversation data according to the instruction of the second prompt word, and performs relevance evaluation on the context semantic features of the extracted content data of each conversation to obtain a context relevance evaluation result, and then determines the logical continuity similarity distribution between each round of conversation content data according to the context relevance evaluation result, that is, the logical continuity correlation intensity distribution, and then determines the logical continuity likelihood estimation probability distribution of the large language model according to the logical continuity correlation intensity distribution.
[0074] It can be understood that the logical continuity likelihood estimation probability distribution of the large language model contains the correlation probabilities between the contextual semantic features of each conversation content data. The probability value of the correlation probability is positively correlated with the correlation strength between the conversation content data. The large language model processes the logical continuity likelihood estimation probability distribution and outputs the corresponding contextual correlation degree evaluation result, which may be the contextual correlation degree evaluation result indicating the logical continuity between each conversation content data, or the contextual correlation degree evaluation result indicating the lack of logical continuity between each conversation content data.
[0075] Step S303: According to the logical continuity estimation probability distribution, obtain the logical likelihood estimation result based on the probability that the logical continuity strength between each conversation content data is less than the preset logical continuity strength threshold.
[0076] In specific implementation, a likelihood estimation function is introduced to estimate the logical continuity estimation probability distribution of the large language model, so as to calculate the probability that the large language model determines that each conversation content data has logical continuity and the probability that the large language model determines that each conversation content data does not have logical continuity in the logical continuity estimation probability distribution of the large language model. Among them, the probability that the large language model determines that each conversation content data has logical continuity refers to the probability that the logical continuity strength between each conversation content data is not less than the preset logical continuity strength threshold, and the probability that the large language model determines that each conversation content data does not have logical continuity refers to the probability that the logical continuity strength between each conversation content data is less than the preset logical continuity strength threshold. Based on the probability that the logical continuity strength between each conversation content data is less than the preset logical continuity strength threshold, obtain the logical likelihood estimation result.
[0077] Step S104: Jointly estimate the topic likelihood estimation result and the logical likelihood estimation result to obtain a joint estimation result representing the topic correlation strength and logical continuity strength between the conversation content data.
[0078] In some embodiments, step S104 includes the following steps: According to the topic likelihood estimation result and the logical likelihood estimation result, calculate the probability that the topic correlation strength between each conversation content data is less than the preset topic correlation strength threshold and the logical continuity strength between each conversation content data is less than the preset logical continuity strength threshold, to obtain the joint estimation result.
[0079] In a specific embodiment, the product of the probability that the topic relevance strength between each conversation content data is less than a preset topic relevance strength threshold and the probability that the logical continuity strength between each conversation content data is less than a preset logical continuity strength threshold is calculated to obtain a joint estimation result. The joint estimation result includes a joint estimation probability, and the joint estimation probability is the probability that the topic relevance strength between the conversation content data is less than the preset topic relevance strength threshold and the logical continuity strength between the conversation content data is less than the preset logical continuity strength threshold.
[0080] It can be understood that in the embodiments of the present application, the probability that the topic relevance strength between each conversation content data is less than the preset topic relevance strength threshold and the logical continuity strength between each conversation content data is less than the preset logical continuity strength threshold is used as the basis for detecting stealing attack behavior. Specifically, if the topic relevance strength between the conversation content data of each round is not less than the preset topic relevance strength threshold and the logical continuity strength between the conversation content data of each round is not less than the preset logical continuity strength threshold, that is, the conversation between the user and the large language model is topic relevant and has logical continuity, it is considered that the formed historical conversation data conforms to the mode of normal human communication and belongs to normal user behavior rather than malicious stealing attack behavior. On the contrary, if the topic relevance strength between the conversation content data of each round is less than the preset topic relevance strength threshold and the logical continuity strength between the conversation content data of each round is less than the preset logical continuity strength threshold, that is, the conversation between the user and the large language model is topic irrelevant and has no logical continuity, it is considered that the formed historical conversation data does not conform to the mode of normal human communication, does not belong to normal user behavior, and belongs to malicious stealing attack behavior. The reason is that normal users usually have a clear purpose or topic when chatting with the large model, and the conversation content will have a certain logical coherence. Even in the case of casual chatting, the communication between normal users can often maintain a certain degree of topic relevance and logical continuity. However, the main purpose of malicious users is to obtain as much knowledge or data of the large language model as possible. To achieve this goal, they will try to use various different inputs to trigger the large language model to output as much information as possible, and these inputs may not have an actual topic or logical continuity.
[0081] Step S105, analyze the joint estimation result to detect stealing attack behavior.
[0082] Figure 4 is a flowchart of the specific method of step S105 provided by the embodiments of the present application. Please refer to Figure 4 , in some embodiments, the method includes but is not limited to steps S401 to S403.
[0083] Step S401: According to the joint estimation result, determine whether the joint estimation probability exceeds a preset probability threshold. If it exceeds, execute Step S402; if not, execute Step S403.
[0084] Among them, the joint estimation probability is the probability that the topic correlation strength between the dialogue content data is less than a preset topic correlation strength threshold and the logical continuity strength between the dialogue content data is less than a preset logical continuity strength threshold.
[0085] Preferably, the probability threshold is set to 0.7.
[0086] Step S402: Generate a prompt message indicating that a stealing attack behavior is detected;
[0087] Step S403: Generate a prompt message indicating that no stealing attack behavior is detected.
[0088] In some embodiments, Step S105 further includes the following steps: Obtain prompt messages generated at multiple different times within a preset duration; When the number of prompt messages indicating that a stealing attack behavior is detected exceeds a preset quantity threshold, determine that there is an attack behavior.
[0089] In specific implementation, it can be to determine whether the joint estimation probability exceeds the preset probability threshold at different times within the preset duration to generate multiple prompt messages. If the number of prompt messages indicating that a stealing attack behavior is detected exceeds the preset quantity threshold, it is determined that there is an attack behavior and an alarm is triggered.
[0090] Please refer to Figure 5 , this application embodiment also provides a stealing attack behavior detection device, which can implement the above stealing attack behavior detection method. The device includes:
[0091] The first module 501 is used to obtain historical dialogue data between the user and the large language model; the historical dialogue data includes multi-round dialogue content data;
[0092] The second module 502 is used to perform a likelihood estimation of topic relevance on the dialogue content data to obtain a topic likelihood estimation result representing the topic correlation strength between the dialogue content data;
[0093] The third module 503 is used to perform a likelihood estimation of logical continuity on the dialogue content data to obtain a logical likelihood estimation result representing the logical continuity strength between the dialogue content data;
[0094] The fourth module 504 is used to perform a joint estimation on the topic likelihood estimation result and the logical likelihood estimation result to obtain a joint estimation result representing the topic correlation strength and the logical continuity strength between the dialogue content data;
[0095] The fifth module 505 is used to analyze the joint estimation result to detect stealing attack behaviors.
[0096] The specific implementation manner of this stealing attack behavior detection device is basically the same as the specific embodiments of the above stealing attack behavior detection method, and will not be described in detail herein.
[0097] Figure 6 It is a block diagram of an electronic device shown according to an exemplary embodiment.
[0098] Next, refer to Figure 6 to describe the electronic device 600 according to this embodiment of the present disclosure. Figure 6 The electronic device 600 shown is only an example, and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0099] As Figure 6 shown, the electronic device 600 is presented in the form of a general-purpose computing device. The components of the electronic device 600 may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different system components (including the storage unit 620 and the processing unit 610), a display unit 640, etc.
[0100] Among them, the storage unit stores program codes, and the program codes can be executed by the processing unit 610, so that the processing unit 610 executes the steps according to various exemplary embodiments of the present disclosure described in the above stealing attack behavior detection method part of this specification.
[0101] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 6201 and / or a cache storage unit 6202, and may further include a read-only storage unit (ROM) 6203.
[0102] The storage unit 620 may further include a program / utilities 6204 having a set (at least one) of program modules 6205. Such program modules 6205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. The implementation of a network environment may be included in each or some combination of these examples.
[0103] The bus 630 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in a variety of bus structures.
[0104] The electronic device 600 can also communicate with one or more external devices 600' (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 600, and / or communicate with any device that enables the electronic device 600 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 650. Moreover, the electronic device 600 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 660. The network adapter 660 can communicate with other modules of the electronic device 600 through the bus 630. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0105] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned theft attack behavior detection method is implemented.
[0106] The theft attack behavior detection method, device, equipment and storage medium provided by the embodiment of the present application obtain a topic likelihood estimation result representing the topic correlation strength between the dialogue content data and a logical likelihood estimation result representing the logical continuity strength between the dialogue content data by performing topic relevance likelihood estimation and logical continuity likelihood estimation on the historical dialogue data between the user and the large language model, and obtain a joint estimation result representing the topic correlation strength and logical continuity strength between the dialogue content data by performing joint estimation on the topic likelihood estimation result and the logical likelihood estimation result, and then detect the theft attack behavior according to the joint estimation result. Since the topic likelihood estimation result and the logical likelihood estimation result between the content data of each round of dialogue in the historical dialogue data are used as the basis for detecting the theft attack behavior, it is possible to accurately identify whether the user's behavior has malicious theft attacks and enhance the security of the large language model.
[0107] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described here can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on the network, including several instructions to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the above-mentioned method according to the embodiments of the present disclosure.
[0108] The program product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0109] A computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable storage medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0110] Those skilled in the art can understand that the above-mentioned modules can be distributed in the device according to the description of the embodiments, or can be correspondingly changed and distributed in one or more devices that are uniquely different from this embodiment. The modules of the above embodiments can be combined into one module, or further split into multiple sub-modules.
[0111] The exemplary embodiments of the present disclosure have been specifically shown and described above. It should be understood that the present disclosure is not limited to the detailed structures, settings, or implementation methods described herein; on the contrary, the present disclosure is intended to cover various modifications and equivalent settings included within the spirit and scope of the appended claims.
Claims
1. A theft attack behavior detection method, characterized in that: include: Obtain historical conversation data between users and large language models; The historical conversation data includes multiple rounds of conversation content data; Performing topic relevance likelihood estimation on the conversation content data to obtain a topic likelihood estimation result representing the topic relevance strength between the conversation content data; Performing logical continuity likelihood estimation on the conversation content data to obtain a logical likelihood estimation result representing the strength of logical continuity between the conversation content data; Jointly estimating the topic likelihood estimation result and the logical likelihood estimation result to obtain a joint estimation result representing the topic correlation strength and the logical continuity strength between the conversation content data; Analyzing the joint estimation result to detect theft attack behavior; The performing of topic relevance likelihood estimation on the conversation content data to obtain a topic likelihood estimation result representing the topic relevance strength between the conversation content data includes: splicing the historical conversation data and a preset first prompt word to obtain first spliced data; Inputting the first concatenated data into the large language model to perform topic relevance likelihood estimation on the conversation content data to obtain a topic likelihood estimation probability distribution of the large language model; According to the topic likelihood estimation probability distribution, the probability that the topic correlation strength between each of the conversation content data is less than a preset topic correlation strength threshold is calculated to obtain the topic likelihood estimation result; The performing logical continuity likelihood estimation on the conversation content data to obtain a logical likelihood estimation result characterizing the logical continuity strength between the conversation content data includes: splicing the historical conversation data and a preset second prompt word to obtain second spliced data; Inputting the second concatenated data into the large language model to perform a logical continuity likelihood estimation on the conversation content data to obtain a logical continuity estimation probability distribution of the large language model; According to the logical continuity estimation probability distribution, the probability that the logical continuity strength between each of the conversation content data is less than a preset logical continuity strength threshold is used to obtain the logical likelihood estimation result.
2. The theft attack behavior detection method according to claim 1, characterized in that: The theft attack behavior detection method further includes: The historical conversation data is updated by adopting a sliding window method, so as to add the latest conversation data to the historical conversation data and remove part of the earliest conversation data.
3. The theft attack behavior detection method according to claim 1, characterized in that: The jointly estimating the topic likelihood estimation result and the logical likelihood estimation result to obtain a joint estimation result characterizing the topic correlation strength and the logical continuity strength between the historical conversation data includes: Based on the topic likelihood estimation result and the logical likelihood estimation result, the probability that the topic correlation strength between the conversation content data is less than a preset topic correlation strength threshold and the logical continuity strength between each conversation content data is less than a preset logical continuity strength threshold is calculated to obtain the joint estimation result.
4. The theft attack behavior detection method according to claim 1, characterized in that: The analyzing the joint estimation result to detect the theft attack behavior includes: According to the joint estimation result, judging whether the joint estimation probability exceeds a preset probability threshold; the joint estimation probability is the probability that the topic correlation strength between the conversation content data is less than a preset topic correlation strength threshold and the logic continuity strength between each of the conversation content data is less than a preset logic continuity strength threshold; If it exceeds, a prompt message is generated indicating that the theft attack behavior is detected; If it does not exceed the limit, a prompt message is generated indicating that no theft attack behavior is detected.
5. The theft attack behavior detection method according to claim 4, characterized in that: The analyzing the joint estimation result to detect the theft attack behavior also includes: Get multiple prompt messages generated at different times within a preset time period; When the number of the prompt information of the detected theft attack behavior exceeds a preset number threshold, it is determined that the attack behavior exists.
6. A theft attack behavior detection device, characterized in that: include: The first module is used to obtain historical conversation data between users and the large language model; The historical conversation data includes multiple rounds of conversation content data; The second module is used to perform topic relevance likelihood estimation on the conversation content data to obtain a topic likelihood estimation result representing the topic relevance strength between the conversation content data; The third module is used to perform a logical continuity likelihood estimation on the conversation content data to obtain a logical likelihood estimation result representing the strength of the logical continuity between the conversation content data; A fourth module is used to jointly estimate the topic likelihood estimation result and the logical likelihood estimation result to obtain a joint estimation result representing the topic correlation strength and the logical continuity strength between the conversation content data; A fifth module is used to analyze the joint estimation result to detect theft attack behavior; The performing of topic relevance likelihood estimation on the conversation content data to obtain a topic likelihood estimation result representing the topic relevance strength between the conversation content data includes: splicing the historical conversation data and a preset first prompt word to obtain first spliced data; Inputting the first concatenated data into the large language model to perform topic relevance likelihood estimation on the conversation content data to obtain a topic likelihood estimation probability distribution of the large language model; According to the topic likelihood estimation probability distribution, the probability that the topic correlation strength between each of the conversation content data is less than a preset topic correlation strength threshold is calculated to obtain the topic likelihood estimation result; The performing logical continuity likelihood estimation on the conversation content data to obtain a logical likelihood estimation result characterizing the logical continuity strength between the conversation content data includes: splicing the historical conversation data and a preset second prompt word to obtain second spliced data; Inputting the second concatenated data into the large language model to perform a logical continuity likelihood estimation on the conversation content data to obtain a logical continuity estimation probability distribution of the large language model; According to the logical continuity estimation probability distribution, the probability that the logical continuity strength between each of the conversation content data is less than a preset logical continuity strength threshold is used to obtain the logical likelihood estimation result.
7. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the theft attack behavior detection method according to any one of claims 1 to 5 when executing the computer program.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the theft attack behavior detection method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Method for estimating relationships between topics, and system
CN104239385A
Prompt word attack detection method and device for large language model
CN118445815A