A method, system, optimized terminal and medium for optimizing prompt detection resources
Through the multi-stage dynamic Bayesian game model optimization, the problems of prompt word security and resource consumption in large language models are solved, and security improvement and resource conservation are achieved.
Patent Information
- Application Number
- CN202510473804.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-16
AI Technical Summary
In the prior art, synonymous replacement or filtering of prompt words in large language models leads to problems such as low output security and high energy consumption.
The multi-stage dynamic Bayesian game model is used to identify the parameters of the prompt word, and optimize the consistency functions of benign users, vicious users and defenders to optimize the security of prompt word and reduce resource consumption, and use edge nodes and vector databases for detection.
It improves the output security of the large language model, reduces system resource consumption and service delays for benign users, and improves user experience.
Smart Images

Figure CN120017420B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network security technology, and in particular, to a method, system, optimized terminal and medium for optimizing prompt detection resources. Background Art
[0002] Large Language Models (LLMs) are widely used. Researchers construct prompts through prompt engineering to ensure that the outputs generated by LLMs can better meet user requirements and improve response quality. However, attackers also use prompt engineering to construct malicious prompts and launch prompt attacks, including injection attacks and jailbreak attacks, resulting in the output of LLMs being insecure content. At the same time, they can also allow unauthorized access to private data. For example, the tools integrated with LLMs may be vulnerable to prompt injection attacks, thus exposing sensitive information. To protect the security of LLM outputs, researchers have proposed some methods for filtering prompts, but these methods do not jointly consider resource consumption and system performance optimization.
[0003] In the existing technical solutions, one is to perform synonym replacement on the prompts. This method may inadvertently change the meaning of the original prompt, resulting in content modification and a decline in the response quality of the LLM. The second is to use the LLM to filter the prompts. However, the LLM may not be able to recognize malicious attacks, and using the LLM for detection consumes a large amount of resources. The third is to use a small model classifier for prompt detection and filtering. This method can achieve a relatively high accuracy rate. However, in reality, the ratio of malicious prompts is small, and filtering all prompts will waste a large amount of resources.
[0004] Filtering all prompts by the existing model will waste a large amount of resources.
[0005] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0006] The main purpose of this application is to provide a method, system, optimized terminal and medium for optimizing prompt detection resources, aiming to solve the problems in the existing technology that synonym replacement or filtering of prompts in large language models leads to low output security and high energy consumption of large language models.
[0007] In the first aspect of the embodiments of the present application, a method for optimizing prompt detection resources is provided. The method for optimizing prompt detection resources includes the following steps: receiving a prompt sent by a user and identifying the prompt parameters of the prompt; constructing a multi-stage dynamic Bayesian game model, and inputting the prompt parameters into the multi-stage dynamic Bayesian game model to obtain a detection probability; detecting the prompt according to the detection probability to obtain a safe prompt; and inputting the safe prompt into a large language model for processing to obtain a response message.
[0008] Optionally, in an embodiment of the present application, the constructing of the multi-stage dynamic Bayesian game model specifically includes: constructing a consistency function corresponding to each of a benign user, a malicious user, and a defender; and minimizing the consistency functions corresponding to the benign user, the malicious user, and the defender respectively to obtain a multi-stage dynamic Bayesian game model.
[0009] Optionally, in an embodiment of the present application, the consistency function of the benign user is expressed as:
[0010] ;
[0011] The consistency function of the malicious user is expressed as:
[0012] ;
[0013] The consistency function of the defender is expressed as: ;
[0014] Wherein, is the consistency function of the benign user, m represents an edge node, represents a set of edge nodes, represents the benign user x 's strategy, is the total delay of whether to detect and whether to send the prompt to the large language model for processing after the prompt of the benign user x is sent to the edge node m , represents the number of rounds of the game, is the total number of rounds of the game; is the consistency function of the malicious user, represents the strategy of the malicious user y , is the F1 score of the detection model, represents the prompt, represents the number of tokens of the prompt, represents the number of floating-point operations per token; is the consistency function for the defender, , represent two weight parameters greater than zero, represents the prompt word, represents the t sequence of prompt words received by the defender in the round, represents the defender's t strategy for detecting each prompt word in the round, represents the total delay for processing each prompt word, represents the defender's belief in each prompt word.
[0015] Optionally, in an embodiment of the present application, the multi-stage dynamic Bayesian game model is expressed as:
[0016] ;
[0017] ;
[0018] ;
[0019] ;
[0020] wherein, represents the strategy of benign users t and malicious users other than the defender x for detecting whether the prompt word y is detected in the round, represents the set of defenders, is the number of malicious prompt words in the t round.
[0021] Optionally, in an embodiment of the present application, after inputting the prompt parameter into the multi-stage dynamic Bayesian game model to obtain the detection probability, the method further includes: updating the beliefs of the benign user, the malicious user, and the defender regarding the security of the prompt word according to the detection probability to obtain updated beliefs; updating the multi-stage dynamic Bayesian game model according to the updated beliefs to obtain an updated multi-stage dynamic Bayesian game model, and inputting the next-round prompt parameter into the updated multi-stage dynamic Bayesian game model to obtain the next-round detection probability.
[0022] Optionally, in an embodiment of the present application, the detecting the prompt word according to the detection probability to obtain a security prompt word specifically includes: when the detection probability meets a preset requirement, detecting the prompt word to obtain a detection result; if the detection result is secure, determining the prompt word as a security prompt word.
[0023] Optionally, in an embodiment of the present application, after detecting the prompt word to obtain a detection result when the detection probability meets a preset requirement, it further includes: if the detection result is insecure, returning the prompt word for the next round of prompt word detection.
[0024] A second aspect of the embodiments of the present application further provides a prompt word detection resource optimization system, where the prompt word detection resource optimization system includes:
[0025] A prompt word receiving module, configured to receive a prompt word sent by a user and identify the prompt parameters of the prompt word;
[0026] A model construction and optimization module, configured to construct a multi-stage dynamic Bayesian game model and input the prompt parameters into the multi-stage dynamic Bayesian game model to obtain a detection probability;
[0027] A detection strategy determination module, configured to detect the prompt word according to the detection probability to obtain a security prompt word;
[0028] A large model output module, configured to input the security prompt word into a large language model for processing to obtain a response message.
[0029] A third aspect of the embodiments of the present application further provides a prompt word detection resource optimization terminal, where the prompt word detection resource optimization terminal includes: a memory, a processor, and a prompt word detection resource optimization program stored on the memory and executable on the processor, and when the prompt word detection resource optimization program is executed by the processor, it implements the steps of the prompt word detection resource optimization method as described above.
[0030] A fourth aspect of the embodiments of the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores a prompt word detection resource optimization program, and when the prompt word detection resource optimization program is executed by a processor, it implements the steps of the prompt word detection resource optimization method as described above.
[0031] Beneficial effects: The present application provides a method, system, optimized terminal and medium for optimizing prompt detection resources. The multi-stage dynamic Bayesian game model of the present application makes a decision on whether to detect each incoming prompt, and according to the strategy of only detecting the prompts that need to be detected based on the model output, thereby reducing system resource consumption and reducing service latency for benign users. The present application optimizes the security of prompts and system performance by identifying users' prompts and obtaining a detection strategy based on the multi-stage dynamic Bayesian game model, so as to effectively identify and prevent malicious prompt detection, improve the security of LLM output, and be able to intelligently select the prompts that need to be detected to avoid unnecessary resource consumption, and further reduce the service latency of benign users to enhance the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0033] Figure 1 is a flowchart of a preferred embodiment of the method for optimizing prompt detection resources of the present application;
[0034] Figure 2 is a schematic diagram of the specific implementation steps of the multi-stage dynamic Bayesian game in a preferred embodiment of the method for optimizing prompt detection resources of the present application;
[0035] Figure 3 is a structural diagram of a preferred embodiment of the system for optimizing prompt detection resources of the present application;
[0036] Figure 4 is a structural diagram of a preferred embodiment of the terminal for optimizing prompt detection resources of the present application.
[0037] Description of the reference numerals:
[0038] 100, prompt receiving module; 200, model construction and optimization module; 300, detection strategy determination module; 400, large model output module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] To make the purpose, technical solution and effects of this application clearer and more definite, the following will describe the technical solutions in the embodiments of this application clearly and completely in conjunction with the accompanying drawings in the embodiments of this application. The described embodiments are only possible technical implementations of this application, not all possible implementations. Based on the embodiments in this application, those skilled in the art can fully combine the embodiments of this application to obtain other embodiments without creative labor, and these embodiments are also within the protection scope of this application.
[0040] First, introduce the terms involved in the embodiments of this application:
[0041] LLM: Large Language Model, a large language model, abbreviated as large model;
[0042] VDB: Vector Database, a vector database.
[0043] Secondly, introduce the system architecture of the embodiments of this application.
[0044] Deploy a vector database (VDB) and a prompt anomaly detector (detect malicious information) on each edge device (server). The prompts sent by the user are detected on the edge device, and then the safe prompts are sent to the LLM in the cloud for service. The dataset downloaded in advance is saved in the VDB, and the saved format is: { }, where is the vectorized prompt, is the label of this prompt, indicates that the prompt is benign, indicates that the prompt is malicious. The prompts sent by the user are first detected on the edge node, and the dataset saved in the VDB is used to identify benign and malicious prompts.
[0045] The method, system, optimization terminal, and medium for optimizing prompt detection resources according to the embodiments of the present application will be described below with reference to the accompanying drawings. In view of the problems in the related art mentioned above, where synonym replacement or filtering of prompts in a large language model leads to low security of the large language model output and high energy consumption, the present application provides a method for optimizing prompt detection resources. In this method, by identifying the user's prompt and obtaining a detection strategy based on a multi-stage dynamic Bayesian game model, the security of the prompt is optimized while the system performance is optimized. Thus, malicious prompt detection can be effectively identified and blocked, the security of the LLM output can be improved, and the prompts that need to be detected can be intelligently selected to avoid unnecessary resource consumption, thereby reducing the service latency of benign users and enhancing the user experience. Therefore, the technical problems in the related art, where synonym replacement or filtering of prompts in a large language model leads to low security of the large language model output and high energy consumption, are solved.
[0046] The present application first proposes to jointly optimize system security, resource consumption, and service latency under prompt attacks, and formalizes the joint prompt detection, latency, and resource optimization problem as a multi-stage incomplete information Bayesian game model. To solve the Bayesian equilibrium at each stage, the present application makes a decision on whether to detect each incoming prompt through a belief update method and a malicious number prediction method. The prompt detection resource optimization method of the present application can improve the security of the LLM system output, reduce system resource consumption, and reduce the service latency of benign users.
[0047] The technical solution of the present application will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0048] The method for optimizing prompt detection resources according to a preferred embodiment of the present application is as Figure 1 shown. The method for optimizing prompt detection resources includes the following steps:
[0049] In step S101, receive the prompt sent by the user and identify the prompt parameters of the prompt.
[0050] Specifically, receive the prompt input by the user through a user interface (such as a web page, application program interface, command line interface, etc.); after preprocessing the prompt, parse the prompt to identify its structure and meaning; based on the parsed prompt, extract the key parameters related to subsequent processing as the input for subsequent processing (such as constructing a multi-stage dynamic Bayesian game model).
[0051] In step S102, construct a multi-stage dynamic Bayesian game model, and input the prompt parameters into the multi-stage dynamic Bayesian game model to obtain a detection probability.
[0052] In a possible implementation, consistency functions corresponding to benign users, malicious users, and defenders are constructed; the consistency functions corresponding to the benign users, the malicious users, and the defenders are minimized to obtain a multi-stage dynamic Bayesian game model.
[0053] It can be understood that this application models system security and performance issues using multi-round incomplete information dynamic Bayesian games, and defines the consistency functions of benign users, malicious users, and defenders.
[0054] Modeling the joint optimization of system security and performance issues using multi-round incomplete information dynamic Bayesian games, the benign user x 's consistency function is:
[0055] ; (1)
[0056] Among them, is the consistency function of the benign user, m represents an edge node, represents the set of edge nodes, represents the strategy of the benign user x (selecting a server), is the total delay of the strategy of whether to detect and whether to send the prompt word of the benign user x to the edge node m after processing by the large language model, represents the number of rounds of the game, is the total number of rounds of the game.
[0057] The consistency function of the malicious user y is: ; (2)
[0058] Among them,
[0059] is the consistency function of the malicious user, represents the strategy of the malicious user , y is the F1 score of the detection model, , represents the prompt word, represents the number of tokens of the prompt word, represents the number of floating-point operations per token.
[0060] The consistency function of the defender n is: ; (3)
[0061] Among them, is the consistency function of the defender, , represent two weight parameters greater than zero, represents the prompt word, represents the t sequence of prompt words received by the defender in the round, represents the defender t for each prompt word in the round whether to detect the strategy, represents the total delay for each prompt word to be processed, represents the defender's belief in each prompt word. x , y , n represent the indices of users or participants, used to distinguish different users or participants; token refers to the smallest unit into which text or data is segmented.
[0062] Therefore, the multi-stage dynamic Bayesian game model representing the minimization of the defender's consistency function, that is, the game problem can be formalized as:
[0063] ;
[0064] ;
[0065] ;
[0066] ; (4)
[0067] Among them, in the t round, except for the defender the benign users x and malicious users y for the prompt word whether to detect the strategy, represents the set of defenders, is the t number of malicious prompt words in the is the abbreviation of "subject to", indicating "subject to" or "meeting the conditions".
[0068] It can be understood that the three consistencies are the goals of benign users, malicious users, and defenders respectively. The goals correspond to the benefits of each player (i.e., benign users, malicious users, and defenders). Through game joint optimization, the benefits of the three are maximized. In formula (4), the variables are , two respectively represent the defendern 's strategy and other players - n 's strategy, each player (i.e., benign user, malicious user, and defender) has a payoff, and their strategies affect each other. The final result should be such that changing the strategy of each player will lead to a decrease in their own payoff. At this time, it is the equilibrium solution, which is the Bayesian equilibrium in the Bayesian game model.
[0069] In one possible implementation, according to the detection probability, update the beliefs of the benign user, the malicious user, and the defender about the security of the prompt word to obtain updated beliefs; use the updated beliefs as beliefs to obtain an updated multi-stage dynamic Bayesian game model, and input the next-round prompt parameters into the updated multi-stage dynamic Bayesian game model to obtain the next-round detection probability.
[0070] Specifically, this application uses the belief update method of incomplete information Bayesian game, and uses historical data and vector similarity to update beliefs. In this step, input data preparation: collect various relevant data from the previous round, including beliefs, delays, detection thresholds, token numbers, expected token numbers, security vector similarity, and insecure vector similarity; benign user processing: judge whether the previous-round reply is benign, set the initial likelihood function value according to the reply result, calculate the difference (diff) between the ratio of the delay to the token number and the detection threshold, and update the likelihood function value according to the positive or negative of diff, and ensure that the value is between 0 and 1; malicious user processing: judge whether the previous-round reply is malicious, set the initial likelihood function value according to the reply result, and update the likelihood function value using the ratio of the expected token number to the actual token number, and ensure that the value is between 0 and 1; defender processing: judge whether an anomaly was detected in the previous round and a malicious reply was given, set the initial likelihood function value according to the detection result, and update the likelihood function value using the vector similarity; belief update: calculate the new-round belief using the updated likelihood function value.
[0071] The belief update method of this application dynamically updates the beliefs of various players (benign users, malicious users, defenders) about the security of the prompt word by considering factors such as the reply result, processing delay, token number, and vector similarity from the previous round. This update mechanism helps the system more accurately judge the security of the current prompt word, so as to make more reasonable decisions.
[0072] Furthermore, the steps of the belief update method are as follows:
[0073] Step K11, input the t belief of the previous -1 round , the processing delay of the prompt word , the detection threshold t , the response of the previous The number of tokens , the number of tokens expected to be output by malicious users , the security vector similarity , the insecure vector similarity .
[0074] Step K12, for benign players, if the previous response was benign, define two likelihood functions , . Otherwise, , . Calculate , if diff is greater than zero, update and . Otherwise, update , . And ensure that the values of the two likelihood functions updated in both ways are between 0 and 1.
[0075] Step K13, for malicious players, if the previous response was malicious, define two likelihood functions , . Otherwise, , . Update , ,, and ensure that the values of the two updated likelihood functions are between 0 and 1.
[0076] Step K14, for the defender, if an anomaly was detected in the previous round and a malicious response was given, define two likelihood functions , , otherwise, , . Update the likelihood functions using the similarity , .
[0077] Step K15, for all types of players, update the belief .
[0078] Above, 、 、 、 、 、 are pre-defined parameters used to adjust the values of the likelihood functions in belief update; 、 respectively represent the likelihood functions used to update the belief; 、 respectively represent the security vector similarity and the insecure vector similarity;
[0079] Specifically, in order to solve formula (4), this application uses a prediction method for the number of malicious prompt words based on the sequential marginal analysis method. In this step, the input data is prepared as follows: collect the prompt word sequence and its length on node m in the round; initialization: set an initial cost parameter t ; loop to solve: traverse the prompt word sequence starting from 0, and solve the consistency function U of the defender for each possible number of malicious prompt words . If is 0, skip the current loop and continue to solve the next . Calculate the difference between the values corresponding to adjacent values . If is greater than the preset cost parameter , stop the loop; output result: return the last that meets the condition as the predicted number of malicious prompt words . .
[0080] The method for predicting the number of malicious prompt words in this application finds the number of malicious prompt words that maximizes the increase in the consistency function value by traversing the possible numbers of malicious prompt words and calculating the consistency function values corresponding to each number. This number is considered the number of malicious prompt words most likely to be encountered by the system in the current round. By setting the preset cost parameter , unnecessary computational overhead can be avoided while ensuring the prediction accuracy.
[0081] Furthermore, the steps of the method for predicting the number of malicious prompt words are as follows:
[0082] Step K21: Input the prompt word sequence t on node m in the round, and the sequence length is
[0083] Step K22: Initialize U .
[0084] Step K23: .
[0085] Step K24: Solve ,
[0086] Step K25: If , , return to the fourth step. Otherwise, calculate U = , , and enter the sixth step.
[0087] Step K26, if , stop the algorithm and output . Otherwise, return to Step K24.
[0088] As described above, represents a cost parameter, which is used as a stopping condition in the malicious prompt word quantity prediction algorithm; As a loop variable, it is used for iterative calculation.
[0089] The multi-stage dynamic Bayesian game model of the present application models the interactions of benign users, malicious users, and defenders, and obtains a detection strategy through a belief update method and a malicious number prediction method, optimizing the security of the prompt words while optimizing the system performance.
[0090] In step S103, according to the detection probability, the prompt word is detected to obtain a safe prompt word.
[0091] In a possible implementation manner, when the detection probability meets the preset requirements, the prompt word is detected to obtain a detection result; if the detection result is safe, it is determined that the prompt word is a safe prompt word. If the detection result is unsafe, the prompt word is returned for the next round of detection. That is, a decision is made based on the obtained detection probability.
[0092] It should be noted that the present application adopts a mixed strategy, and the obtained probability is generally 0 or 1; if not, a random number is used, a standard value is randomly generated by a computer, if the obtained value is greater than the standard value, the prompt word is detected, and if the obtained value is less than the standard value, the prompt word is not detected. It can be understood that generally, the result of the mixed strategy is a probability, which can be any number between 0 and 1. Therefore, in the present application, by predicting the malicious quantity, the sum of the probabilities is made equal to this number, so each probability becomes 1.
[0093] In step S104, the safe prompt word is input into a large language model for processing to obtain a response message.
[0094] Specifically, the prompt words confirmed to be safe after edge node detection are sent to the LLM (Large Language Model) in the cloud for processing.
[0095] The present application jointly optimizes LLM security and system resources, uses a multi-round incomplete information dynamic Bayesian model to model the problem, which fits the actual problem; in order to reduce the inaccurate detection problem caused by simply using the prompt word detection model, the present application combines the Bayesian belief update method of VDB and historical data, which can improve the accuracy of the detection decision; in order to optimize system resources, the present application adopts a malicious prompt word quantity prediction method and uses the sequential marginal analysis method, making the system resource consumption controllable.
[0096] As Figure 2 shown below, the above-mentioned prompt word detection resource optimization method of the present application will be further described through specific embodiments: In the first round of initializing beliefs, receive the prompt word, match the prompt word with the database, and determine whether to perform prompt word detection through the edge device. If so (i.e., safe), input it into the large model for response. If not, then judge whether the number of rounds t is less than the preset maximum number of rounds t max , if it is less, update the beliefs, detect the updated prompt word, and end if the number of rounds is equal to or greater than the preset maximum number of rounds.
[0097] Next, refer to the drawings to describe the prompt word detection resource optimization system according to the embodiments of the present application.
[0098] Figure 3 is the structural diagram of the prompt word detection resource optimization system according to the embodiments of the present application.
[0099] As Figure 3 shown, the prompt word detection resource optimization system includes: a prompt word receiving module 100, a model construction and optimization module 200, a detection strategy determination module 300, and a large model output module 400.
[0100] Specifically, the prompt word receiving module 100 is used to receive the prompt word sent by the user and identify the prompt parameters of the prompt word;
[0101] The model construction and optimization module 200 is used to construct a multi-stage dynamic Bayesian game model and input the prompt parameters into the multi-stage dynamic Bayesian game model to obtain the detection probability;
[0102] The detection strategy determination module 300 is used to detect the prompt word according to the detection probability to obtain a safe prompt word;
[0103] The large model output module 400 is used to input the safe prompt word into the large language model for processing to obtain response information.
[0104] Figure 4 is the structural diagram of the prompt word detection resource optimization terminal provided by the embodiments of the present application. The prompt word detection resource optimization terminal may include:
[0105] A memory 501, a processor 502, and a computer program stored on the memory 501 and executable on the processor 502.
[0106] When the processor 502 executes the program, it implements the prompt word detection resource optimization method provided in the above embodiments.
[0107] Further, the prompt word detection resource optimization terminal further includes:
[0108] A communication interface 503 for communication between the memory 501 and the processor 502.
[0109] A memory 501 for storing a computer program that can run on the processor 502.
[0110] The memory 501 may include high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory.
[0111] If the memory 501, the processor 502, and the communication interface 503 are implemented independently, the communication interface 503, the memory 501, and the processor 502 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component (PCI) bus, or an Extended Industry Standard Architecture (EIS) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0112] Optionally, in a specific implementation, if the memory 501, the processor 502, and the communication interface 503 are integrated on a chip, the memory 501, the processor 502, and the communication interface 503 can communicate with each other through an internal interface.
[0113] The processor 502 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0114] This embodiment also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned prompt word detection resource optimization method is implemented.
[0115] An embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the prompt word detection resource optimization method provided by any embodiment of the corresponding embodiment of the present application is implemented. Figure 1 The prompt word detection resource optimization method provided by any embodiment of the corresponding embodiment of the present application.
[0116] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0117] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0118] Any process or method description in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or N executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of this application includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in the reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art of the embodiments of this application.
[0119] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable storage medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable storage medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable storage media include the following: an electrical connection part (electronic device) having one or N wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable storage medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.
[0120] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.
[0121] Those of ordinary skill in the art of the present technology can understand that all or part of the steps carried by the methods of the above-described embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0122] In addition, each functional unit in various embodiments of the present application may be integrated into a processing module, or each unit may exist physically alone, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0123] The above-mentioned storage medium may be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
[0124] It should be understood that the application of the present application is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present application.
[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent substitutions for some or all of the technical features. And these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for optimizing prompt detection resources, characterized in that The method for optimizing prompt detection resources includes: Receiving a prompt sent by a user and identifying the prompt parameters of the prompt; Constructing a multi-stage dynamic Bayesian game model and inputting the prompt parameters into the multi-stage dynamic Bayesian game model to obtain a detection probability; Detecting the prompt according to the detection probability to obtain a safe prompt; Inputting the safe prompt into a large language model for processing to obtain a response message; The multi-stage dynamic Bayesian game model is expressed as: ; ; ; ; Among them, is the consistency function of the defender, indicating the defender 's strategy for whether to detect each prompt word in the t r-th round, indicates the strategy of benign users and malicious users other than the defender in the t r-th round for whether to detect the prompt word x y t represents the constraint condition, represents the set of defenders, represents the prompt word, represents the sequence of prompt words received by the defender in the t r-th round, is the number of malicious prompt words in the t r-th round, represents the defender's belief in each prompt word; The step of detecting the prompt according to the detection probability to obtain a safe prompt specifically includes: when the detection probability meets the preset requirements, detecting the prompt to obtain a detection result; if the detection result is safe, determining that the prompt is a safe prompt; The step of constructing the multi-stage dynamic Bayesian game model specifically includes: Constructing the respective consistency functions of a benign user, a malicious user, and a defender; Minimizing the respective consistency functions of the benign user, the malicious user, and the defender to obtain a multi-stage dynamic Bayesian game model; After inputting the prompt parameters into the multi-stage dynamic Bayesian game model to obtain a detection probability, the following steps are further included: Updating the beliefs of the benign user, the malicious user, and the defender about the safety of the prompt according to the detection probability to obtain updated beliefs; After model update according to the updated beliefs, obtaining an updated multi-stage dynamic Bayesian game model and inputting the next-round prompt parameters into the updated multi-stage dynamic Bayesian game model to obtain the next-round detection probability; The step of updating the beliefs of the benign user, the malicious user, and the defender about the safety of the prompt according to the detection probability to obtain updated beliefs specifically includes: Determining the reply result of the previous round according to the detection probability and setting an initial likelihood function value according to the reply result; If the reply result is confirmed to be benign, calculating the difference between the ratio of the delay to the number of tokens and the detection threshold, and updating the initial likelihood function value according to the positive or negative of the difference to obtain an updated likelihood function value; If the reply result is confirmed to be malicious, calculating the ratio of the expected number of tokens to the actual number of tokens, and updating the initial likelihood function value according to the ratio to obtain an updated likelihood function value; If the reply result is confirmed to be an abnormal and malicious reply, calculating the vector similarity, and updating the initial likelihood function value according to the vector similarity to obtain an updated likelihood function value; Obtaining updated beliefs according to the updated likelihood function value.
2. The prompt word detection resource optimization method according to claim 1, wherein The consistency function of the benign user is expressed as: ; The consistency function of the malicious user is expressed as: ; The consistency function of the defender is expressed as: ; Among them, is the consistency function of benign users, m represents an edge node, represents a set of edge nodes, represents a benign user x 's strategy, is to send the prompt words of benign users x to the edge nodes m and the total delay of the strategy of whether to detect and whether to send to the large language model for post-prompt word processing, represents the number of rounds of the game, is the total number of rounds of the game; is the consistency function of malicious users, represents a malicious user y 's strategy, is the F1 score of the detection model, represents the number of tokens of the prompt words, represents the number of floating-point operations per token; , represent two weight parameters greater than zero, represents the total delay of each prompt word being processed.
3. The prompting word detection resource optimization method according to claim 1, wherein After the step that when the detection probability meets the preset requirements, detecting the prompt to obtain a detection result, the following steps are further included: If the detection result is unsafe, returning the prompt for the next-round detection of the prompt.
4. A prompt detection resource optimization system, characterized in that, The prompt word detection resource optimization system is applied to the prompt word detection resource optimization method described in any one of claims 1-3; the prompt word detection resource optimization system includes: A prompt word receiving module, configured to receive a prompt word sent by a user and identify the prompt parameters of the prompt word; A model construction and optimization module, configured to construct a multi-stage dynamic Bayesian game model and input the prompt parameters into the multi-stage dynamic Bayesian game model to obtain a detection probability; A detection strategy determination module, configured to detect the prompt word according to the detection probability to obtain a safe prompt word; A large model output module, configured to input the safe prompt word into a large language model for processing to obtain a response message.
5. A prompt detection resource optimization terminal, characterized in that, The prompt word detection resource optimization terminal includes: a memory, a processor, and a prompt word detection resource optimization program stored on the memory and executable on the processor. When the prompt word detection resource optimization program is executed by the processor, the steps of the prompt word detection resource optimization method described in any one of claims 1-3 are implemented.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a prompt word detection resource optimization program. When the prompt word detection resource optimization program is executed by a processor, the steps of the prompt word detection resource optimization method described in any one of claims 1-3 are implemented.
Citation Information
Patent Citations
Sensitive question identification method and device based on large language model, equipment and medium
CN118821789A
Cited By
Bayesian automatic cue word optimization method based on meta cue word
CN120832540A
A bayesian automatic prompt word optimization method based on meta prompt words
CN120832540B