Vulnerability detection method and system for shielding word overflow of generative language model
By building topological networks and cluster analysis, combined with the inducing dialogue generation method, the problem of difficult reproducing and detecting blocked word overflow vulnerabilities in the generative language model is solved, and efficient vulnerability detection and analysis is achieved.
Patent Information
- Application Number
- CN202510538665.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-27
AI Technical Summary
The prior art is difficult to reproduce and locate the masked word overflow vulnerability in the generative language model, making it difficult to effectively detect and repair.
By building a topological network, analyzing the alternating text of the user and the language model, extracting the overflow path of the masked word, and achieving efficient detection of overflow of the masked word through clustering and inducing dialogue generation methods.
Effectively capture and reproduce the overflow path of masked words, improve the accuracy and efficiency of vulnerability detection, and enable dynamic modeling and analysis in complex scenarios.
Smart Images

Figure CN120046163A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and more specifically, it relates to a method and system for detecting vulnerabilities of blocked word overflow in a generative language model. Background Art
[0002] With the rapid development of generative language models, their wide applications in the fields of natural language generation, dialogue systems, text creation, etc. have put forward higher requirements for language standardization. However, since it is difficult to reproduce the situation of blocked word overflow in many cases, the vulnerabilities cannot be located. Therefore, there is a problem that it is difficult to reproduce blocked word overflow. Summary of the Invention
[0003] The present invention provides a method and system for detecting vulnerabilities of blocked word overflow in a generative language model to solve the technical problems raised in the background art.
[0004] The present invention provides a system for detecting vulnerabilities of blocked word overflow in a generative language model, including:
[0005] A data acquisition module for obtaining characteristic conversations of blocked word overflow of several users; wherein, the characteristic conversations include alternating texts of several rounds between users and the language model;
[0006] A text processing module for preprocessing each characteristic conversation to obtain several entity word vectors;
[0007] A topology construction module for constructing a topology network based on the entity word vectors;
[0008] A topology analysis module for analyzing each topology network to obtain the overflow paths of the blocked words;
[0009] A path clustering module for clustering the overflow paths to obtain characteristic overflow paths;
[0010] An overflow testing module for generating induced conversations for the language model based on the characteristic overflow paths; and generating vulnerability detection results of the blocked words according to the induced conversations.
[0011] Further, preprocessing each characteristic conversation includes:
[0012] Deleting the punctuation marks of the alternating texts in the characteristic conversations to obtain pure text data;
[0013] Performing part-of-speech recognition on the pure text data, and annotating the part-of-speech attributes of each word through natural language processing technology, including nouns, verbs, and adjectives, to obtain a part-of-speech recognition result;
[0014] Perform word segmentation on the plain text data based on the part-of-speech recognition results, and divide the plain text data into several independent word units or phrase units;
[0015] Map the word unit or phrase unit to the corresponding word vector representation to obtain the entity word vector.
[0016] Furthermore, construct a topological network based on the entity word vector. The topological network includes topological nodes, topological edges, and edge weights:
[0017] Step 301, obtain the masking words corresponding to the feature dialogue, and map the masking words to the topological nodes of the topological network; among them, the masking word vector corresponding to the masking word is used as the feature vector of the corresponding topological node;
[0018] Step 302, map each alternating text in the feature dialogue to the topological nodes of the topological network, and use the entity word vector corresponding to the alternating text as the feature vector of the corresponding topological node;
[0019] Step 303, calculate the correlation degree between the topological node of the masking word and the topological node of each alternating text. The calculation formula of the correlation degree is as follows: ;
[0020] Among them, represents the correlation degree between the topological node of the masking word and the topological node of the alternating text, represents the masking word vector of the masking word, represents the number of entity word vectors of the alternating text, represents the index of represents the th entity word vector in the alternating text, represents the th entity word vector in the alternating text and the cosine similarity between the masking word vector of the masking word, represents the operation of obtaining the modulus length of the vector, represents the non-linear term weight, represents the index of represents the th entity word vector and the th entity word vector in the alternating text and the cosine similarity, represents the th entity word vector and the cosine similarity between the masking word vector of the masking word in the alternating text, represents the activation function;
[0021] Step 304: Construct topological edges between the topological nodes of the alternating text at the t-th moment and the (t + 1)-th moment, and use the correlation degree between the topological nodes of the alternating text at the t-th moment and the topological nodes of the shielding words as the edge weights of the corresponding topological edges.
[0022] Further, analyze each topological network, including:
[0023] Step 401: Obtain the edge weights of any two adjacent topological edges in the topological network;
[0024] Step 402: If the difference between the edge weights is less than or equal to a preset threshold, splice the corresponding topological edges and delete the topological nodes between the corresponding topological edges;
[0025] Step 403: If the difference between the edge weights is greater than the preset threshold, keep the topological edges and topological nodes unchanged;
[0026] Step 404: Process the topological network based on Steps 402 and 403 to obtain the overflow path.
[0027] Further, cluster the overflow paths, including:
[0028] Cluster multiple overflow paths of the same shielding word to obtain several overflow paths corresponding to the shielding word;
[0029] Obtain the edge weight vectors of several overflow paths of the shielding word. The edge weight vector ; where represents the edge weight of the n-th topological edge in the overflow path;
[0030] Based on the edge weight vector, calculate the clustering degree of any two overflow paths. The calculation formula of the clustering degree is as follows: ; ; ;
[0031] where represents the clustering degree of any two overflow paths, represents the path similarity matrix, represents the element in the g-th row and h-th column of the path similarity matrix, represents the symmetric matrix of the path similarity matrix, represents the maximum eigenvalue of represents the degree matrix of the path similarity matrix, represents the operation of obtaining the Frobenius, represents ; represents the natural base represents the number of overflow paths corresponding to the masking word represents index of represents index of represents the edge weight vector of the g-th overflow path represents the edge weight vector of the k-th overflow path represents the transpose operation represents the operation of summing the vector elements;
[0032] If the clustering degree is greater than the preset clustering threshold, the corresponding overflow paths are grouped into one category to obtain several categories of overflow paths;
[0033] If the number of overflow paths in one category is greater than the preset number threshold, the corresponding category of overflow paths is used as the feature overflow paths.
[0034] Further, generating an induced dialogue for the language model includes:
[0035] Calculating the mean vector of the edge weight vectors corresponding to the overflow paths in several feature overflow paths, and the calculation formula of the mean vector is as follows: ; ;
[0036] where represents the mean vector of several feature overflow paths represents the n-th element in the mean vector represents the number of the n-th element of the mean vector of several feature overflow paths represents index of represents the n-th element of the k-th mean vector;
[0037] Constructing an induced dialogue based on the mean vector of several feature overflow paths.
[0038] Further, constructing an induced dialogue includes:
[0039] Step 701, obtaining the mean vector corresponding to several feature overflow paths of the masking word;
[0040] Step 702, taking the average value of the first elements in all the mean vectors as the correlation degree between the first text of the induced dialogue and the masking word, and randomly generating the first text;
[0041] Step 703: Obtain the second text which is the output of the language model for the first text, calculate the correlation degree between the second text and the masked word; and match the second element in all the mean vectors based on the correlation degree. Specifically, if the difference between the correlation degree between the second text and the masked word and the second element of the s-th mean vector is less than or equal to the preset difference, then the s-th mean vector is classified into the first induced set;
[0042] Step 704: Take the average value of the third element of the first induced set as the correlation degree between the third text of the induced dialogue and the masked word, and randomly generate the third text;
[0043] Step 705: Obtain the fourth text which is the output of the language model for the third text, calculate the correlation degree between the fourth text and the masked word; and match the fourth element in all the mean vectors based on the correlation degree. Specifically, if the difference between the correlation degree between the fourth text and the masked word and the fourth element of the s-th mean vector is less than or equal to the preset difference, then the s-th mean vector is classified into the second induced set;
[0044] Step 706: Repeat Step 704 and Step 705 until the u-th text or the r-th text is obtained; where the u-th text represents the text with the masked word output by the language model, and the r-th text represents the dimension of the mean vector in the (r - 1)-th induced set;
[0045] Step 707: If there is a masked word in any one of the u-th text or the r-th text, then obtain the vulnerability detection result.
[0046] A vulnerability detection method for masked word overflow of a generative language model, applied to the vulnerability detection system for masked word overflow of the generative language model, includes:
[0047] Step 1: Obtain the characteristic dialogues of masked word overflow of several users; where the characteristic dialogue includes several rounds of alternating texts between the user and the language model;
[0048] Step 2: Preprocess each characteristic dialogue to obtain several entity word vectors;
[0049] Step 3: Construct a topological network based on the entity word vectors;
[0050] Step 4: Analyze each topological network to obtain the overflow path of the masked word;
[0051] Step 5: Cluster the overflow paths to obtain the characteristic overflow paths;
[0052] Step 6: Generate an induced dialogue for the language model based on the characteristic overflow paths; generate the vulnerability detection result of the masked word according to the induced dialogue.
[0053] The beneficial effects of the present invention are as follows: By constructing the dialogue text with blocked word overflow into a topological network, multiple overflow paths in the dialogue text can be effectively captured. Subsequently, through cluster analysis of the overflow paths, characteristic overflow paths are extracted, and an induced dialogue with blocked word overflow is further designed, thereby realizing a detection method that can efficiently reproduce the blocked word overflow vulnerability. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 is a module diagram of a vulnerability detection system for blocked word overflow of a generative language model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] Now, the subject matter described herein will be discussed with reference to exemplary embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Without departing from the protection scope of the content of this specification, changes can be made to the functions and arrangements of the elements discussed. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described relative to some examples can also be combined in other examples.
[0056] As Figure 1 shown, a vulnerability detection system for blocked word overflow of a generative language model includes:
[0057] A data collection module for obtaining characteristic dialogues of blocked word overflow of several users; wherein, the characteristic dialogues include alternating texts of several rounds between users and the language model.
[0058] A text processing module for preprocessing each characteristic dialogue to obtain several entity word vectors.
[0059] A topology construction module for constructing a topological network based on the entity word vectors.
[0060] A topology analysis module for analyzing each topological network to obtain the overflow paths of the blocked words.
[0061] A path clustering module for clustering the overflow paths to obtain characteristic overflow paths.
[0062] An overflow test module for generating an induced dialogue for the language model based on the characteristic overflow paths; and generating a vulnerability detection result of the blocked word according to the induced dialogue.
[0063] In an embodiment of the present invention, preprocessing each characteristic dialogue includes:
[0064] Deleting the punctuation marks of the alternating texts in the characteristic dialogue to obtain pure text data.
[0065] Perform part-of-speech recognition on pure text data, and label the part-of-speech attributes of each word through natural language processing technology, including nouns, verbs, and adjectives, to obtain the part-of-speech recognition result;
[0066] Perform word segmentation on the pure text data based on the part-of-speech recognition result, and divide the pure text data into several independent word units or phrase units;
[0067] Map the word unit or phrase unit to the corresponding word vector representation to obtain the entity word vector.
[0068] Specifically, first, for the alternating text in the feature conversation, delete all punctuation marks (such as commas, periods, quotation marks, etc.) to clean up the interference of meaningless symbols, so as to obtain pure text data consisting only of words or phrases. The purpose of this step is to simplify the text structure and ensure that subsequent processing is more efficient and focused. Apply natural language processing technology to the pure text data for part-of-speech tagging to identify the part-of-speech attributes of each word, such as nouns, verbs, adjectives, etc. For example, through part-of-speech recognition technology, "blocking word" can be tagged as a noun, "trigger" as a verb, and "potential" as an adjective. The goal of this step is to extract the language attribute features in the text for subsequent word segmentation and semantic analysis. Based on the result of part-of-speech recognition, divide the pure text data into several independent word units or phrase units. For example, the sentence "Blocking words may be triggered" is segmented into three units: "blocking word", "may", and "be triggered". The main purpose of word segmentation processing is to convert continuous text into operable basic language units, laying the foundation for subsequent vectorization operations. Map the word units or phrase units after word segmentation processing to a high-dimensional vector representation (i.e., word vector), and convert the language information in the text into a computable numerical form through a pre-trained word embedding model (such as Word2Vec, GloVe, or BERT, etc.). The result of this step is to generate the entity word vector corresponding to each word unit or phrase unit for the construction and analysis of the subsequent topology network.
[0069] In an embodiment of the present invention, a topology network is constructed based on the entity word vector. The topology network includes topology nodes, topology edges, and edge weights:
[0070] Step 301, obtain the blocking words corresponding to the feature conversation, and map the blocking words to the topology nodes of the topology network; wherein, the blocking word vector corresponding to the blocking word is used as the feature vector of the corresponding topology node;
[0071] Step 302, map each alternating text in the feature conversation to the topology nodes of the topology network, and use the entity word vector corresponding to the alternating text as the feature vector of the corresponding topology node;
[0072] Step 303, calculate the correlation degree between the topology node of the blocking word and the topology nodes of each alternating text. The calculation formula of the correlation degree is as follows: ;
[0073] Wherein, represents the correlation degree between the topological node of the shielding word and the topological node of the alternative text, represents the shielding word vector of the shielding word, represents the number of entity word vectors of the alternative text, represents the index of represents the th entity word vector in the alternative text, represents the th entity word vector in the alternative text and the cosine similarity between the shielding word vectors of the shielding word, represents the operation of obtaining the modulus of the vector, represents the non-linear term weight, represents the index of represents the th entity word vector in the alternative text and the th entity word vector and the cosine similarity between them, represents the th entity word vector in the alternative text and the cosine similarity between the shielding word vectors of the shielding word, represents the activation function;
[0074] Step 304, construct a topological edge between the topological nodes of the alternative text at the t-th moment and the (t + 1)-th moment, and use the correlation degree between the topological node of the alternative text at the t-th moment and the topological node of the shielding word as the edge weight of the corresponding topological edge.
[0075] Specifically, the construction of the topological network has the following remarkable beneficial effects:
[0076] Enhance the representational ability of the masked word overflow path: Through the form of nodes and edges, the topological network explicitly structures the word units and their semantic relationships in the text. This structure can more intuitively capture the potential patterns of the masked word overflow path. Compared with simple word vector analysis or context statistical methods, the topological network can reflect higher-level semantic associations, providing a richer and more comprehensive data basis for subsequent path analysis. Improve the accuracy of masked word overflow detection: By constructing a topological network, the explicit masked word overflow paths in the text can be captured. Implement dynamic modeling in complex scenarios: Regardless of the complexity of the dialogue text, this method can transform the language information in the dialogue text into a unified topological structure, facilitating subsequent clustering and analysis. Provide strong support for path clustering and overflow detection: The constructed topological network is not only used to capture overflow paths but also lays a foundation for subsequent clustering analysis and feature extraction of paths. The weights of each node and edge in the network can accurately quantify the semantic characteristics of word units and their interactions. This high-precision representation improves the reliability of the masked word overflow detection system.
[0077] In one embodiment of the present invention, analyzing each topological network includes:
[0078] Step 401, obtaining the edge weights of any two adjacent topological edges in the topological network;
[0079] Step 402, if the difference between the edge weights is less than or equal to a preset threshold, splicing the corresponding topological edges and deleting the topological nodes between the corresponding topological edges;
[0080] Step 403, if the difference between the edge weights is greater than the preset threshold, keeping the topological edges and topological nodes unchanged;
[0081] Step 404, processing the topological network based on Steps 402 and 403 to obtain the overflow path.
[0082] For example, there is a topological network "E1→E2→E3→E4→E5"; where "→" represents a topological edge, and "E1, E2, E3, E4, E5" all represent topological nodes; among them, the edge weights of the two topological edges in "→E2→" are 0.4 and 0.43 respectively, and the difference is 0.03 which is less than the threshold 0.05, so splicing is performed and E2 is deleted to obtain "E1→E3→E4→E5", and the edge weight of the topological edge in "E1→E3" is 0.415.
[0083] In one embodiment of the present invention, clustering the overflow paths includes:
[0084] Clustering multiple overflow paths of the same masked word to obtain several overflow paths corresponding to the masked word;
[0085] Obtain the edge weight vectors of several overflow paths of the stop words, and the edge weight vectors ; where represents the edge weight of the nth topological edge in the overflow path;
[0086] Based on the edge weight vectors, calculate the clustering degree of any two overflow paths. The calculation formula of the clustering degree is as follows: ; ; ;
[0087] where represents the clustering degree of any two overflow paths, represents the path similarity matrix, represents the element in the gth row and hth column of the path similarity matrix, represents the symmetric matrix of the path similarity matrix, represents the maximum eigenvalue of, represents the degree matrix of the path similarity matrix, represents the operation of obtaining the Frobenius, represents , represents the natural base, represents the number of overflow paths corresponding to the stop words, represents the index of, represents the index of, represents the edge weight vector of the gth overflow path, represents the th edge weight vector of the overflow path, represents the transpose operation, represents the operation of obtaining the sum of vector elements;
[0088] If the clustering degree is greater than the preset clustering threshold, then classify the corresponding overflow paths into one category to obtain several categories of overflow paths;
[0089] If the number of overflow paths in one category is greater than the preset number threshold, then use the corresponding category of overflow paths as the characteristic overflow paths.
[0090] For example, for the blocked word "nightclub", "nightclub" has a total of 10 overflow paths. Clustering the 10 overflow paths, if the clustering degrees of overflow path 1 and overflow path 2 are 0.8, and 0.8 is greater than the preset clustering threshold of 0.75, then overflow path 1 and overflow path 2 are grouped into one category. If the clustering degrees of overflow path 3 with overflow path 1 and overflow path 2 are 0.87 and 0.43 respectively, then overflow path 3 and overflow path 1 are grouped into one category. Then two clusters are obtained, one cluster is overflow path 1 and overflow path 2, and one cluster is overflow path 1 and overflow path 3. If the clustering degrees of overflow path 3 with overflow path 1 and overflow path 2 are 0.87 and 0.77 respectively, then one cluster is obtained, and the cluster includes overflow path 1, overflow path 2, and overflow path 3. Through the feature extraction and clustering analysis of the overflow paths, representative characteristic overflow paths are summarized, improving the systematicness and efficiency of overflow detection.
[0091] In one embodiment of the present invention, generating an induced dialogue for a language model includes:
[0092] Calculating the mean vector of the edge weight vectors corresponding to the overflow paths among several characteristic overflow paths, and the calculation formula of the mean vector is as follows: ; ;
[0093] Wherein, represents the mean vector of several characteristic overflow paths, represents the nth element in the mean vector, represents the number of the nth element of the mean vector of several characteristic overflow paths, represents the index of, represents the nth element of the kth mean vector;
[0094] Based on the mean vector of several characteristic overflow paths, constructing an induced dialogue.
[0095] Specifically, if the edge weights of the topological edges of "E1→E3→E4→E5" are "0.415, 0.46, 0.65, 0.82" respectively, then the edge weight vector is . Calculating the mean vector based on the edge weight vector. For example, the characteristic overflow paths of the blocked word include 2 overflow paths, and the edge weight vectors are respectively and ; then the mean vector is .
[0096] In one embodiment of the present invention, constructing an induced dialogue includes:
[0097] Step 701: Obtain the mean vectors corresponding to several feature overflow paths of the masking words;
[0098] Step 702: Take the average of the first elements in all the mean vectors as the correlation degree between the first text of the induced dialogue and the masking words, and randomly generate the first text;
[0099] Step 703: Obtain the second text output by the language model for the first text, and calculate the correlation degree between the second text and the masking words; and match the second elements in all the mean vectors based on the correlation degree. Specifically, if the difference between the correlation degree between the second text and the masking words and the second element of the s-th mean vector is less than or equal to the preset difference, then the s-th mean vector is classified into the first induced set;
[0100] Step 704: Take the average of the third elements in the first induced set as the correlation degree between the third text of the induced dialogue and the masking words, and randomly generate the third text;
[0101] Step 705: Obtain the fourth text output by the language model for the third text, and calculate the correlation degree between the fourth text and the masking words; and match the fourth elements in all the mean vectors based on the correlation degree. Specifically, if the difference between the correlation degree between the fourth text and the masking words and the fourth element of the s-th mean vector is less than or equal to the preset difference, then the s-th mean vector is classified into the second induced set;
[0102] Step 706: Repeat Step 704 and Step 705 until the u-th text or the r-th text is obtained; where the u-th text represents the text with masking words output by the language model, and the r-th text represents the dimension of the mean vector in the (r - 1)-th induced set;
[0103] Step 707: If there is a masking word in any of the u-th text or the r-th text, obtain the vulnerability detection result.
[0104] Specifically, based on 4 mean vectors of the masking words, construct an induced dialogue as follows: ; ;
[0105] Establish the first text: Determine that the correlation degree between the first text and the masking words is ; Generate the first text based on the correlation degree, which can be constructed by experts or a third-party language model.
[0106] Take the first text as the input of the user, and then obtain the second text fed back by the language model.
[0107] Calculate the correlation between the second text and the masking word. If the correlation is 0.44, then the second mean vector, the third mean vector, and the fourth mean vector are all used as the first induction set.
[0108] Establish a third text and determine that the correlation between the third text and the masking word is ; Generate the third text based on the correlation.
[0109] If the correlation between the fourth text feedback by the language model and the masking word is then the second induction set includes the second mean vector and the third mean vector.
[0110] Establish a fifth text and determine that the correlation between the fifth text and the masking word is ; Generate the fifth text based on the correlation.
[0111] If the sixth text has a masking word, then a detection of the vulnerability is obtained.
[0112] A method for detecting vulnerabilities of masking word overflow in a generative language model, which is applied to the system for detecting vulnerabilities of masking word overflow in the generative language model, includes:
[0113] Step 1: Obtain characteristic dialogues of masking word overflow of several users; wherein, the characteristic dialogues include alternating texts of users and the language model in several rounds.
[0114] Step 2: Preprocess each characteristic dialogue to obtain several entity word vectors.
[0115] Step 3: Construct a topological network based on the entity word vectors.
[0116] Step 4: Analyze each topological network to obtain the overflow path of the masking word.
[0117] Step 5: Cluster the overflow paths to obtain characteristic overflow paths.
[0118] Step 6: Generate an induced dialogue for the language model based on the characteristic overflow paths; generate a vulnerability detection result of the masking word according to the induced dialogue.
[0119] The above has described the embodiments of this example, but this example is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of this example, those of ordinary skill in the art can also make many forms, all of which fall within the protection scope of this example.
Claims
1. A vulnerability detection system for overflow of blocked words in a generative language model, characterized in that: include: A data collection module is used to obtain characteristic dialogues of blocked word overflows of several users; wherein the characteristic dialogues include alternating texts of several rounds of users and language models; The text processing module is used to pre-process each feature dialogue to obtain several entity word vectors; Topology building module, used to build a topology network based on entity word vectors; The topological analysis module is used to analyze each topological network and obtain the overflow path of the blocked words; A path clustering module is used to cluster the overflow paths and obtain characteristic overflow paths; The overflow test module is used to generate induced dialogues for the language model based on the feature overflow path; and to generate vulnerability detection results of blocked words based on the induced dialogues.
2. According to claim 1, a generative language model shielding word overflow vulnerability detection system is characterized in that: Preprocess each feature dialogue, including: Remove punctuation marks from alternate text in feature dialogues to obtain plain text data; Perform part-of-speech recognition on plain text data, and use natural language processing technology to annotate the part-of-speech attributes of each word, including nouns, verbs, and adjectives, to obtain part-of-speech recognition results; Based on the part-of-speech recognition results, the plain text data is segmented into several independent word units or phrase units; Map word units or phrase units to corresponding word vector representations to obtain entity word vectors.
3. According to claim 2, a generative language model shielding word overflow vulnerability detection system is characterized in that: A topological network is constructed based on entity word vectors. The topological network includes topological nodes, topological edges, and edge weights: Step 301, obtaining the shielded words corresponding to the feature dialogue, and mapping the shielded words to the topological nodes of the topological network; wherein the shielded word vector corresponding to the shielded word is used as the feature vector of the corresponding topological node; Step 302, mapping each alternate text in the feature dialogue to a topological node of the topological network, and using the entity word vector corresponding to the alternate text as the feature vector of the corresponding topological node; Step 303, calculate the correlation between the topological node of the blocked word and the topological node of each alternate text, and the calculation formula of the correlation is as follows: ; in, The correlation between the topological node of the blocked word and the topological node of the alternate text, The masked word vector representing the masked word, The number of entity word vectors representing alternate texts, express The index of Indicates the first entity word vectors, Indicates the first The cosine similarity between the entity word vector and the shielded word vector of the shielded word, represents the operation of finding the modulus length of the orientation vector, represents the weight of nonlinear term, express The index of Indicates the first entity word vector and The cosine similarity of entity word vectors, Indicates the first The cosine similarity between the entity word vector and the shielded word vector of the shielded word, express Activation function; Step 304, construct a topological edge between the topological nodes of the alternating text at the tth moment and the t+1th moment, and use the correlation between the topological node of the alternating text at the tth moment and the topological node of the blocked word as the edge weight of the corresponding topological edge.
4. According to claim 3, a generative language model shielding word overflow vulnerability detection system is characterized in that: Each topology network is analyzed, including: Step 401, obtaining edge weights of any two adjacent topological edges in the topological network; Step 402: if the difference between the edge weights is less than or equal to a preset threshold, the corresponding topological edges are spliced, and the topological nodes between the corresponding topological edges are deleted; Step 403: if the difference between the edge weights is greater than a preset threshold, the topological edge and the topological node are kept unchanged; Step 404: Process the topology network based on steps 402 and 403 to obtain an overflow path.
5. According to claim 4, a generative language model shielding word overflow vulnerability detection system is characterized in that: Clustering of overflow paths, including: Clustering multiple overflow paths of the same blocked word to obtain several overflow paths of the corresponding blocked word; Get the edge weight vectors of several overflow paths of the blocked words, the edge weight vector ;in, Represents the edge weight of the nth topological edge in the overflow path; Based on the edge weight vector, the clustering degree of any two overflow paths is calculated. The calculation formula of the clustering degree is as follows: ; ; ; in, represents the clustering degree of any two overflow paths, represents the path similarity matrix, represents the element in the gth row and hth column in the path similarity matrix, A symmetric matrix representing the path similarity matrix, express The maximum eigenvalue of The degree matrix representing the path similarity matrix, represents the operation of obtaining Frobenius. express , represents the natural base, Indicates the number of overflow paths corresponding to the blocked words, express The index of express The index of represents the edge weight vector of the g-th overflow path, Indicates The edge weight vector of the overflow path, represents the transpose operation, represents the operation of finding the sum of vector elements; If the clustering degree is greater than the preset clustering threshold, the corresponding overflow paths are classified into one category to obtain several categories of overflow paths; If the number of overflow paths of a type is greater than a preset number threshold, the corresponding overflow path of the type is used as a characteristic overflow path.
6. According to claim 5, a generative language model shielding word overflow vulnerability detection system is characterized in that: Generate induced dialogue for the language model, including: Calculate the mean vector of the edge weight vectors corresponding to the overflow paths in several feature overflow paths. The calculation formula of the mean vector is as follows: ; ; in, represents the mean vector of several feature overflow paths, represents the nth element in the mean vector, Represents the number of n-th elements of the mean vector of several feature overflow paths, express The index of represents the nth element of the kth mean vector; Based on the mean vector of several feature overflow paths, an induced dialogue is constructed.
7. The system for detecting a vulnerability of a generative language model for overflow of blocked words according to claim 6, characterized in that: Construct an inducing dialogue that includes: Step 701, obtaining mean vectors corresponding to several feature overflow paths of the blocked words; Step 702, taking the average of the first elements in all mean vectors as the correlation between the first text of the induced dialogue and the blocked words, and randomly generating the first text; Step 703, obtaining a second text output by the language model for the first text, calculating the relevance between the second text and the blocked word; and matching the second elements in all mean vectors based on the relevance, specifically including: if the difference between the relevance between the second text and the blocked word and the second element of the s-th mean vector is less than or equal to a preset difference, then the s-th mean vector is classified into the first induced set; Step 704, taking the average value of the third element of the first induced set as the correlation between the third text of the induced dialogue and the blocked word, and randomly generating the third text; Step 705, obtaining a fourth text output by the language model for the third text, calculating the relevance between the fourth text and the blocked word; and matching the fourth elements in all mean vectors based on the relevance, specifically including: if the difference between the relevance between the fourth text and the blocked word and the fourth element of the s-th mean vector is less than or equal to a preset difference, then the s-th mean vector is classified into the second induced set; Step 706, repeating steps 704 and 705 until the uth text or the rth text is obtained; wherein the uth text represents the text output by the language model with blocked words, and the rth text represents the dimension of the mean vector in the r-1th induced set; Step 707: If the blocked word exists in any of the u-th text or the r-th text, a vulnerability detection result is obtained.
8. A method for detecting a vulnerability of a generative language model for overflow of a shielding word, applied to a system for detecting a vulnerability of a generative language model for overflow of a shielding word as claimed in any one of claims 1 to 7, characterized in that: include: Step 1, obtaining characteristic dialogues of blocked word overflow of several users; wherein the characteristic dialogues include alternating texts of several rounds of users and language models; Step 2: Preprocess each feature dialogue to obtain several entity word vectors; Step 3, construct a topological network based on entity word vectors; Step 4, analyze each topological network to obtain the overflow path of the blocked words; Step 5, clustering the overflow paths to obtain characteristic overflow paths; Step 6: Generate an induced dialogue for the language model based on the feature overflow path; and generate a vulnerability detection result of the blocked words based on the induced dialogue.
Citation Information
Patent Citations
Buffer overflow loophole automatic detection method based on symbolic execution path pruning
CN104732152A
Data security analysis method and intelligent calculation data security workstation
CN119249440A
Nested named entity recognition method based on part-of-speech awareness, device and storage medium therefor
US20240111956A1
Methods, apparatuses and computer program products for use in evaluating the coverage provided by base stations of a cellular radio access network
WO2023247928A1