Vulnerability Detection Method and System for Mask Word Overflow of a Generative Language Model

By building topological networks and path clustering analysis, induced dialogues are generated, and the problem of difficult reproducibility of blocked word overflow in the generative language model is solved, and efficient vulnerability detection is achieved.

CN120046163BActive Publication Date: 2025-07-01ZHEJIANG QIANGUA INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510538665.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-01
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

In the generative language model, it is difficult to reproduce the overflow of blocked words, which leads to difficulty in positioning the vulnerability.

Method used

The user's blocked word overflow feature dialogue is obtained through the data acquisition module, text processing and part-of-speech recognition, topological network is built, overflow paths are analyzed, and path clustering is performed, and induced dialogue is finally generated to detect vulnerabilities.

Benefits of technology

It realizes vulnerability detection of efficient reproduction of blocking word overflow, improving detection accuracy and systematicity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046163B_ABST
    Figure CN120046163B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and discloses a method and system for detecting vulnerabilities of blocked word overflow in a generative language model. The data acquisition module is used to obtain characteristic conversations of blocked word overflow of a number of users; wherein, the characteristic conversations include alternating texts of a number of rounds between users and the language model. The text processing module is used to preprocess each characteristic conversation to obtain a number of entity word vectors. The topology construction module is used to construct a topology network based on the entity word vectors. The topology analysis module is used to analyze each topology network to obtain the overflow path of the blocked word. The path clustering module is used to cluster the overflow paths to obtain characteristic overflow paths. The overflow test module is used to generate an induced conversation for the language model based on the characteristic overflow paths, and generate a vulnerability detection result of the blocked word according to the induced conversation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and more specifically, it relates to a method and system for detecting vulnerabilities of blocked word overflow in a generative language model. Background Art

[0002] With the rapid development of generative language models, their wide applications in the fields of natural language generation, dialogue systems, text creation, etc. have put forward higher requirements for language norms. However, since it is difficult to reproduce the situation of blocked word overflow in many cases, the vulnerabilities cannot be located. Therefore, there is a problem that it is difficult to reproduce blocked word overflow. Summary of the Invention

[0003] The present invention provides a method and system for detecting vulnerabilities of blocked word overflow in a generative language model, and solves the technical problems raised in the background art.

[0004] The present invention provides a system for detecting vulnerabilities of blocked word overflow in a generative language model, including:

[0005] A data collection module, configured to obtain characteristic conversations of blocked word overflow of a number of users; wherein, the characteristic conversations include alternating texts of a number of rounds between users and the language model;

[0006] A text processing module, configured to preprocess each characteristic conversation to obtain a number of entity word vectors;

[0007] A topology construction module, configured to construct a topology network based on the entity word vectors;

[0008] A topology analysis module, configured to analyze each topology network to obtain the overflow path of the blocked word;

[0009] A path clustering module, configured to cluster the overflow paths to obtain characteristic overflow paths;

[0010] An overflow test module, configured to generate induced conversations for the language model based on the characteristic overflow paths; and generate a vulnerability detection result of the blocked word according to the induced conversations.

[0011] Further, preprocessing each characteristic conversation includes:

[0012] Deleting the punctuation marks of the alternating texts in the characteristic conversations to obtain pure text data;

[0013] Performing part-of-speech recognition on the pure text data, and annotating the part-of-speech attributes of each word through natural language processing technology, including nouns, verbs, and adjectives, to obtain a part-of-speech recognition result;

[0014] Perform word segmentation on the plain text data based on the part-of-speech recognition results, and divide the plain text data into several independent word units or phrase units;

[0015] Map the word units or phrase units to corresponding word vector representations to obtain entity word vectors.

[0016] Furthermore, construct a topological network based on the entity word vectors. The topological network includes topological nodes, topological edges, and edge weights:

[0017] Step 301: Obtain the masked words corresponding to the feature dialogue, and map the masked words to the topological nodes of the topological network; among them, the masked word vectors corresponding to the masked words are used as the feature vectors of the corresponding topological nodes;

[0018] Step 302: Map each alternating text in the feature dialogue to the topological nodes of the topological network, and use the entity word vectors corresponding to the alternating texts as the feature vectors of the corresponding topological nodes;

[0019] Step 303: Calculate the correlation degree between the topological nodes of the masked words and the topological nodes of each alternating text. The calculation formula for the correlation degree is as follows:

[0020] ;

[0021] Among them, represents the correlation degree between the topological nodes of the masked words and the topological nodes of the alternating texts, represents the masked word vector of the masked words, represents the number of entity word vectors of the alternating texts, represents the index of represents the th entity word vector in the alternating text, represents the th entity word vector in the alternating text and the cosine similarity between the masked word vector of the masked words, represents the operation of obtaining the modulus length of the vector, represents the non-linear term weight, represents the index of represents the th entity word vector in the alternating text and the th entity word vector and the cosine similarity, represents the th entity word vector in the alternating text and the cosine similarity between the masked word vector of the masked words, represents the activation function;

[0022] Step 304, construct a topological edge between the topological nodes of the alternating text at the tth moment and the t+1th moment, and use the correlation between the topological node of the alternating text at the tth moment and the topological node of the blocked word as the edge weight of the corresponding topological edge.

[0023] Furthermore, each topology network is analyzed, including:

[0024] Step 401, obtaining edge weights of any two adjacent topological edges in the topological network;

[0025] Step 402: if the difference between the edge weights is less than or equal to a preset threshold, the corresponding topological edges are spliced, and the topological nodes between the corresponding topological edges are deleted;

[0026] Step 403: if the difference between the edge weights is greater than a preset threshold, the topological edge and the topological node are kept unchanged;

[0027] Step 404: Process the topology network based on steps 402 and 403 to obtain an overflow path.

[0028] Furthermore, the overflow paths are clustered, including:

[0029] Clustering multiple overflow paths of the same blocked word to obtain several overflow paths of the corresponding blocked word;

[0030] Get the edge weight vectors of several overflow paths of the blocked words, the edge weight vector ;in, Represents the edge weight of the nth topological edge in the overflow path;

[0031] Based on the edge weight vector, the clustering degree of any two overflow paths is calculated. The calculation formula of the clustering degree is as follows:

[0032] ;

[0033] ;

[0034] ;

[0035] in, represents the clustering degree of any two overflow paths, represents the path similarity matrix, represents the element in the gth row and hth column in the path similarity matrix, A symmetric matrix representing the path similarity matrix, express The maximum eigenvalue of The degree matrix representing the path similarity matrix, represents the operation of obtaining Frobenius. denote , denotes the natural base, denotes the number of overflow paths corresponding to the masked word, denote the index of denote the index of denotes the edge weight vector of the g-th overflow path, denote the -th overflow path's edge weight vector, denotes the transpose operation, denotes the operation of obtaining the sum of vector elements;

[0036] If the clustering degree is greater than the preset clustering threshold, the corresponding overflow paths are grouped into one category to obtain several categories of overflow paths;

[0037] If the number of overflow paths in a category is greater than the preset quantity threshold, the corresponding category of overflow paths is used as the feature overflow paths.

[0038] Furthermore, generating an induced dialogue for the language model includes:

[0039] Calculating the mean vector of the edge weight vectors corresponding to the overflow paths in several feature overflow paths, and the calculation formula of the mean vector is as follows:

[0040] ;

[0041] ;

[0042] where, denotes the mean vector of several feature overflow paths, denotes the n-th element in the mean vector, denotes the number of the n-th element of the mean vector of several feature overflow paths, denote the index of denotes the n-th element of the k-th mean vector;

[0043] Constructing an induced dialogue based on the mean vector of several feature overflow paths.

[0044] Furthermore, constructing an induced dialogue includes:

[0045] Step 701, obtaining the mean vector corresponding to several feature overflow paths of the masked word;

[0046] Step 702, taking the average of the first elements in all the mean vectors as the correlation degree between the first text of the induced dialogue and the masked word, and randomly generating the first text;

[0047] Step 703: Obtain a second text which is the output of the language model for the first text, and calculate the correlation between the second text and the masked word; and match the second element in all the mean vectors based on the correlation. Specifically, if the difference between the correlation between the second text and the masked word and the second element of the sth mean vector is less than or equal to a preset difference, then the sth mean vector is classified into the first induced set;

[0048] Step 704: Take the average value of the third element of the first induced set as the correlation between the third text of the induced conversation and the masked word, and randomly generate the third text;

[0049] Step 705: Obtain a fourth text which is the output of the language model for the third text, and calculate the correlation between the fourth text and the masked word; and match the fourth element in all the mean vectors based on the correlation. Specifically, if the difference between the correlation between the fourth text and the masked word and the fourth element of the sth mean vector is less than or equal to a preset difference, then the sth mean vector is classified into the second induced set;

[0050] Step 706: Repeat Step 704 and Step 705 until the u-th text or the r-th text is obtained; where the u-th text represents the text with the masked word output by the language model, and the r-th text represents the dimension of the mean vector in the (r - 1)-th induced set;

[0051] Step 707: If there is a masked word in any one of the u-th text or the r-th text, then obtain the vulnerability detection result.

[0052] A vulnerability detection method for masked word overflow of a generative language model, which is applied to the vulnerability detection system for masked word overflow of the generative language model, includes:

[0053] Step 1: Obtain the characteristic conversations of masked word overflow of several users; where the characteristic conversations include several rounds of alternating texts between the user and the language model;

[0054] Step 2: Preprocess each characteristic conversation to obtain several entity word vectors;

[0055] Step 3: Construct a topological network based on the entity word vectors;

[0056] Step 4: Analyze each topological network to obtain the overflow path of the masked word;

[0057] Step 5: Cluster the overflow paths to obtain the characteristic overflow paths;

[0058] Step 6: Generate an induced conversation for the language model based on the characteristic overflow paths; and generate a vulnerability detection result of the masked word according to the induced conversation.

[0059] The beneficial effects of the present invention are as follows: By constructing the dialogue text with blocked word overflow into a topological network, multiple overflow paths in the dialogue text can be effectively captured. Subsequently, through cluster analysis of the overflow paths, characteristic overflow paths are extracted, and an induced dialogue with blocked word overflow is further designed, thereby realizing a detection method capable of efficiently reproducing the blocked word overflow vulnerability. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 It is a module diagram of a vulnerability detection system for blocked word overflow of a generative language model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0061] Now, the subject matter described herein will be discussed with reference to exemplary embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Without departing from the scope of protection of the content of this specification, changes can be made to the functions and arrangements of the elements discussed. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described relative to some examples can also be combined in other examples.

[0062] As Figure 1 shown, a vulnerability detection system for blocked word overflow of a generative language model includes:

[0063] A data acquisition module for obtaining characteristic dialogues with blocked word overflow of several users; wherein, the characteristic dialogues include alternating texts of several rounds between users and the language model;

[0064] A text processing module for preprocessing each characteristic dialogue to obtain several entity word vectors;

[0065] A topology construction module for constructing a topological network based on the entity word vectors;

[0066] A topology analysis module for analyzing each topological network to obtain the overflow paths of the blocked words;

[0067] A path clustering module for clustering the overflow paths to obtain characteristic overflow paths;

[0068] An overflow test module for generating an induced dialogue for the language model based on the characteristic overflow paths; and generating a vulnerability detection result of the blocked words according to the induced dialogue.

[0069] In an embodiment of the present invention, preprocessing each characteristic dialogue includes:

[0070] Deleting the punctuation marks of the alternating texts in the characteristic dialogue to obtain pure text data;

[0071] Perform part-of-speech recognition on the plain text data, and label the part-of-speech attributes of each word through natural language processing technology, including nouns, verbs, and adjectives, to obtain the part-of-speech recognition result;

[0072] Perform word segmentation on the plain text data based on the part-of-speech recognition result, and divide the plain text data into several independent word units or phrase units;

[0073] Map the word unit or phrase unit to the corresponding word vector representation to obtain the entity word vector.

[0074] Specifically, first, for the alternating text in the feature dialogue, delete all punctuation marks (such as commas, periods, quotation marks, etc.) to clean up the interference of meaningless symbols, so as to obtain plain text data consisting only of words or phrases. The purpose of this step is to simplify the text structure and ensure that subsequent processing is more efficient and focused. Apply natural language processing technology to the plain text data for part-of-speech tagging to identify the part-of-speech attributes of each word, such as nouns, verbs, adjectives, etc. For example, through part-of-speech recognition technology, "blocking word" can be tagged as a noun, "trigger" as a verb, and "potential" as an adjective. The goal of this step is to extract the language attribute features in the text for subsequent word segmentation and semantic analysis. Based on the result of part-of-speech recognition, divide the plain text data into several independent word units or phrase units. For example, the sentence "Blocking words may be triggered" is segmented into three units: "blocking word", "may", and "be triggered". The main purpose of word segmentation processing is to convert continuous text into operable basic language units, laying a foundation for subsequent vectorization operations. Map the word units or phrase units after word segmentation processing to high-dimensional vector representations (i.e., word vectors), and convert the language information in the text into a computable numerical form through a pre-trained word embedding model (such as Word2Vec, GloVe, or BERT, etc.). The result of this step is to generate the entity word vector corresponding to each word unit or phrase unit for the construction and analysis of the subsequent topological network.

[0075] In an embodiment of the present invention, a topological network is constructed based on the entity word vector. The topological network includes topological nodes, topological edges, and edge weights:

[0076] Step 301, obtain the blocking words corresponding to the feature dialogue, and map the blocking words to the topological nodes of the topological network; wherein, the blocking word vector corresponding to the blocking word is used as the feature vector of the corresponding topological node;

[0077] Step 302, map each alternating text in the feature dialogue to the topological nodes of the topological network, and use the entity word vector corresponding to the alternating text as the feature vector of the corresponding topological node;

[0078] Step 303, calculate the correlation degree between the topological nodes of the blocking words and the topological nodes of each alternating text. The calculation formula of the correlation degree is as follows:

[0079] ;

[0080] Among them, represents the correlation degree between the topological nodes of the masked words and the topological nodes of the alternative text, represents the masked word vector of the masked words, represents the number of entity word vectors of the alternative text, represents the index of represents the th entity word vector in the alternative text, represents the th entity word vector in the alternative text and the cosine similarity between the masked word vector of the masked words, represents the operation of obtaining the norm of the vector, represents the weight of the non - linear term, represents the index of represents the th entity word vector and the th entity word vector in the alternative text and the cosine similarity between them, represents the th entity word vector in the alternative text and the cosine similarity between the masked word vector of the masked words, represents the activation function;

[0081] Step 304, construct a topological edge between the topological nodes of the alternative text at the t - th moment and the (t + 1)-th moment, and use the correlation degree between the topological nodes of the alternative text at the t - th moment and the topological nodes of the masked words as the edge weight of the corresponding topological edge.

[0082] Specifically, the construction of the topological network has the following remarkable beneficial effects:

[0083] Enhance the representation ability of the masked word overflow path: The topological network explicitly structures the word units and their semantic relationships in the text in the form of nodes and edges. This structure can more intuitively capture the potential patterns of the masked word overflow path. Compared with simple word vector analysis or context statistical methods, the topological network can reflect higher-level semantic associations, providing a richer and more comprehensive data basis for subsequent path analysis. Improve the accuracy of masked word overflow detection: By constructing a topological network, the explicit masked word overflow paths in the text can be captured. Achieve dynamic modeling in complex scenarios: Regardless of the complexity of the dialogue text, this method can transform the language information in the dialogue text into a unified topological structure, facilitating subsequent clustering and analysis. Provide strong support for path clustering and overflow detection: The constructed topological network is not only used to capture overflow paths but also lays a foundation for subsequent clustering analysis and feature extraction of paths. The weights of each node and edge in the network can accurately quantify the semantic characteristics of word units and their interactions. This high-precision representation improves the reliability of the masked word overflow detection system.

[0084] In one embodiment of the present invention, analyzing each topological network includes:

[0085] Step 401, obtaining the edge weights of any two adjacent topological edges in the topological network;

[0086] Step 402, if the difference between the edge weights is less than or equal to a preset threshold, splicing the corresponding topological edges and deleting the topological nodes between the corresponding topological edges;

[0087] Step 403, if the difference between the edge weights is greater than the preset threshold, keeping the topological edges and topological nodes unchanged;

[0088] Step 404, processing the topological network based on Steps 402 and 403 to obtain the overflow path.

[0089] For example, there is a topological network "E1→E2→E3→E4→E5"; where "→" represents a topological edge, and "E1, E2, E3, E4, E5" all represent topological nodes; among them, the edge weights of the two topological edges in "→E2→" are 0.4 and 0.43 respectively, and the difference is 0.03 which is less than the threshold 0.05, so splicing is performed and E2 is deleted to obtain "E1→E3→E4→E5", and the edge weight of the topological edge in "E1→E3" is 0.415.

[0090] In one embodiment of the present invention, clustering the overflow paths includes:

[0091] Clustering multiple overflow paths of the same masked word to obtain several overflow paths of the corresponding masked word;

[0092] Obtain the edge weight vectors of several overflow paths of the stop words, and the edge weight vectors ; where represents the edge weight of the nth topological edge in the overflow path;

[0093] Based on the edge weight vectors, calculate the clustering degree of any two overflow paths. The calculation formula of the clustering degree is as follows:

[0094] ;

[0095] ;

[0096] ;

[0097] where represents the clustering degree of any two overflow paths, represents the path similarity matrix, represents the element in the gth row and hth column of the path similarity matrix, represents the symmetric matrix of the path similarity matrix, represents the largest eigenvalue of, represents the degree matrix of the path similarity matrix, represents the operation of obtaining the Frobenius, represents , represents the natural base, represents the number of overflow paths corresponding to the stop words, represents the index of, represents the index of, represents the edge weight vector of the gth overflow path, represents the th edge weight vector of the overflow path, represents the transpose operation, represents the operation of obtaining the sum of vector elements;

[0098] If the clustering degree is greater than the preset clustering threshold, then classify the corresponding overflow paths into one category to obtain several categories of overflow paths;

[0099] If the number of overflow paths in one category is greater than the preset number threshold, then use the corresponding category of overflow paths as the characteristic overflow paths.

[0100] For example, for the blocked word "nightclub", "nightclub" has a total of 10 overflow paths. Clustering the 10 overflow paths, if the clustering degrees of overflow path 1 and overflow path 2 are 0.8, and 0.8 is greater than the preset clustering threshold of 0.75, then overflow path 1 and overflow path 2 are grouped into one category. If the clustering degrees of overflow path 3 with overflow path 1 and overflow path 2 are 0.87 and 0.43 respectively, then overflow path 3 and overflow path 1 are grouped into one category. Then two clusters are obtained, one cluster is overflow path 1 and overflow path 2, and one cluster is overflow path 1 and overflow path 3. If the clustering degrees of overflow path 3 with overflow path 1 and overflow path 2 are 0.87 and 0.77 respectively, then one cluster is obtained, and the cluster includes overflow path 1, overflow path 2, and overflow path 3. By extracting the characteristics of the overflow paths and performing clustering analysis, representative characteristic overflow paths are summarized, improving the systematicness and efficiency of overflow detection.

[0101] In one embodiment of the present invention, generating an induced dialogue for a language model includes:

[0102] Calculating the mean vector of the edge weight vectors corresponding to the overflow paths among several characteristic overflow paths, and the calculation formula of the mean vector is as follows:

[0103] ;

[0104] ;

[0105] Wherein, represents the mean vector of several characteristic overflow paths, represents the nth element in the mean vector, represents the number of the nth element of the mean vector of several characteristic overflow paths, represents the index of, represents the nth element of the kth mean vector;

[0106] Based on the mean vector of several characteristic overflow paths, constructing an induced dialogue.

[0107] Specifically, if the edge weights of the topological edges of "E1→E3→E4→E5" are "0.415, 0.46, 0.65, 0.82" respectively, then the edge weight vector is . Calculating the mean vector based on the edge weight vector. For example, the characteristic overflow paths of the blocked word include 2 overflow paths, and the edge weight vectors are respectively and ; then the mean vector is .

[0108] In one embodiment of the present invention, constructing an induced dialogue includes:

[0109] Step 701: Obtain the mean vectors corresponding to several characteristic overflow paths of the shielding words;

[0110] Step 702: Take the average of the first elements in all the mean vectors as the correlation degree between the first text of the induced dialogue and the shielding words, and randomly generate the first text;

[0111] Step 703: Obtain the second text of the output of the language model for the first text, calculate the correlation degree between the second text and the shielding words; and match the second elements in all the mean vectors based on the correlation degree, specifically including: if the difference between the correlation degree between the second text and the shielding words and the second element of the s-th mean vector is less than or equal to the preset difference, then the s-th mean vector is classified into the first induced set;

[0112] Step 704: Take the average of the third elements in the first induced set as the correlation degree between the third text of the induced dialogue and the shielding words, and randomly generate the third text;

[0113] Step 705: Obtain the fourth text of the output of the language model for the third text, calculate the correlation degree between the fourth text and the shielding words; and match the fourth elements in all the mean vectors based on the correlation degree, specifically including: if the difference between the correlation degree between the fourth text and the shielding words and the fourth element of the s-th mean vector is less than or equal to the preset difference, then the s-th mean vector is classified into the second induced set;

[0114] Step 706: Repeat Step 704 and Step 705 until the u-th text or the r-th text is obtained; where the u-th text represents the text with shielding words output by the language model, and the r-th text represents the dimension of the mean vectors in the (r - 1)-th induced set;

[0115] Step 707: If there are shielding words in any one of the u-th text or the r-th text, obtain the vulnerability detection result.

[0116] Specifically, based on 4 mean vectors of the shielding words, construct an induced dialogue as follows:

[0117] ;

[0118] ;

[0119] Establish the first text: Determine that the correlation degree between the first text and the shielding words is ; Generate the first text based on the correlation degree, which can be constructed by experts or a third-party language model.

[0120] Use the first text as the input of the user, and then obtain the second text feedback by the language model.

[0121] Calculate the correlation between the second text and the masked word. If the correlation is 0.44, then the second mean vector, the third mean vector, and the fourth mean vector are all used as the first induced set.

[0122] Establish the third text and determine that the correlation between the third text and the masked word is ; Generate the third text based on the correlation.

[0123] If the correlation between the fourth text feedback by the language model and the masked word is then the second induced set includes the second mean vector and the third mean vector.

[0124] Establish the fifth text and determine that the correlation between the fifth text and the masked word is ; Generate the fifth text based on the correlation.

[0125] If the sixth text has a masked word, then a detection of the vulnerability is obtained.

[0126] A method for detecting vulnerability of masked word overflow in a generative language model, which is applied to the system for detecting vulnerability of masked word overflow in the generative language model, includes:

[0127] Step 1: Obtain the characteristic conversations of masked word overflow of several users; wherein, the characteristic conversations include several rounds of alternating texts between users and the language model.

[0128] Step 2: Preprocess each characteristic conversation to obtain several entity word vectors.

[0129] Step 3: Construct a topological network based on the entity word vectors.

[0130] Step 4: Analyze each topological network to obtain the overflow path of the masked word.

[0131] Step 5: Cluster the overflow paths to obtain the characteristic overflow paths.

[0132] Step 6: Generate induced conversations for the language model based on the characteristic overflow paths; generate the vulnerability detection result of the masked word according to the induced conversations.

[0133] The above has described the embodiments of this example, but this example is not limited to the above specific implementation manners. The above specific implementation manners are only illustrative and not restrictive. Under the inspiration of this example, those of ordinary skill in the art can also make many forms, all of which fall within the protection scope of this example.

Claims

1. A vulnerability detection system for overflow of blocked words in a generative language model, characterized in that: include: A data collection module is used to obtain characteristic dialogues of blocked word overflows of several users; wherein the characteristic dialogues include alternating texts of several rounds of users and language models; The text processing module is used to pre-process each feature dialogue to obtain several entity word vectors; Topology building module, used to build a topology network based on entity word vectors; The topological analysis module is used to analyze each topological network and obtain the overflow path of the blocked words; A path clustering module is used to cluster the overflow paths and obtain characteristic overflow paths; The overflow test module is used to generate induced dialogues for the language model based on the feature overflow path; and to generate vulnerability detection results of blocked words based on the induced dialogues.

2. According to claim 1, a generative language model shielding word overflow vulnerability detection system is characterized in that: Preprocess each feature dialogue, including: Remove punctuation marks from alternate text in feature dialogues to obtain plain text data; Perform part-of-speech recognition on plain text data, and use natural language processing technology to annotate the part-of-speech attributes of each word, including nouns, verbs, and adjectives, to obtain part-of-speech recognition results; Based on the part-of-speech recognition results, the plain text data is segmented into several independent word units or phrase units; Map word units or phrase units to corresponding word vector representations to obtain entity word vectors.

3. According to claim 2, a generative language model shielding word overflow vulnerability detection system is characterized in that: A topological network is constructed based on entity word vectors. The topological network includes topological nodes, topological edges, and edge weights: Step 301, obtaining the shielded words corresponding to the feature dialogue, and mapping the shielded words to the topological nodes of the topological network; wherein the shielded word vector corresponding to the shielded word is used as the feature vector of the corresponding topological node; Step 302, mapping each alternate text in the feature dialogue to a topological node of the topological network, and using the entity word vector corresponding to the alternate text as the feature vector of the corresponding topological node; Step 303, calculate the correlation between the topological node of the blocked word and the topological node of each alternate text, and the calculation formula of the correlation is as follows: ; in, The correlation between the topological node of the blocked word and the topological node of the alternate text, The masked word vector representing the masked word, The number of entity word vectors representing alternate texts, express The index of Indicates the first entity word vectors, Indicates the first The cosine similarity between the entity word vector and the shielded word vector of the shielded word, represents the operation of finding the modulus length of the orientation vector, represents the weight of nonlinear term, express The index of Indicates the first entity word vector and The cosine similarity of entity word vectors, Indicates the first The cosine similarity between the entity word vector and the shielded word vector of the shielded word, express Activation function; Step 304, construct a topological edge between the topological nodes of the alternating text at the tth moment and the t+1th moment, and use the correlation between the topological node of the alternating text at the tth moment and the topological node of the blocked word as the edge weight of the corresponding topological edge.

4. According to claim 3, a generative language model shielding word overflow vulnerability detection system is characterized in that: Each topology network is analyzed, including: Step 401, obtaining edge weights of any two adjacent topological edges in the topological network; Step 402: if the difference between the edge weights is less than or equal to a preset threshold, the corresponding topological edges are spliced, and the topological nodes between the corresponding topological edges are deleted; Step 403: if the difference between the edge weights is greater than a preset threshold, the topological edge and the topological node are kept unchanged; Step 404: Process the topology network based on steps 402 and 403 to obtain an overflow path.

5. According to claim 4, a generative language model shielding word overflow vulnerability detection system is characterized in that: Clustering of overflow paths, including: Clustering multiple overflow paths of the same blocked word to obtain several overflow paths of the corresponding blocked word; Get the edge weight vectors of several overflow paths of the blocked words, the edge weight vector ;in, Represents the edge weight of the nth topological edge in the overflow path; Based on the edge weight vector, the clustering degree of any two overflow paths is calculated. The calculation formula of the clustering degree is as follows: ; ; ; in, represents the clustering degree of any two overflow paths, represents the path similarity matrix, represents the element in the gth row and hth column in the path similarity matrix, A symmetric matrix representing the path similarity matrix, express The maximum eigenvalue of The degree matrix representing the path similarity matrix, represents the operation of obtaining Frobenius. express , represents the natural base, Indicates the number of overflow paths corresponding to the blocked words, express The index of express The index of represents the edge weight vector of the g-th overflow path, Indicates The edge weight vector of the overflow path, represents the transpose operation, represents the operation of finding the sum of vector elements; If the clustering degree is greater than the preset clustering threshold, the corresponding overflow paths are classified into one category to obtain several categories of overflow paths; If the number of overflow paths of a type is greater than a preset number threshold, the corresponding overflow path of the type is used as a characteristic overflow path.

6. According to claim 5, a generative language model shielding word overflow vulnerability detection system is characterized in that: Generate induced dialogue for the language model, including: Calculate the mean vector of the edge weight vectors corresponding to the overflow paths in several feature overflow paths. The calculation formula of the mean vector is as follows: ; ; in, represents the mean vector of several feature overflow paths, represents the nth element in the mean vector, Represents the number of n-th elements of the mean vector of several feature overflow paths, express The index of represents the nth element of the kth mean vector; Based on the mean vector of several feature overflow paths, an induced dialogue is constructed.

7. The system for detecting a vulnerability of a generative language model for overflow of blocked words according to claim 6, characterized in that: Construct an inducing dialogue that includes: Step 701, obtaining mean vectors corresponding to several feature overflow paths of the blocked words; Step 702, taking the average of the first elements in all mean vectors as the correlation between the first text of the induced dialogue and the blocked words, and randomly generating the first text; Step 703, obtaining a second text output by the language model for the first text, calculating the relevance between the second text and the blocked word; and matching the second elements in all mean vectors based on the relevance, specifically including: if the difference between the relevance between the second text and the blocked word and the second element of the s-th mean vector is less than or equal to a preset difference, then the s-th mean vector is classified into the first induced set; Step 704, taking the average value of the third element of the first induced set as the correlation between the third text of the induced dialogue and the blocked word, and randomly generating the third text; Step 705, obtaining a fourth text output by the language model for the third text, calculating the relevance between the fourth text and the blocked word; and matching the fourth elements in all mean vectors based on the relevance, specifically including: if the difference between the relevance between the fourth text and the blocked word and the fourth element of the s-th mean vector is less than or equal to a preset difference, then the s-th mean vector is classified into the second induced set; Step 706, repeating steps 704 and 705 until the uth text or the rth text is obtained; wherein the uth text represents the text output by the language model with blocked words, and the rth text represents the dimension of the mean vector in the r-1th induced set; Step 707: If the blocked word exists in any of the u-th text or the r-th text, a vulnerability detection result is obtained.

8. A method for detecting a vulnerability of a generative language model for overflow of a shielding word, applied to a system for detecting a vulnerability of a generative language model for overflow of a shielding word as claimed in any one of claims 1 to 7, characterized in that: include: Step 1, obtaining characteristic dialogues of blocked word overflow of several users; wherein the characteristic dialogues include alternating texts of several rounds of users and language models; Step 2: Preprocess each feature dialogue to obtain several entity word vectors; Step 3, construct a topological network based on entity word vectors; Step 4, analyze each topological network to obtain the overflow path of the blocked words; Step 5, clustering the overflow paths to obtain characteristic overflow paths; Step 6: Generate an induced dialogue for the language model based on the feature overflow path; and generate a vulnerability detection result of the blocked words based on the induced dialogue.

Citation Information

Patent Citations

  • Buffer overflow loophole automatic detection method based on symbolic execution path pruning

    CN104732152A

  • Data security analysis method and intelligent calculation data security workstation

    CN119249440A